Skip to content

Cap DeepSeek V4 Flash max output at 262,144 on Posit AI Pass - #125

Merged
wch merged 1 commit into
mainfrom
ds-v4-flash-update
Sep 29, 2026
Merged

wch merged 1 commit into
mainfrom
ds-v4-flash-update

Conversation

@wch

@wch wch commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator

Using DeepSeek V4 Flash through Posit AI Pass fails with:

Error making request (HTTP 400): Invalid request: ['max_tokens (384000): Input should be less than or equal to 262144']

The Model APIs capability entry for deepseek-ai/DeepSeek-V4-Flash-0731 used maxOutputTokens: 384_000, which is DeepSeek's native API limit. Baseten Model APIs caps this model at 262,144, so this lowers the entry to 262_144 and adds a comment explaining the discrepancy.

The direct DeepSeek provider entries in deepseek-helpers.ts are unchanged, since they target DeepSeek's own API where 384K is the documented limit.

ai-config model-capabilities tests pass.

Baseten Model APIs rejects max_tokens above 262,144 for
deepseek-ai/DeepSeek-V4-Flash-0731, so the 384K limit copied from
DeepSeek's native API caused HTTP 400 errors.
@wch
wch enabled auto-merge (squash) September 29, 2026 21:50
@wch
wch merged commit ee84da0 into main Sep 29, 2026
4 checks passed
@wch
wch deleted the ds-v4-flash-update branch September 29, 2026 21:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant