Conversation
A streamed response that stops at `max_tokens` ended in a `RUN_ERROR` with no `usage`. The `max_tokens` branch in `processAnthropicStream` did not call `buildAnthropicUsage(event.usage)`, while the `tool_use` and default branches did. Anthropic bills these tokens, so code that counts `usage` undercounted every call that ended this way. Attach `usage` to that `RUN_ERROR`. Core already carries `usage` on `RUN_ERROR` to the stream and to `onChunk`. Fixes TanStack#1597 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… omits them `processAnthropicStream` built usage from the closing `message_delta` only. The SDK types its input and cache counts as nullable, and some Anthropic-compatible servers send only `output_tokens` there. aimock is one of them, so against aimock every stream reported 0 input tokens. Keep the `message_start` usage, and take each input or cache count that the delta leaves null or out from it. Counts that the delta sends still win, because they are cumulative. This is the same merge that the SDK's MessageStream does. It applies to every stop reason, so the `max_tokens` `RUN_ERROR` from the previous commit carries the input count too. The E2E spec drives a `max_tokens` stop through aimock and needs both commits to pass. Refs TanStack#1597 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The SDK's MessageStream also keeps `server_tool_use` from `message_start` when the closing `message_delta` has it null, and `buildAnthropicUsage` reads it. Fill it the same way as the input and cache counts. The SDK carries `iterations` too, but the adapter does not read it, so it stays as the delta sends it. Fold the two changesets of this branch into one. Refs TanStack#1597 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
🦋 Changeset detectedLatest commit: e3e19b8 The changes in this PR will be included in the next version bump. This PR includes changesets to release 1 package
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (1)
Included review availability: This review used your included allowance. Your plan provides up to 8 included reviews per hour; 7 remain after this review. 📝 WalkthroughWalkthroughThe Anthropic adapter now merges usage from stream start and delta events for terminal chunks. A max-token ChangesAnthropic stream usage
Priority: ➖ Normal Estimated code review effort: 3 (Moderate) | ~20 minutes Change: Bug fix · Severity of issue fixed: Medium Merge Risk: ⚪ Minimal · up to Supported Anthropic streams retain available usage in max-token error chunks, and the added end-to-end coverage checks the reported totals. No concrete merge-blocking risk remains; the change is ready for normal checks. Security Architecture ReviewSecurity architecture risk: 🔵 Low · up to The change improves token reporting without changing terminal error handling or provider permissions. The new endpoint uses a fixed test request and dummy credential. No introduced security concern was identified, but its exposure outside the local test setup is not established. Retained concerns Security review detailsSecurity Blast Radius
Trust Boundaries and Controls
Resilience and Maintainability Implications
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Out of Scope Changes checkExplanation Changes since the previous review add unrelated work with no connection to [
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Thanks for the PR, @ArVaViT! 🙌 @AlemTuzlak will take a look. Automated pre-review checks
Automated triage — a human review follows. |
…ropic-max-tokens-usage # Conflicts: # testing/e2e/src/routeTree.gen.ts
A streamed Anthropic call that stops at
max_tokensended in aRUN_ERRORwith nousage, so code that counts tokens missed every such call. This PR putsusageon thatRUN_ERROR. It also keeps the input, cache, and server tool counts frommessage_startwhen the closingmessage_deltaleaves them out, as the Anthropic SDK does. aimock is one server that leaves them out, so without that part the newusagestill said 0 input tokens.🎯 Changes
max_tokensstop (commit 1). Themax_tokensbranch ofprocessAnthropicStreamnow setsusageon itsRUN_ERROR, the same as thetool_useand default branches set it onRUN_FINISHED.message_startcounts (commits 2 and 3). The stream handler keeps the usage ofmessage_start. If the closingmessage_deltahas no input, cache, or server tool count, the handler takes that count frommessage_start. The SDK'sMessageStreamdoes the same for these counts. It also carriesiterations, which the adapter does not read. A count on the delta still wins, because the delta counts are cumulative. The new helpermergeAnthropicStreamUsageinsrc/usage.tsdoes this. It is not exported from the package.docs/adapters/anthropic.mdsays that themax_tokensRUN_ERRORcarriesusage, and thatonUsagedoes not fire for it.docs/config.jsonhas the newupdatedAt.packages/ai-anthropic/tests/usage-extraction.test.ts, and the E2E specanthropic-max-tokens-usage.spec.tswith its aimock fixture and route.@tanstack/ai-anthropicpatch changeset with one paragraph for each fix.✅ Checklist
pnpm run test:pr, or these tests do not apply to this pull request.docs/for this change, or this change is not user-facing.pnpm changeset), or this PR does not change a published package.Not ticked:
pnpm test:pr: not run. I ran the checks of@tanstack/ai-anthropicand the Anthropic E2E specs (see Testing). CI runs the full set.🚀 Release Impact
Root cause
Issue. A streamed Anthropic call that stops at
max_tokensends inRUN_ERRORwithcode: 'max_tokens'. That chunk had nousage, but Anthropic bills the call. Against servers that send onlyoutput_tokenson the closingmessage_delta, every stop reason also reported 0 input and cache tokens, and no server tool counts.Cause.
processAnthropicStream, themax_tokenscase of themessage_deltaswitch did not callbuildAnthropicUsage(event.usage). Thetool_useand default cases did.message_deltausage only.buildAnthropicUsagereads a missinginput_tokensas 0. The SDK types the input, cache, and server tool counts on the delta as nullable, and the SDK's ownMessageStreamkeeps themessage_startvalues for them.Fix.
max_tokenscase setsusageon itsRUN_ERROR. Core already keepsusageonRUN_ERROR: it is a spec key for that event, andchat()passes it to the stream and toonChunk.message_startusage. Before it buildsusagefor any stop reason, it fills each null or missing input, cache, or server tool count on the delta frommessage_start.Possible alternatives
max_tokensstop withRUN_FINISHEDandfinishReason: 'length'. Bedrock Converse and OpenAI Chat Completions do this, andonUsagewould then fire. But it changes the event that Anthropic users get today (RUN_ERROR,code: 'max_tokens'), so it is a larger change than this bug needs.onUsagefor aRUN_ERRORwithusage. The docs sayonUsagefires forRUN_FINISHED. The subagent path already puts the usage of child runs on a failed parentRUN_ERROR, so this needs its own check for double counts. That is a core change for its own issue.message_startmerge (commits 2 and 3). The Anthropic API repeats the input counts on the closing delta, so commit 1 alone works against it. But aimock, the mock that this repo's E2E suite uses, sends onlyoutput_tokensthere. With commit 1 only, the E2E spec getspromptTokens: 0(transcript below). Commits 2 and 3 can be dropped together if you want a smaller PR.Testing
Gate 1 repro (agent-written). I wrote the unit tests and the E2E spec. I ran them on
main(d31e4ebb9) with only the test files added, then on this branch.Unit tests,
packages/ai-anthropic/tests/usage-extraction.test.ts, onmain:On this branch:
E2E,
testing/e2e/tests/anthropic-max-tokens-usage.spec.tsagainst aimock. Onmain:With commit 1 only:
With all commits:
1 passed.Mutation check. I removed each part of the fix in turn and ran the unit tests. Each of the 15 mutations made at least one test fail:
usageon themax_tokensRUN_ERROR.message_startusage, or return the delta usage unchanged.input_tokens,cache_creation_input_tokens,cache_read_input_tokens, orserver_tool_use. Or letmessage_startwin over the delta for one of them.tool_use,max_tokens, or default case.Commands run (macOS, one at a time):
vitest runinpackages/ai-anthropic: 173 passed.tsc(test:types) andoxlint src --type-aware(test:oxlint) inpackages/ai-anthropic: pass, no new warnings.publint --strict: pass.tsc --noEmitintesting/e2e: pass.tests/anthropic-*.spec.ts(8 files): 16 passed.pnpm test:prand the full E2E suite.Manual test.
main, runchat()withcreateAnthropicChat('claude-haiku-4-5', key),modelOptions: { max_tokens: 3 }, and the prompt "Say hello in five words.". Print theRUN_ERRORchunk. It has nousage.RUN_ERRORhasusage, for example{ promptTokens: 13, completionTokens: 3, totalTokens: 16 }.pnpm --filter @tanstack/ai-e2e test:e2e -- tests/anthropic-max-tokens-usage.spec.ts. Expect 1 passed.How this PR makes testing easy. The unit tests in
usage-extraction.test.tscover each stop reason. The E2E route/api/anthropic-max-tokens-usageand the aimock fixturefixtures/max-tokens-usage/basic.jsonneed no API key.Linked issues
Closes #1597
Risk / rollback
message_startvalue. This is the correct count, but a dashboard can show a jump.usageonly onRUN_FINISHED, or only throughonUsage, still does not see themax_tokensusage. The docs say where to read it.Not in this PR:
structuredOutput()path reads the fullresponse.usage, so it has no gap. On amax_tokenscut it throws, and the usage is lost there. fix(ai-anthropic, ai-gemini, ai-mistral, ai-ollama, ai-bedrock): report max-token truncation in structuredOutput #1548 tracks that truncation.anthropicVertexTextandanthropicSummarizeuse the sameAnthropicTextAdapter, so this fix covers them. Bedrock Converse has its own stream code and already reports usage onmax_tokens.usageon aRUN_FINISHEDafter itsMAX_TOKENSRUN_ERROR, andchat()stops reading at theRUN_ERROR.🤖 Generated with Claude Code
Summary by CodeRabbit
Bug Fixes
RUN_ERRORresult, including input, cache, and server-tool counts when available.Documentation
onUsageruns only forRUN_FINISHED; token-limit error usage is available from theRUN_ERRORchunk or anonChunkmiddleware.