Skip to content

fix(ai-anthropic): report token usage when a stream stops at max_tokens - #1609

Open
ArVaViT wants to merge 4 commits into
TanStack:mainfrom
ArVaViT:fix/anthropic-max-tokens-usage
Open

ArVaViT wants to merge 4 commits into
TanStack:mainfrom
ArVaViT:fix/anthropic-max-tokens-usage

Conversation

@ArVaViT

@ArVaViT ArVaViT commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

A streamed Anthropic call that stops at max_tokens ended in a RUN_ERROR with no usage, so code that counts tokens missed every such call. This PR puts usage on that RUN_ERROR. It also keeps the input, cache, and server tool counts from message_start when the closing message_delta leaves them out, as the Anthropic SDK does. aimock is one server that leaves them out, so without that part the new usage still said 0 input tokens.

🎯 Changes

  • Usage on the max_tokens stop (commit 1). The max_tokens branch of processAnthropicStream now sets usage on its RUN_ERROR, the same as the tool_use and default branches set it on RUN_FINISHED.
  • Keep the message_start counts (commits 2 and 3). The stream handler keeps the usage of message_start. If the closing message_delta has no input, cache, or server tool count, the handler takes that count from message_start. The SDK's MessageStream does the same for these counts. It also carries iterations, which the adapter does not read. A count on the delta still wins, because the delta counts are cumulative. The new helper mergeAnthropicStreamUsage in src/usage.ts does this. It is not exported from the package.
  • Docs. docs/adapters/anthropic.md says that the max_tokens RUN_ERROR carries usage, and that onUsage does not fire for it. docs/config.json has the new updatedAt.
  • Tests. Six new unit tests in packages/ai-anthropic/tests/usage-extraction.test.ts, and the E2E spec anthropic-max-tokens-usage.spec.ts with its aimock fixture and route.
  • Changeset. One @tanstack/ai-anthropic patch changeset with one paragraph for each fix.

✅ Checklist

  • I have followed the steps in the Contributing guide.
  • I have tested code changes locally with pnpm run test:pr, or these tests do not apply to this pull request.
  • I fully understand the code in this pull request, including any code generated with AI assistance.
  • Docs: I updated docs/ for this change, or this change is not user-facing.
  • Changeset: I added a changeset (pnpm changeset), or this PR does not change a published package.

Not ticked:

  • pnpm test:pr: not run. I ran the checks of @tanstack/ai-anthropic and the Anthropic E2E specs (see Testing). CI runs the full set.

🚀 Release Impact

  • This change affects published code, and I have generated a changeset.
  • This change is docs/CI/dev-only (no release).

Root cause

Issue. A streamed Anthropic call that stops at max_tokens ends in RUN_ERROR with code: 'max_tokens'. That chunk had no usage, but Anthropic bills the call. Against servers that send only output_tokens on the closing message_delta, every stop reason also reported 0 input and cache tokens, and no server tool counts.

Cause.

  • In processAnthropicStream, the max_tokens case of the message_delta switch did not call buildAnthropicUsage(event.usage). The tool_use and default cases did.
  • The handler built usage from the message_delta usage only. buildAnthropicUsage reads a missing input_tokens as 0. The SDK types the input, cache, and server tool counts on the delta as nullable, and the SDK's own MessageStream keeps the message_start values for them.

Fix.

  • The max_tokens case sets usage on its RUN_ERROR. Core already keeps usage on RUN_ERROR: it is a spec key for that event, and chat() passes it to the stream and to onChunk.
  • The handler keeps the message_start usage. Before it builds usage for any stop reason, it fills each null or missing input, cache, or server tool count on the delta from message_start.

Possible alternatives

  • End a max_tokens stop with RUN_FINISHED and finishReason: 'length'. Bedrock Converse and OpenAI Chat Completions do this, and onUsage would then fire. But it changes the event that Anthropic users get today (RUN_ERROR, code: 'max_tokens'), so it is a larger change than this bug needs.
  • Make core fire onUsage for a RUN_ERROR with usage. The docs say onUsage fires for RUN_FINISHED. The subagent path already puts the usage of child runs on a failed parent RUN_ERROR, so this needs its own check for double counts. That is a core change for its own issue.
  • Leave out the message_start merge (commits 2 and 3). The Anthropic API repeats the input counts on the closing delta, so commit 1 alone works against it. But aimock, the mock that this repo's E2E suite uses, sends only output_tokens there. With commit 1 only, the E2E spec gets promptTokens: 0 (transcript below). Commits 2 and 3 can be dropped together if you want a smaller PR.

Testing

Gate 1 repro (agent-written). I wrote the unit tests and the E2E spec. I ran them on main (d31e4ebb9) with only the test files added, then on this branch.

Unit tests, packages/ai-anthropic/tests/usage-extraction.test.ts, on main:

× reports usage on the RUN_ERROR of a max_tokens stop (#1597)
× keeps the message_start counts when the end_turn message_delta sends only output_tokens
× keeps the message_start counts when the tool_use message_delta sends only output_tokens
× keeps the message_start counts when the max_tokens message_delta sends only output_tokens
AssertionError: expected undefined to deeply equal { promptTokens: 13, …(3) }
AssertionError: expected { promptTokens: +0, …(2) } to deeply equal { promptTokens: 13, …(3) }
Tests  4 failed | 7 passed (11)

On this branch:

✓ reports usage on the RUN_ERROR of a max_tokens stop (#1597)
✓ keeps the message_start counts when the end_turn message_delta sends only output_tokens
✓ keeps the message_start counts when the tool_use message_delta sends only output_tokens
✓ keeps the message_start counts when the max_tokens message_delta sends only output_tokens
✓ takes the message_delta counts over message_start when both are present
✓ omits usage when the message_delta reports none
Tests  11 passed (11)

E2E, testing/e2e/tests/anthropic-max-tokens-usage.spec.ts against aimock. On main:

Expected: {"completionTokens": 3, "promptTokens": 13, "totalTokens": 16}
Received: undefined
1 failed

With commit 1 only:

  Object {
    "completionTokens": 3,
-   "promptTokens": 13,
-   "totalTokens": 16,
+   "promptTokens": 0,
+   "totalTokens": 3,
  }
1 failed

With all commits: 1 passed.

Mutation check. I removed each part of the fix in turn and ran the unit tests. Each of the 15 mutations made at least one test fail:

  • Drop usage on the max_tokens RUN_ERROR.
  • Never keep the message_start usage, or return the delta usage unchanged.
  • Drop the fallback for input_tokens, cache_creation_input_tokens, cache_read_input_tokens, or server_tool_use. Or let message_start win over the delta for one of them.
  • Use the unmerged usage in the tool_use, max_tokens, or default case.
  • Drop the guard for a delta with no usage.

Commands run (macOS, one at a time):

  1. vitest run in packages/ai-anthropic: 173 passed.
  2. tsc (test:types) and oxlint src --type-aware (test:oxlint) in packages/ai-anthropic: pass, no new warnings. publint --strict: pass.
  3. oxfmt check of the changed files: pass. tsc --noEmit in testing/e2e: pass.
  4. Playwright, tests/anthropic-*.spec.ts (8 files): 16 passed.
  5. Not run: pnpm test:pr and the full E2E suite.

Manual test.

  1. On main, run chat() with createAnthropicChat('claude-haiku-4-5', key), modelOptions: { max_tokens: 3 }, and the prompt "Say hello in five words.". Print the RUN_ERROR chunk. It has no usage.
  2. On this branch, run the same code. The RUN_ERROR has usage, for example { promptTokens: 13, completionTokens: 3, totalTokens: 16 }.
  3. Run pnpm --filter @tanstack/ai-e2e test:e2e -- tests/anthropic-max-tokens-usage.spec.ts. Expect 1 passed.

How this PR makes testing easy. The unit tests in usage-extraction.test.ts cover each stop reason. The E2E route /api/anthropic-max-tokens-usage and the aimock fixture fixtures/max-tokens-usage/basic.json need no API key.

Linked issues

Closes #1597

Risk / rollback

  • Usage from Anthropic-compatible servers that leave input counts off the delta changes from 0 to the message_start value. This is the correct count, but a dashboard can show a jump.
  • Code that reads usage only on RUN_FINISHED, or only through onUsage, still does not see the max_tokens usage. The docs say where to read it.
  • Rollback: revert the PR. Commits 2 and 3 can also be reverted together.

Not in this PR:

  • The non-streaming structuredOutput() path reads the full response.usage, so it has no gap. On a max_tokens cut it throws, and the usage is lost there. fix(ai-anthropic, ai-gemini, ai-mistral, ai-ollama, ai-bedrock): report max-token truncation in structuredOutput #1548 tracks that truncation.
  • anthropicVertexText and anthropicSummarize use the same AnthropicTextAdapter, so this fix covers them. Bedrock Converse has its own stream code and already reports usage on max_tokens.
  • By code reading (not run), the Gemini text adapter has a similar gap: it puts usage on a RUN_FINISHED after its MAX_TOKENS RUN_ERROR, and chat() stops reading at the RUN_ERROR.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Bug Fixes

    • Streamed responses that reach Anthropic’s token limit now include the call’s token usage in the RUN_ERROR result, including input, cache, and server-tool counts when available.
    • Usage counts from the closing stream event take precedence when present.
  • Documentation

    • Clarified that onUsage runs only for RUN_FINISHED; token-limit error usage is available from the RUN_ERROR chunk or an onChunk middleware.

ArVaViT and others added 3 commits October 2, 2026 18:52
A streamed response that stops at `max_tokens` ended in a `RUN_ERROR`
with no `usage`. The `max_tokens` branch in `processAnthropicStream`
did not call `buildAnthropicUsage(event.usage)`, while the `tool_use`
and default branches did. Anthropic bills these tokens, so code that
counts `usage` undercounted every call that ended this way.

Attach `usage` to that `RUN_ERROR`. Core already carries `usage` on
`RUN_ERROR` to the stream and to `onChunk`.

Fixes TanStack#1597

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… omits them

`processAnthropicStream` built usage from the closing `message_delta`
only. The SDK types its input and cache counts as nullable, and some
Anthropic-compatible servers send only `output_tokens` there. aimock is
one of them, so against aimock every stream reported 0 input tokens.

Keep the `message_start` usage, and take each input or cache count that
the delta leaves null or out from it. Counts that the delta sends still
win, because they are cumulative. This is the same merge that the SDK's
MessageStream does. It applies to every stop reason, so the
`max_tokens` `RUN_ERROR` from the previous commit carries the input
count too.

The E2E spec drives a `max_tokens` stop through aimock and needs both
commits to pass.

Refs TanStack#1597

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The SDK's MessageStream also keeps `server_tool_use` from
`message_start` when the closing `message_delta` has it null, and
`buildAnthropicUsage` reads it. Fill it the same way as the input and
cache counts. The SDK carries `iterations` too, but the adapter does
not read it, so it stays as the delta sends it.

Fold the two changesets of this branch into one.

Refs TanStack#1597

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@changeset-bot

changeset-bot Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: e3e19b8

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 1 package
Name Type
@tanstack/ai-anthropic Patch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@coderabbitai

coderabbitai Bot commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: TanStack/ai/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: f1512bc2-9b39-451c-9bd2-e8a35d94fe01
📥 Commits

Reviewing files that changed from the base of the PR and between bb936b5 and e3e19b8.

📒 Files selected for processing (1)
  • testing/e2e/src/routeTree.gen.ts

Included review availability: This review used your included allowance. Your plan provides up to 8 included reviews per hour; 7 remain after this review.


📝 Walkthrough

Walkthrough

The Anthropic adapter now merges usage from stream start and delta events for terminal chunks. A max-token RUN_ERROR can include usage. Unit and end-to-end tests cover this behavior.

Changes

Anthropic stream usage

Layer / File(s) Summary
Merge and emit terminal usage
packages/ai-anthropic/src/usage.ts, packages/ai-anthropic/src/adapters/text.ts, packages/ai-anthropic/tests/usage-extraction.test.ts, docs/adapters/anthropic.md, .changeset/anthropic-max-tokens-usage.md, docs/config.json
The adapter merges message_start and message_delta usage for terminal events. Delta values take precedence, with start values retained when delta fields are absent or null. Tests cover max-token errors, other stop reasons, and missing delta usage. The documentation and changeset describe the usage behavior.
Layer / File(s) Summary
Verify max-token usage end to end
testing/e2e/fixtures/max-tokens-usage/*, testing/e2e/src/routes/api.anthropic-max-tokens-usage.ts, testing/e2e/src/routeTree.gen.ts, testing/e2e/tests/anthropic-max-tokens-usage.spec.ts
A new route streams a max-token response and returns the run error code and usage. The end-to-end test checks for max_tokens and usage totals of 13 prompt tokens, 3 completion tokens, and 16 total tokens.

Priority: ➖ Normal

Estimated code review effort: 3 (Moderate) | ~20 minutes

Change: Bug fix · Severity of issue fixed: Medium

Merge Risk: ⚪ Minimal · up to e3e19

Supported Anthropic streams retain available usage in max-token error chunks, and the added end-to-end coverage checks the reported totals. No concrete merge-blocking risk remains; the change is ready for normal checks.

Security Architecture Review

Security architecture risk: 🔵 Low · up to e3e19

The change improves token reporting without changing terminal error handling or provider permissions. The new endpoint uses a fixed test request and dummy credential. No introduced security concern was identified, but its exposure outside the local test setup is not established.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • inferred — In the checked configuration, the new endpoint adds reachability to a fixed local mock call, not to tenant data or a real-provider credential. If the E2E application is exposed elsewhere, callers can repeat that fixed call; the effective service exposure then depends on deployment and LLMOCK_URL configuration that was not established here.

Trust Boundaries and Controls

  • observed — The POST handler contains no authentication check, but it also accepts no request-controlled prompt, provider URL, model, credential, or token limit. Requesters can initiate the probe; only deployment configuration selects its destination. This counterevidence rejects a request-controlled privileged-call concern for the inspected handler.

Resilience and Maintainability Implications

  • observed — Attaching usage does not convert truncation into success or alter core early termination. RUN_ERROR still records finalization failure, while usage remains diagnostic/accounting data rather than a new retry or authorization decision.
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Out of Scope Changes check ⚠️ Warning Changes since the previous review add unrelated work with no connection to [#1597]. Examples include response-body cancellation documentation in docs/chat/connection-adapters.md and `docs/media/gene… Remove the unrelated documentation and example package changes from this pull request, or move them to a separate pull request linked to their own objectives.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly summarizes the main change: reporting token usage when an Anthropic stream stops at max_tokens.
Description check ✅ Passed The description covers the changes, checklist, release impact, testing, root cause, alternatives, linked issue, and risks. It explains that pnpm test:pr and the full E2E suite were not run, while list…
Linked Issues check ✅ Passed [#1597] The previously reviewed implementation adds Anthropic token usage to the max_tokens RUN_ERROR. It also falls back to message_start input, cache, and server-tool counts when `message_delt…
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 3 functions across 6 files.
Full details: Out of Scope Changes check

Explanation

Changes since the previous review add unrelated work with no connection to [#1597]. Examples include response-body cancellation documentation in docs/chat/connection-adapters.md and docs/media/generations.md, batching documentation in docs/resumable-streams/advanced.md, cancellation documentation in docs/sandbox/providers.md, a server-tools documentation change, and dependency/version updates across example packages.

  • Fix all pre-merge checks with AI
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Thanks for the PR, @ArVaViT! 🙌 @AlemTuzlak will take a look.

Automated pre-review checks

  • ✅ CI passing
  • ✅ No merge conflicts
  • ✅ Changeset present
  • ✅ E2E test changes included

Automated triage — a human review follows.

@github-actions github-actions Bot added the waiting-on: maintainer The ball is in the maintainers’ court label Oct 3, 2026
@tombeckenham
tombeckenham self-requested a review October 3, 2026 06:36
@github-actions github-actions Bot added merge-conflicts Conflicts with the base branch — needs a rebase waiting-on: author Waiting for the author to respond or update and removed waiting-on: maintainer The ball is in the maintainers’ court labels Oct 3, 2026
…ropic-max-tokens-usage

# Conflicts:
#	testing/e2e/src/routeTree.gen.ts
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

merge-conflicts Conflicts with the base branch — needs a rebase waiting-on: author Waiting for the author to respond or update

Projects

None yet

Development

Successfully merging this pull request may close these issues.

@tanstack/ai-anthropic: a max_tokens stop emits RUN_ERROR with no token usage

3 participants