Skip to content

Fix OpenRouter reasoning usage reporting in TanStack AI - #1563

Open
DrewHoo wants to merge 2 commits into
TanStack:mainfrom
DrewHoo:task-OZpW3nW1r-fix-openrouter-reasoning
Open

DrewHoo wants to merge 2 commits into
TanStack:mainfrom
DrewHoo:task-OZpW3nW1r-fix-openrouter-reasoning

Conversation

@DrewHoo

@DrewHoo DrewHoo commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

OpenRouter drops a reported zero reasoning-token count and loses received usage when structuredOutputStream() fails. Preserve zero and attach token usage and cost to the existing terminal RUN_ERROR event. Successful calls still report usage once on RUN_FINISHED.

🎯 Changes

Use a nullish presence check for reasoning tokens. Normalize received structured-stream usage before parsing and retain it for parse, truncation, empty-response, and SDK errors. Requests, schemas, structured results, and error codes stay unchanged.

This lets consumers account for failed requests without a provider-specific custom usage event or a downstream dependency patch. Includes regression tests, public chat() middleware E2E coverage, documentation, and a patch changeset. Two shipped @tanstack/ai skills now clarify that failed-call usage reaches onChunk, while onUsage receives RUN_FINISHED usage. The changeset covers the adapter fix and shipped core documentation.

✅ Checklist

  • I have followed the steps in the Contributing guide.
  • I have tested code changes locally with pnpm run test:pr, or these tests do not apply to this pull request.
  • I fully understand the code in this pull request, including any code generated with AI assistance. (Prepared and verified with Codex; human contributor review is pending.)
  • Docs: I updated docs/ for this change, or this change is not user-facing.
  • Changeset: I added a changeset (pnpm changeset), or this PR does not change a published package.

🚀 Release Impact

  • This change affects published code, and I have generated a changeset.
  • This change is docs/CI/dev-only (no release).

Root cause

Issue. Consumers cannot distinguish zero reasoning tokens from an unknown count, or account for received usage after structured-output failure.

Cause. buildOpenRouterUsage() checks truthiness. structuredOutputStream() emits its retained usage only after content checks and JSON parsing succeed.

Fix. Check presence and retain normalized usage for either terminal outcome. No extra usage event is emitted.

Possible alternatives

  • A custom openrouter.usage event introduces another accounting source. The existing RUN_ERROR.usage contract supports the required data.
  • Emitting RUN_FINISHED before parsing would label a failed structured response as successful.

Testing

Verified on macOS arm64, Node 26.0.0, pnpm 11.9.0. Base: 62bec34bb78a2f2d0d283c8ea2e9dc39fbd12d2c; implementation: b56990643e088394bf868872e8aa5fb526e9283e; final reviewed head: d2464b56e0275f456479b02d6c498d8f7e75741a. The cleanup commit changes only shipped skill prose and release metadata; runtime reproduction and E2E results remain applicable.

Commands passed:

  • NX_DAEMON=false NX_BASE=62bec34bb78a2f2d0d283c8ea2e9dc39fbd12d2c pnpm test:pr: 398 Nx tasks after cleanup, including affected builds, tests, types, lint, docs, and snippet checks (394 cached); React Native smoke and declaration scan passed.
  • pnpm --filter @tanstack/ai-openrouter test:lib: 255 tests passed. Adapter typecheck and lint passed; lint retains existing warnings.
  • E2E_PROVIDERS=openrouter pnpm test:e2e: all spec files ran, with 357 passed and 3 skipped. This includes all six new usage cases. This is the OpenRouter-filtered suite, not the full provider matrix.
  • E2E typecheck, pnpm build:all, pnpm test:docs, and git diff --check passed.

The new tests cover positive, zero, null/missing reasoning counts; malformed, truncated and empty output; success accounting; SDK failures and aborts before/after usage; and provider-iterator cleanup when the consumer stops. E2E exercises the real SDK decoder, public promise API, and middleware using deterministic SSE responses, without provider credentials. It uses a model routed through structuredOutputStream, not the separate combined tools/schema path.

Six independent cleanup audits retained the implementation and tests. The skill audit found the two passages corrected in the cleanup commit. Final verification run.

GitHub workflows require maintainer approval (action_required). Socket checks passed. CodeRabbit is pending with no substantive findings at the last read.

Reproduction transcripts

The same agent-written regression file ran against detached clean main and the committed fix. Command from each checkout's packages/ai-openrouter directory:

node ../../node_modules/vitest/vitest.mjs run tests/structured-usage.test.ts

Clean main:

preserves reasoning count 0
AssertionError: expected {} to deeply equal { reasoningTokens: +0 }
reports received usage once on malformed / truncated / empty
AssertionError: expected [] to have a length of 1 but got +0
Test Files  1 failed (1)
Tests       4 failed | 4 passed (8)

Committed fix:

Test Files  1 passed (1)
Tests       8 passed (8)

Reviewer test path

  1. Copy the new regression file to a checkout of the base revision and run the command above. Expect the four failures shown.
  2. Run the same command on this branch. Expect eight passes.
  3. Run the E2E command above to inspect public chat() success and failure accounting.

The regression file, E2E route and E2E assertions are included on the branch. No UI behavior changes; event assertions demonstrate this adapter contract.

Risk / rollback

The error event now contains usage that the provider already supplied. Consumers can read it through middleware onChunk. The core onUsage hook still runs for RUN_FINISHED only; this PR does not change that hook or the combined tools/schema path. A consumer that stops iterating cannot receive later events. Revert this PR to restore the prior behavior.

No released version contains this commit yet. The changeset requests an adapter patch release; the exact version remains maintainer-controlled. DrewHoo owns contribution follow-up and downstream release tracking.

Public API change

No caller signature changes. Middleware can now account for received usage on failed structured calls through the existing RUN_ERROR.usage field.

Before — onUsage accounts for successful calls. Failed structured calls omit received usage.

const usageObserver = {
  onUsage: (_ctx, usage) => recordUsage(usage),
} satisfies ChatMiddleware

After — keep successful-call accounting and add failed-call accounting through onChunk.

const usageObserver = {
  onUsage: (_ctx, usage) => recordUsage(usage),
  onChunk: (_ctx, chunk) => {
    if (
      chunk.type === 'RUN_ERROR' &&
      chunk.usage &&
      !Array.isArray(chunk.usage)
    ) {
      recordUsage(chunk.usage)
    }
  },
} satisfies ChatMiddleware

Both examples use ChatMiddleware from @tanstack/ai and the caller's recordUsage function. Pass usageObserver in chat({ middleware: [usageObserver], ... }). The two hooks handle separate terminal outcomes, so successful calls are counted once.

Summary by CodeRabbit

  • Bug Fixes

    • Preserved reported reasoning-token counts of zero.
    • Structured-output errors now include token usage and cost when those details were received before the error.
  • Documentation

    • Clarified how to access usage for failed and successful calls, when usage may be unavailable, and how stopping stream consumption affects later events.

@coderabbitai

coderabbitai Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: TanStack/ai/.coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 2c0a5c61-b643-4b0a-99c9-167fddd9c543

📥 Commits

Reviewing files that changed from the base of the PR and between 62bec34 and d2464b5.

📒 Files selected for processing (12)
  • .changeset/openrouter-structured-usage.md
  • docs/adapters/openrouter.md
  • docs/config.json
  • packages/ai-openrouter/src/adapters/text.ts
  • packages/ai-openrouter/src/usage.ts
  • packages/ai-openrouter/tests/openrouter-adapter.test.ts
  • packages/ai-openrouter/tests/structured-usage.test.ts
  • packages/ai/skills/ai-core/middleware/SKILL.md
  • packages/ai/skills/ai-core/structured-outputs/SKILL.md
  • testing/e2e/src/routeTree.gen.ts
  • testing/e2e/src/routes/api.openrouter-structured-usage.ts
  • testing/e2e/tests/structured-output-stream.spec.ts

Included review availability: This review used your included allowance. Your plan provides up to 8 included reviews per hour; 7 remain after this review.


📝 Walkthrough

Walkthrough

The OpenRouter adapter now retains normalized, cost-enriched usage and includes it on terminal error events when available. It preserves zero reasoning-token counts. New unit and E2E tests cover stream outcomes, middleware observations, and reasoning-token values.

Changes

OpenRouter structured usage reporting

Layer / File(s) Summary
Normalize usage and attach it to terminal events
packages/ai-openrouter/src/*, packages/ai-openrouter/tests/*, docs/adapters/openrouter.md, docs/config.json, packages/ai/skills/ai-core/*, .changeset/openrouter-structured-usage.md
The adapter normalizes usage and adds cost as chunks arrive, then includes retained usage on terminal success and error events. Reasoning-token value 0 is preserved. Unit tests cover error and interruption cases; documentation and release notes describe usage reporting and middleware hooks.
Exercise usage reporting through the E2E route
testing/e2e/src/routes/api.openrouter-structured-usage.ts, testing/e2e/src/routeTree.gen.ts, testing/e2e/tests/structured-output-stream.spec.ts
A new test route mocks structured responses and usage. Generated route registration and E2E cases cover valid and malformed output with positive, zero, or missing reasoning-token counts.

Priority: ➖ Normal

Estimated code review effort: 3 (Moderate) | ~25 minutes

Change: Bug fix

Suggested reviewers: alemtuzlak, tombeckenham

Merge Risk: ⚪ Minimal · up to d2464

No actionable issue remains in the supplied evidence; the PR is mergeable after normal checks.

Security Architecture Review

Security architecture risk: 🔵 Low · up to d2464

The change affects how callers account for failed requests, but no new production credential, provider request, or privilege exposure was demonstrated. The new test endpoint uses synthetic responses; whether the test app is externally reachable remains unverified.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • inferred — The new failure-event usage is visible to consumers of OpenRouter structured-output streams. The inspected route exposes only synthetic usage and results, not live provider credentials or charges; its runtime reachability is unknown.

Trust Boundaries and Controls

  • observed — Provider usage crosses into terminal events through token mapping and finite-number cost extraction, rather than by spreading the entire raw usage object. Existing prompt-token details can be passed through within the usage contract.

Resilience and Maintainability Implications

  • observed — Run identity and retained usage are invocation-local, and the inspected terminal branches return or end after emission. No cross-request usage association or duplicate terminal usage emission is apparent in this path.

Hardening Proposals

  • proposed — If the E2E app is served outside controlled test infrastructure, restrict access to its routes at the deployment boundary. The available source does not establish its ingress policy.
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 1 functions across 7 files. (5 skipped: 5 …
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly summarizes the primary change: fixing OpenRouter reasoning-token and usage reporting in TanStack AI.
Description check ✅ Passed The description follows the required template and provides clear change details, root cause, testing evidence, release impact, risk, rollback, and public API information. One checklist item remains un…
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the PR, @DrewHoo! 🙌 @jherr will take a look.

Automated pre-review checks

  • ✅ CI passing
  • ✅ No merge conflicts
  • ✅ Changeset present
  • ✅ E2E test changes included

Automated triage — a human review follows.

@github-actions github-actions Bot added waiting-on: maintainer The ball is in the maintainers’ court merge-conflicts Conflicts with the base branch — needs a rebase waiting-on: author Waiting for the author to respond or update and removed waiting-on: maintainer The ball is in the maintainers’ court labels Sep 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

merge-conflicts Conflicts with the base branch — needs a rebase waiting-on: author Waiting for the author to respond or update

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants