Skip to content

fix(usage): preserve DeepSeek prompt cache hits - #5817

Closed
ramapitecusment wants to merge 1 commit into
router-for-me:devfrom
ramapitecusment:fix/deepseek-cache-usage
Closed

ramapitecusment wants to merge 1 commit into
router-for-me:devfrom
ramapitecusment:fix/deepseek-cache-usage

Conversation

@ramapitecusment

Copy link
Copy Markdown
Contributor

DeepSeek reports cached prompt tokens in usage.prompt_cache_hit_tokens. The OpenAI-compatible usage parser currently ignores this field, so a response with 100 input tokens, 80 cache hits and 10 output tokens is recorded as 100 uncached input tokens and zero cache reads. This loses the cache discount information used by downstream cost accounting.

Read the DeepSeek field as a fallback after the existing Chat Completions and Responses cache fields. The existing token breakdown then records 20 uncached input + 80 cached input + 10 output = 110 total. Explicit zero in a standard cache field remains authoritative.

The production change is three lines in the shared parser used by normal responses and SSE. It uses existing usage fields and requires no SDK, queue, journal or desktop changes. This covers documented responses with input/output counters; handling incomplete cache-only payloads remains outside this patch.

Validation:

  • TDD: the new tests failed on unmodified dev for mixed/full cache in both response and streaming modes, then passed with the fallback.
  • 16 regression subcases cover mixed/full/no cache, absent cache, both standard aliases, explicit-zero precedence, canonical bucket totals, and preservation through content/final-usage/[DONE] stream chunks.
  • go test ./internal/runtime/executor/helps ./sdk/cliproxy/usage -count=1 passed.
  • go build -o test-output ./cmd/server passed (Go 1.26.5, macOS arm64).
  • The wider executor suite encounters TestOpenAICompatExecutorToolResultContentByInputModalities; the same failure reproduces on clean upstream bb20fa2d, without this patch.
  • Independent review found no actionable issues in this small diff.

References: DeepSeek usage schema and cache semantics.

Extracted from #5765 so this correction can be reviewed independently of persistent delivery and broader accounting changes. Related to the earlier #3802 proposal; this patch only maps the documented counter into the existing schema and does not add new cache-hit/cache-miss SDK fields.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 14, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-14T05:38:28.131289Z 910da5d PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@sususu98

Copy link
Copy Markdown
Collaborator

Thanks for the PR!

Just a quick note from live verification: the official DeepSeek API (api.deepseek.com) currently already returns usage.prompt_tokens_details.cached_tokens alongside prompt_cache_hit_tokens (with identical values) in both streaming and non-streaming responses, matching their documented schema.

That said, adding prompt_cache_hit_tokens as a fallback is still very valuable for third-party gateways, self-hosted deployments (vLLM/SGLang), and custom proxies (such as reported in #3358) where the nested prompt_tokens_details object is omitted.

The fallback precedence and explicit-zero semantics look clean and well-tested. LGTM!

@sususu98

Copy link
Copy Markdown
Collaborator

Closing as not planned.

After further architectural consideration, we have decided to keep the generic OpenAI-compatible usage parser strictly aligned with the canonical OpenAI specification (prompt_tokens_details.cached_tokens). As verified, official DeepSeek (api.deepseek.com) and modern open-source engines (e.g. llama.cpp, vLLM) already emit standard prompt_tokens_details.cached_tokens natively.

Introducing vendor-specific non-standard fields (prompt_cache_hit_tokens) into the shared parser creates dialect bloat and technical debt. Non-compliant intermediate proxies or gateways should align with the standard specification instead.

Thank you again for the effort and the well-crafted test suite!

@sususu98 sususu98 closed this Sep 14, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants