Skip to content

feat(pricing): DeepSeek V4.1-Flash + time-of-day peak/off-peak pricing - #3

Merged
jsaigou merged 1 commit into
mainfrom
deepseek-peak-pricing-and-compressor-fixes
Sep 12, 2026
Merged

jsaigou merged 1 commit into
mainfrom
deepseek-peak-pricing-and-compressor-fixes

Conversation

@jsaigou

@jsaigou jsaigou commented Sep 12, 2026

Copy link
Copy Markdown
Owner

Summary

Two independent threads, both live-verified before opening this PR:

  • DeepSeek V4.1-Flash + time-of-day peak/off-peak pricing — a new internal/pricing
    package (with proper IANA timezone/DST support), additive schema for peak price tiers,
    and the catalog/router/frontend wiring to price a request at the tier in force when it
    completes rather than a single flat rate. Full design rationale in docs/adr/0013.
  • Compressor investigation — the remote compression arm was found to be a net loss in
    production (large latency cost for a small dollar saving) and is now bypassed; real ONNX
    batching was implemented and verified correct; a dead local savings metric was fixed.

See the commit message for full detail on both.

Test plan

  • go build/go vet/go test clean across every touched package (re-verified against
    this exported tree, including a fresh npm ci + tsc -b for the frontend)
  • -race clean for the new internal/pricing package
  • Migration dry-run against a real copy of the production DB before deploying
  • Deployed and live-verified in production: a real completion priced and recorded
    correctly against the real clock; an invalid timezone rejected before any write; a
    scratch schedule correctly computed live peak/off-peak state, then reverted
  • Browser-verified in both themes

🤖 Generated with Claude Code

https://claude.ai/code/session_01TwE7KuNQQ2oiyuMo8LeoFh

DeepSeek retired V4-Flash and V4-Flash-Vision-Exp in favor of a single
natively-multimodal V4.1-Flash (wire model deepseek-flash), and switched to
peak/off-peak UTC pricing (peak = 2x, weekdays 01:00-04:00 + 06:00-10:00).
The catalog only ever stored one flat price triple per offering, so roughly
half of all DeepSeek spend would have been mispriced regardless of which
number was entered — built peak/off-peak support first rather than pick a
permanently-wrong tier.

- New internal/pricing package: time-window evaluation (half-open
  [start,end), midnight-wrap support), shared by the router and the
  historical savings estimators so a request is billed and a past event is
  re-priced by the exact same rule.
- Optional per-schedule IANA timezone (e.g. "America/Los_Angeles", "" =
  UTC): most providers quote peak hours in one local zone, and a hand UTC
  conversion goes silently wrong by an hour across every DST transition.
  Evaluated via time.Time.In against the real tzdata (memoized — this runs
  on the remote-request hot path), not a stored offset, so the same
  schedule stays correct across DST with no re-entry needed, and correctly
  shifts which calendar day/hour applies in zones far from UTC. Proven with
  tests, not just parsed: the same window is active at 16:30 UTC in July
  (PDT) but not at the same UTC clock time in January (PST) for a fixed
  09:00-17:00 Pacific schedule; a "Monday 00:00-04:00 Asia/Tokyo" window is
  correctly active during Sunday afternoon UTC.
- New tables/columns (all additive, nullable, no backfill): peak price
  triple on offerings (per-field nil falls back to the base rate),
  peak_windows on providers (provider-level — a peak schedule applies to
  every model that provider serves), price_tier recorded on usage events.
- Tier is resolved once, at usage-record time, not route-selection time:
  nothing upstream uses price for routing, and resolving early would let a
  boundary-crossing request's stored cost disagree with the tier its own
  stored timestamp implies. Historical savings estimators weight each past
  event at its own recorded (or re-derived) tier instead of blending at
  today's base rate. Full reasoning in docs/adr/0013.
- API: peak price fields validated for sign only, never against the base
  rate (that's the provider's own policy); peak_windows + a
  server-computed peak_active_now on providers, so the frontend never
  re-derives the window math itself.
- FE: an offering's peak-pricing section with a "peak = 2x off-peak" fill
  helper; a provider's schedule editor gets a labeled Timezone field (a
  curated common-zones dropdown + free-text IANA escape hatch) plus a
  validated JSON windows box with a human-readable summary; an offering
  enable/disable toggle switched from a hand-enumerated PUT body to a
  spread of the existing record (was a latent hazard — would have silently
  dropped any field added after it was written).

Verified: go build/vet/test (incl. -race on the new package) and
tsc -b/vite build/oxlint all clean. Live-verified against a real deployment
before this PR: one real completion through the router to deepseek-flash
confirmed cost and price tier computed and stored correctly against the
real clock; an invalid timezone is rejected before any write; a scratch
schedule on a second provider correctly computed peak_active_now against
the real time, then was reverted; checked in the browser in both themes.

fix(compressor): bypass a net-negative remote arm, real ONNX batching, fix
a dead savings metric

Investigation found the compressor's per-chunk ONNX scoring was fully
sequential batch-of-1 with no economy of scale: against a real
high-volume remote provider's traffic (tens of thousands of tokens per
request on average), that meant hundreds of sequential native calls per
request — tens of seconds of mean overhead against a low-single-digit-
second real upstream time-to-first-byte, saving a small amount of API
cost per year at the expense of hundreds of hours per year of added
latency. Local traffic's real payoff (avoided GPU prefill time, not a
dollar figure) had been invisible the whole time: the existing local
time-saved estimator was gated on a counter the compressor binary never
actually incremented.

- Bypass the net-negative remote compressor arm (a live config flip).
- GOMEMLIMIT set under the existing cgroup memory ceiling — Go's GC had no
  awareness of it, which was the real root cause of an observed OOM-kill.
- New local time-saved estimate keyed off the real tokens-saved counter
  instead of the dead request-cached path.
- New per-outcome/size-tier message counters and overhead percentile
  metrics wired end-to-end so compression's real value (and where it
  doesn't pay off) is visible from production traffic instead of a
  misleading mean.
- Real ONNX batching: the scorer now tokenizes up front and scores in
  configurable batches instead of one sequential call per chunk.
  Correctness verified against the unbatched path on real (including CJK)
  samples with zero decision mismatches; real measured speedup is modest
  (~1.2x) since this workload turned out to be compute-bound rather than
  per-call-overhead-bound — corrected an earlier, unverified assumption
  that batching would be much faster.
- Fixed a max-inflight setting that was never actually applied, silently
  stuck at a low hardcoded default.
- Promoted the dashboard's percentile/median helpers to a shared
  internal/statutil package rather than duplicating them.
- README corrected: the compression feature bullet had overstated
  remote-arm savings as measured/real when the remote arm was actually a
  net loss.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TwE7KuNQQ2oiyuMo8LeoFh
@jsaigou
jsaigou merged commit 1cec72a into main Sep 12, 2026
2 checks passed
@jsaigou
jsaigou deleted the deepseek-peak-pricing-and-compressor-fixes branch September 12, 2026 09:36
jsaigou added a commit that referenced this pull request Sep 16, 2026
#3)

DeepSeek retired V4-Flash and V4-Flash-Vision-Exp in favor of a single
natively-multimodal V4.1-Flash (wire model deepseek-flash), and switched to
peak/off-peak UTC pricing (peak = 2x, weekdays 01:00-04:00 + 06:00-10:00).
The catalog only ever stored one flat price triple per offering, so roughly
half of all DeepSeek spend would have been mispriced regardless of which
number was entered — built peak/off-peak support first rather than pick a
permanently-wrong tier.

- New internal/pricing package: time-window evaluation (half-open
  [start,end), midnight-wrap support), shared by the router and the
  historical savings estimators so a request is billed and a past event is
  re-priced by the exact same rule.
- Optional per-schedule IANA timezone (e.g. "America/Los_Angeles", "" =
  UTC): most providers quote peak hours in one local zone, and a hand UTC
  conversion goes silently wrong by an hour across every DST transition.
  Evaluated via time.Time.In against the real tzdata (memoized — this runs
  on the remote-request hot path), not a stored offset, so the same
  schedule stays correct across DST with no re-entry needed, and correctly
  shifts which calendar day/hour applies in zones far from UTC. Proven with
  tests, not just parsed: the same window is active at 16:30 UTC in July
  (PDT) but not at the same UTC clock time in January (PST) for a fixed
  09:00-17:00 Pacific schedule; a "Monday 00:00-04:00 Asia/Tokyo" window is
  correctly active during Sunday afternoon UTC.
- New tables/columns (all additive, nullable, no backfill): peak price
  triple on offerings (per-field nil falls back to the base rate),
  peak_windows on providers (provider-level — a peak schedule applies to
  every model that provider serves), price_tier recorded on usage events.
- Tier is resolved once, at usage-record time, not route-selection time:
  nothing upstream uses price for routing, and resolving early would let a
  boundary-crossing request's stored cost disagree with the tier its own
  stored timestamp implies. Historical savings estimators weight each past
  event at its own recorded (or re-derived) tier instead of blending at
  today's base rate. Full reasoning in docs/adr/0013.
- API: peak price fields validated for sign only, never against the base
  rate (that's the provider's own policy); peak_windows + a
  server-computed peak_active_now on providers, so the frontend never
  re-derives the window math itself.
- FE: an offering's peak-pricing section with a "peak = 2x off-peak" fill
  helper; a provider's schedule editor gets a labeled Timezone field (a
  curated common-zones dropdown + free-text IANA escape hatch) plus a
  validated JSON windows box with a human-readable summary; an offering
  enable/disable toggle switched from a hand-enumerated PUT body to a
  spread of the existing record (was a latent hazard — would have silently
  dropped any field added after it was written).

Verified: go build/vet/test (incl. -race on the new package) and
tsc -b/vite build/oxlint all clean. Live-verified against a real deployment
before this PR: one real completion through the router to deepseek-flash
confirmed cost and price tier computed and stored correctly against the
real clock; an invalid timezone is rejected before any write; a scratch
schedule on a second provider correctly computed peak_active_now against
the real time, then was reverted; checked in the browser in both themes.

fix(compressor): bypass a net-negative remote arm, real ONNX batching, fix
a dead savings metric

Investigation found the compressor's per-chunk ONNX scoring was fully
sequential batch-of-1 with no economy of scale: against a real
high-volume remote provider's traffic (tens of thousands of tokens per
request on average), that meant hundreds of sequential native calls per
request — tens of seconds of mean overhead against a low-single-digit-
second real upstream time-to-first-byte, saving a small amount of API
cost per year at the expense of hundreds of hours per year of added
latency. Local traffic's real payoff (avoided GPU prefill time, not a
dollar figure) had been invisible the whole time: the existing local
time-saved estimator was gated on a counter the compressor binary never
actually incremented.

- Bypass the net-negative remote compressor arm (a live config flip).
- GOMEMLIMIT set under the existing cgroup memory ceiling — Go's GC had no
  awareness of it, which was the real root cause of an observed OOM-kill.
- New local time-saved estimate keyed off the real tokens-saved counter
  instead of the dead request-cached path.
- New per-outcome/size-tier message counters and overhead percentile
  metrics wired end-to-end so compression's real value (and where it
  doesn't pay off) is visible from production traffic instead of a
  misleading mean.
- Real ONNX batching: the scorer now tokenizes up front and scores in
  configurable batches instead of one sequential call per chunk.
  Correctness verified against the unbatched path on real (including CJK)
  samples with zero decision mismatches; real measured speedup is modest
  (~1.2x) since this workload turned out to be compute-bound rather than
  per-call-overhead-bound — corrected an earlier, unverified assumption
  that batching would be much faster.
- Fixed a max-inflight setting that was never actually applied, silently
  stuck at a low hardcoded default.
- Promoted the dashboard's percentile/median helpers to a shared
  internal/statutil package rather than duplicating them.
- README corrected: the compression feature bullet had overstated
  remote-arm savings as measured/real when the remote arm was actually a
  net loss.


Claude-Session: https://claude.ai/code/session_01TwE7KuNQQ2oiyuMo8LeoFh

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant