feat(pricing): DeepSeek V4.1-Flash + time-of-day peak/off-peak pricing - #3
Merged
Merged
Conversation
DeepSeek retired V4-Flash and V4-Flash-Vision-Exp in favor of a single natively-multimodal V4.1-Flash (wire model deepseek-flash), and switched to peak/off-peak UTC pricing (peak = 2x, weekdays 01:00-04:00 + 06:00-10:00). The catalog only ever stored one flat price triple per offering, so roughly half of all DeepSeek spend would have been mispriced regardless of which number was entered — built peak/off-peak support first rather than pick a permanently-wrong tier. - New internal/pricing package: time-window evaluation (half-open [start,end), midnight-wrap support), shared by the router and the historical savings estimators so a request is billed and a past event is re-priced by the exact same rule. - Optional per-schedule IANA timezone (e.g. "America/Los_Angeles", "" = UTC): most providers quote peak hours in one local zone, and a hand UTC conversion goes silently wrong by an hour across every DST transition. Evaluated via time.Time.In against the real tzdata (memoized — this runs on the remote-request hot path), not a stored offset, so the same schedule stays correct across DST with no re-entry needed, and correctly shifts which calendar day/hour applies in zones far from UTC. Proven with tests, not just parsed: the same window is active at 16:30 UTC in July (PDT) but not at the same UTC clock time in January (PST) for a fixed 09:00-17:00 Pacific schedule; a "Monday 00:00-04:00 Asia/Tokyo" window is correctly active during Sunday afternoon UTC. - New tables/columns (all additive, nullable, no backfill): peak price triple on offerings (per-field nil falls back to the base rate), peak_windows on providers (provider-level — a peak schedule applies to every model that provider serves), price_tier recorded on usage events. - Tier is resolved once, at usage-record time, not route-selection time: nothing upstream uses price for routing, and resolving early would let a boundary-crossing request's stored cost disagree with the tier its own stored timestamp implies. Historical savings estimators weight each past event at its own recorded (or re-derived) tier instead of blending at today's base rate. Full reasoning in docs/adr/0013. - API: peak price fields validated for sign only, never against the base rate (that's the provider's own policy); peak_windows + a server-computed peak_active_now on providers, so the frontend never re-derives the window math itself. - FE: an offering's peak-pricing section with a "peak = 2x off-peak" fill helper; a provider's schedule editor gets a labeled Timezone field (a curated common-zones dropdown + free-text IANA escape hatch) plus a validated JSON windows box with a human-readable summary; an offering enable/disable toggle switched from a hand-enumerated PUT body to a spread of the existing record (was a latent hazard — would have silently dropped any field added after it was written). Verified: go build/vet/test (incl. -race on the new package) and tsc -b/vite build/oxlint all clean. Live-verified against a real deployment before this PR: one real completion through the router to deepseek-flash confirmed cost and price tier computed and stored correctly against the real clock; an invalid timezone is rejected before any write; a scratch schedule on a second provider correctly computed peak_active_now against the real time, then was reverted; checked in the browser in both themes. fix(compressor): bypass a net-negative remote arm, real ONNX batching, fix a dead savings metric Investigation found the compressor's per-chunk ONNX scoring was fully sequential batch-of-1 with no economy of scale: against a real high-volume remote provider's traffic (tens of thousands of tokens per request on average), that meant hundreds of sequential native calls per request — tens of seconds of mean overhead against a low-single-digit- second real upstream time-to-first-byte, saving a small amount of API cost per year at the expense of hundreds of hours per year of added latency. Local traffic's real payoff (avoided GPU prefill time, not a dollar figure) had been invisible the whole time: the existing local time-saved estimator was gated on a counter the compressor binary never actually incremented. - Bypass the net-negative remote compressor arm (a live config flip). - GOMEMLIMIT set under the existing cgroup memory ceiling — Go's GC had no awareness of it, which was the real root cause of an observed OOM-kill. - New local time-saved estimate keyed off the real tokens-saved counter instead of the dead request-cached path. - New per-outcome/size-tier message counters and overhead percentile metrics wired end-to-end so compression's real value (and where it doesn't pay off) is visible from production traffic instead of a misleading mean. - Real ONNX batching: the scorer now tokenizes up front and scores in configurable batches instead of one sequential call per chunk. Correctness verified against the unbatched path on real (including CJK) samples with zero decision mismatches; real measured speedup is modest (~1.2x) since this workload turned out to be compute-bound rather than per-call-overhead-bound — corrected an earlier, unverified assumption that batching would be much faster. - Fixed a max-inflight setting that was never actually applied, silently stuck at a low hardcoded default. - Promoted the dashboard's percentile/median helpers to a shared internal/statutil package rather than duplicating them. - README corrected: the compression feature bullet had overstated remote-arm savings as measured/real when the remote arm was actually a net loss. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TwE7KuNQQ2oiyuMo8LeoFh
jsaigou
added a commit
that referenced
this pull request
Sep 16, 2026
#3) DeepSeek retired V4-Flash and V4-Flash-Vision-Exp in favor of a single natively-multimodal V4.1-Flash (wire model deepseek-flash), and switched to peak/off-peak UTC pricing (peak = 2x, weekdays 01:00-04:00 + 06:00-10:00). The catalog only ever stored one flat price triple per offering, so roughly half of all DeepSeek spend would have been mispriced regardless of which number was entered — built peak/off-peak support first rather than pick a permanently-wrong tier. - New internal/pricing package: time-window evaluation (half-open [start,end), midnight-wrap support), shared by the router and the historical savings estimators so a request is billed and a past event is re-priced by the exact same rule. - Optional per-schedule IANA timezone (e.g. "America/Los_Angeles", "" = UTC): most providers quote peak hours in one local zone, and a hand UTC conversion goes silently wrong by an hour across every DST transition. Evaluated via time.Time.In against the real tzdata (memoized — this runs on the remote-request hot path), not a stored offset, so the same schedule stays correct across DST with no re-entry needed, and correctly shifts which calendar day/hour applies in zones far from UTC. Proven with tests, not just parsed: the same window is active at 16:30 UTC in July (PDT) but not at the same UTC clock time in January (PST) for a fixed 09:00-17:00 Pacific schedule; a "Monday 00:00-04:00 Asia/Tokyo" window is correctly active during Sunday afternoon UTC. - New tables/columns (all additive, nullable, no backfill): peak price triple on offerings (per-field nil falls back to the base rate), peak_windows on providers (provider-level — a peak schedule applies to every model that provider serves), price_tier recorded on usage events. - Tier is resolved once, at usage-record time, not route-selection time: nothing upstream uses price for routing, and resolving early would let a boundary-crossing request's stored cost disagree with the tier its own stored timestamp implies. Historical savings estimators weight each past event at its own recorded (or re-derived) tier instead of blending at today's base rate. Full reasoning in docs/adr/0013. - API: peak price fields validated for sign only, never against the base rate (that's the provider's own policy); peak_windows + a server-computed peak_active_now on providers, so the frontend never re-derives the window math itself. - FE: an offering's peak-pricing section with a "peak = 2x off-peak" fill helper; a provider's schedule editor gets a labeled Timezone field (a curated common-zones dropdown + free-text IANA escape hatch) plus a validated JSON windows box with a human-readable summary; an offering enable/disable toggle switched from a hand-enumerated PUT body to a spread of the existing record (was a latent hazard — would have silently dropped any field added after it was written). Verified: go build/vet/test (incl. -race on the new package) and tsc -b/vite build/oxlint all clean. Live-verified against a real deployment before this PR: one real completion through the router to deepseek-flash confirmed cost and price tier computed and stored correctly against the real clock; an invalid timezone is rejected before any write; a scratch schedule on a second provider correctly computed peak_active_now against the real time, then was reverted; checked in the browser in both themes. fix(compressor): bypass a net-negative remote arm, real ONNX batching, fix a dead savings metric Investigation found the compressor's per-chunk ONNX scoring was fully sequential batch-of-1 with no economy of scale: against a real high-volume remote provider's traffic (tens of thousands of tokens per request on average), that meant hundreds of sequential native calls per request — tens of seconds of mean overhead against a low-single-digit- second real upstream time-to-first-byte, saving a small amount of API cost per year at the expense of hundreds of hours per year of added latency. Local traffic's real payoff (avoided GPU prefill time, not a dollar figure) had been invisible the whole time: the existing local time-saved estimator was gated on a counter the compressor binary never actually incremented. - Bypass the net-negative remote compressor arm (a live config flip). - GOMEMLIMIT set under the existing cgroup memory ceiling — Go's GC had no awareness of it, which was the real root cause of an observed OOM-kill. - New local time-saved estimate keyed off the real tokens-saved counter instead of the dead request-cached path. - New per-outcome/size-tier message counters and overhead percentile metrics wired end-to-end so compression's real value (and where it doesn't pay off) is visible from production traffic instead of a misleading mean. - Real ONNX batching: the scorer now tokenizes up front and scores in configurable batches instead of one sequential call per chunk. Correctness verified against the unbatched path on real (including CJK) samples with zero decision mismatches; real measured speedup is modest (~1.2x) since this workload turned out to be compute-bound rather than per-call-overhead-bound — corrected an earlier, unverified assumption that batching would be much faster. - Fixed a max-inflight setting that was never actually applied, silently stuck at a low hardcoded default. - Promoted the dashboard's percentile/median helpers to a shared internal/statutil package rather than duplicating them. - README corrected: the compression feature bullet had overstated remote-arm savings as measured/real when the remote arm was actually a net loss. Claude-Session: https://claude.ai/code/session_01TwE7KuNQQ2oiyuMo8LeoFh Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Two independent threads, both live-verified before opening this PR:
internal/pricingpackage (with proper IANA timezone/DST support), additive schema for peak price tiers,
and the catalog/router/frontend wiring to price a request at the tier in force when it
completes rather than a single flat rate. Full design rationale in
docs/adr/0013.production (large latency cost for a small dollar saving) and is now bypassed; real ONNX
batching was implemented and verified correct; a dead local savings metric was fixed.
See the commit message for full detail on both.
Test plan
go build/go vet/go testclean across every touched package (re-verified againstthis exported tree, including a fresh
npm ci+tsc -bfor the frontend)-raceclean for the newinternal/pricingpackagecorrectly against the real clock; an invalid timezone rejected before any write; a
scratch schedule correctly computed live peak/off-peak state, then reverted
🤖 Generated with Claude Code
https://claude.ai/code/session_01TwE7KuNQQ2oiyuMo8LeoFh