diff --git a/docs/adr/0013-peak-pricing-tier-resolved-at-record-time.md b/docs/adr/0013-peak-pricing-tier-resolved-at-record-time.md new file mode 100644 index 0000000..db8dc87 --- /dev/null +++ b/docs/adr/0013-peak-pricing-tier-resolved-at-record-time.md @@ -0,0 +1,83 @@ +# Peak/off-peak pricing tier is resolved once, at usage-record time + +Status: accepted. + +DeepSeek switched to time-of-day (peak/off-peak) API pricing on 2026-09-10 — peak hours cost +2x off-peak, currently 01:00–04:00 and 06:00–10:00 UTC, Monday–Friday. `store.Offering` held +exactly one flat price triple, so before this change roughly half of all DeepSeek spend was +mispriced no matter what number was entered. + +## Decision + +Each offering carries an optional peak-tier price triple (`price_in_per_1m_peak`/ +`price_out_per_1m_peak`/`price_cached_in_per_1m_peak`, per-field nil = falls back to the base +rate — the same convention `PriceCachedInPer1M` already used); the schedule itself +(`internal/pricing.Windows`) lives on the **provider**, not the offering, since a peak window is +a fact about the provider's billing policy shared by every model it serves, not a per-model +property. `router/routing.go`'s `offeringChain` copies both triples plus the parsed schedule +onto `Backend`/`ResolvedBackend` at its single existing copy point. + +**The tier itself is picked exactly once, in `computeCostNative` (`router/usage.go`), at the +same instant `recordExternalUsage` stamps the event's timestamp** — not earlier, at +route-selection/request-start time. Concretely: `now := time.Now()` is captured once in +`recordExternalUsage` and used for both `ev.TS` and the tier lookup; the tier is recorded on the +event (`usage_events.price_tier`). + +## Why not resolve the tier at request start + +Nothing upstream of cost computation uses price for a routing decision — `select.go` sorts +candidate offerings by `priority` only, so resolving the tier early buys nothing operationally. +Resolving it early would actively create a correctness problem: a request whose response +completes after a tier boundary would then have a stored cost that *disagrees with the tier +derivable from the event's own stored timestamp* — and `compressor_summary_handlers.go`'s +historical savings estimators (`estimateRemoteCacheDiscountSaved`/ +`estimateRemoteCompressionSaved`) re-price past events from exactly that timestamp. Pinning the +tier at record time means the ledger is self-consistent by construction: one function +(`computeCostNative`), one time input, shared by both the live-billing path and the historical +estimator (which prefers the event's own recorded `price_tier`, falling back to evaluating the +schedule against `ev.TS` only for rows written before this migration). + +**Consequence, disclosed rather than hidden:** a request that starts in one tier and whose +response completes after a boundary is billed entirely at the tier in force at completion, not +a split or the starting tier. A 2x rate delta on a long streaming completion is exactly the +boundary-crossing case where this matters most — a reproducible, auditable rule (visible via the +recorded `price_tier`) beats guessing DeepSeek's own internal accounting for the same case. + +## Other decisions folded in here + +- **Half-open `[start, end)` window boundaries** — a request landing exactly on the window's end + time is off-peak, pinned by `internal/pricing`'s own tests. +- **No "has peak pricing" boolean.** A flag can disagree with the data it describes; the + per-field-nil convention is one source of truth. +- **Peak price validated for sign only, never `peak >= base`.** Whether a provider's peak tier + costs more or less than off-peak is the provider's own policy, not this app's to enforce. +- **`peak_active_now` is computed server-side only**, both on the repeatedly-polled + `GET /api/v1/providers` list (via `internal/providers.Service`'s existing injectable clock) and + on the one-off create/update echo (`httpapi.peakActiveNow`, plain `time.Now()`) — the window + math is never ported to TypeScript. + +## Correction, 2026-09-13: timezone support + +The original decision above said "UTC only, no timezone field — a timezone field the code +doesn't honor is worse than no field." That was right for DeepSeek's own published schedule (it +really is UTC), but wrong as a general policy: most providers publish peak hours in one local +zone ("9am–5pm Pacific"), and requiring the operator to hand-convert to UTC is exactly the trap +the original reasoning was trying to avoid — a fixed UTC offset entered once goes silently wrong +by an hour across the next DST transition, twice a year, with nothing to catch it. + +**Fixed properly, not worked around:** `Windows` gained an optional `TZ` field (IANA zone name, +`""` = UTC, backward-compatible with DeepSeek's already-stored no-`tz` schedule). `Active` +converts the evaluation instant into that zone via `time.Time.In` — Go's real IANA tzdata, not a +stored offset — before any day-of-week/time-of-day comparison, so the same schedule is correct +in January and July without ever being re-entered, and a schedule quoted in a zone far from UTC +correctly shifts which calendar day it evaluates against (e.g. "Monday 00:00–04:00 Asia/Tokyo" is +active during Sunday afternoon UTC, because it's already Monday in Tokyo). `time.LoadLocation` is +memoized (`internal/pricing`'s package-level `locCache`) since `Active` runs on the remote-request +hot path and the stdlib doesn't cache zoneinfo parsing itself. No new migration — `peak_windows` +was already a JSON `TEXT` column; `tz` is just a new key in that same blob. + +Frontend: `ProviderKeys`' `PeakWindowsEditor` gained a labeled Timezone field (a curated +common-zones dropdown + a "Custom…" free-text IANA name escape hatch) separate from the windows +JSON textarea, composing/splitting the two into the one stored JSON string — the timezone is a +single flat value worth a real input, unlike the windows list itself (still JSON; still no +bespoke hour-range picker, per the original reasoning, which stands). diff --git a/go/cmd/forge-compress/config.go b/go/cmd/forge-compress/config.go index 7c3c1e7..f415b30 100644 --- a/go/cmd/forge-compress/config.go +++ b/go/cmd/forge-compress/config.go @@ -113,6 +113,11 @@ func loadConfig() (config, error) { } else { c.Compress.ByteThreshold = v } + if v, err := intEnv("FORGE_COMPRESS_BATCH_SIZE", c.Compress.BatchSize); err != nil { + return config{}, err + } else { + c.Compress.BatchSize = v + } if v, err := intEnv("FORGE_COMPRESS_FAILOPEN_BUDGET_MS", c.FailOpenBudgetMS); err != nil { return config{}, err } else { diff --git a/go/cmd/forge-compress/messages.go b/go/cmd/forge-compress/messages.go index 850c10f..50f2fb6 100644 --- a/go/cmd/forge-compress/messages.go +++ b/go/cmd/forge-compress/messages.go @@ -18,12 +18,18 @@ import ( // succeed. Returns the real ModernBERT-tokenizer token counts summed // across every message actually run through the engine (0 for either if // nothing in the request was tokenizable/compressible), for the caller to -// fold into the tokens_saved metric. -func compressMessages(engine *compress.Engine, body map[string]any, budget time.Duration) (originalTokens, compressedTokens int64, failOpenTimeout, failOpenError int64) { +// fold into the tokens_saved metric, plus outcomeSize: a per-message count +// of "outcome:size_tier" composite labels (see messageOutcomeSize) for the +// caller to fold into messagesByOutcomeSize — added 2026-09-11 so the +// bimodal "compression barely matters on chatty messages, matters enormously +// on huge ones" shape (the operator's own early-testing observation) is +// directly visible in production telemetry instead of inferred from a mean. +func compressMessages(engine *compress.Engine, body map[string]any, budget time.Duration) (originalTokens, compressedTokens int64, failOpenTimeout, failOpenError int64, outcomeSize map[string]int64) { messagesRaw, ok := body["messages"].([]any) if !ok { - return 0, 0, 0, 0 + return 0, 0, 0, 0, nil } + outcomeSize = make(map[string]int64) for _, mRaw := range messagesRaw { msg, ok := mRaw.(map[string]any) if !ok { @@ -43,8 +49,52 @@ func compressMessages(engine *compress.Engine, body map[string]any, budget time. case "error": failOpenError++ } + outcomeSize[messageOutcomeSize(res, reason, len(content))]++ + } + return originalTokens, compressedTokens, failOpenTimeout, failOpenError, outcomeSize +} + +// messageOutcomeSize classifies one message into a composite +// "outcome:size_tier" label value. +// +// outcome distinguishes gated_passthrough (below Engine.Config's +// MinWords/ByteThreshold — the engine skips tokenization entirely, so +// compress.Result.OriginalTokens stays 0, per that field's own doc comment) +// from alldrop_passthrough (tokenization and scoring both ran, every word +// scored below threshold — OriginalTokens is real/nonzero) using that +// existing documented invariant, rather than adding a new field to +// compress.Result just for this. +func messageOutcomeSize(res compress.Result, reason string, contentBytes int) string { + var outcome string + switch { + case reason == "timeout": + outcome = "failopen_timeout" + case reason == "error": + outcome = "failopen_error" + case res.Passthrough && res.OriginalTokens == 0: + outcome = "gated_passthrough" + case res.Passthrough: + outcome = "alldrop_passthrough" + default: + outcome = "compressed" + } + return outcome + ":" + sizeTier(contentBytes) +} + +// sizeTier buckets a message's raw content size — deliberately coarse (4 +// tiers) so the resulting label cardinality stays small regardless of +// traffic volume. +func sizeTier(bytes int) string { + switch { + case bytes < 2*1024: + return "small" + case bytes < 20*1024: + return "medium" + case bytes < 200*1024: + return "large" + default: + return "huge" } - return originalTokens, compressedTokens, failOpenTimeout, failOpenError } // compressOne runs engine.Compress with a wall-clock fail-open budget and diff --git a/go/cmd/forge-compress/metrics.go b/go/cmd/forge-compress/metrics.go index 91a7a81..bcecfa8 100644 --- a/go/cmd/forge-compress/metrics.go +++ b/go/cmd/forge-compress/metrics.go @@ -8,6 +8,8 @@ import ( "os" "sort" "sync" + + "github.com/jsaigou/the-forge/internal/statutil" ) // selfRSSBytes reads this process's resident set from /proc/self/statm @@ -56,18 +58,41 @@ type metrics struct { ttfb histogram latency histogram overhead histogram + // overheadRing backs compress_overhead_ms_p50/p90/p99 — see sampleRing's + // doc comment for why overhead specifically gets a percentile view and + // ttfb/latency don't (yet): overhead is the one number this session's + // investigation found was actively misleading as a mean (dominated by a + // small share of huge messages, hiding that most real traffic barely + // pays the tax). + overheadRing *sampleRing failOpenTimeout counter failOpenError counter requestsByProvider labelCounter requestsByModel labelCounter + // messagesByOutcomeSize is keyed by a composite "outcome:size_tier" + // label value (e.g. "compressed:huge") rather than two independent + // label dimensions — this repo's label-sample storage + // (internal/store's compressor_label_samples) is a flat + // (label_key, label_value, metric) shape with one dimension per row, so + // a composite value is how a second dimension rides along without a + // schema change. outcome ∈ {compressed, gated_passthrough, + // alldrop_passthrough, failopen_timeout, failopen_error}; size_tier ∈ + // {small, medium, large, huge} — see messageOutcomeSize in messages.go. + // Added 2026-09-11 to answer whether compression's real value is + // concentrated in a few huge messages (the operator's own early-testing + // finding) or spread evenly — something the prior mean-only metrics + // couldn't show. + messagesByOutcomeSize labelCounter } func newMetrics() *metrics { return &metrics{ - requestsByProvider: newLabelCounter(), - requestsByModel: newLabelCounter(), + overheadRing: newSampleRing(overheadRingCapacity), + requestsByProvider: newLabelCounter(), + requestsByModel: newLabelCounter(), + messagesByOutcomeSize: newLabelCounter(), } } @@ -106,9 +131,11 @@ func (m *metrics) WriteTo(w io.Writer) (int64, error) { writeHistogram(write, "compress_ttfb_ms", &m.ttfb) writeHistogram(write, "compress_latency_ms", &m.latency) writeHistogram(write, "compress_overhead_ms", &m.overhead) + writePercentiles(write, "compress_overhead_ms", m.overheadRing) writeLabelCounter(write, "compress_requests_by_provider", "provider", &m.requestsByProvider) writeLabelCounter(write, "compress_requests_by_model", "model", &m.requestsByModel) + writeLabelCounter(write, "compress_messages_total", "outcome_size", &m.messagesByOutcomeSize) return n, nil } @@ -179,6 +206,80 @@ func (h *histogram) snapshot() (count int64, sum, min, max float64) { return h.count, h.sum, h.min, h.max } +const ( + // overheadRingCapacity mirrors the collector's existing 120-sample + // sparkline-ring pattern (internal/collector/run.go's rings field) — + // this process has no other precedent for bounding an otherwise + // unbounded-lifetime sample set. + overheadRingCapacity = 256 + // percentileMinSamples is this repo's established floor for trusting a + // percentile computed from a raw sample set — see + // internal/httpapi/cost_handlers.go's activeSingleSlotWallW gate and + // compressor_summary_handlers.go's prefillObservedMinSamples, both + // named "10" for the same reason: a couple of noisy early observations + // shouldn't produce a misleadingly-precise-looking figure. + percentileMinSamples = 10 +) + +// sampleRing is a fixed-capacity, thread-safe ring buffer of recent +// float64 samples. This binary's histogram accumulators are lifetime-since- +// process-start (never reset — see histogram's doc comment), so a plain +// growing []float64 isn't safe for a long-running process; a bounded ring +// gives "percentile of recent traffic" instead, which is what actually +// answers "is this request typical" during an incident. +type sampleRing struct { + mu sync.Mutex + buf []float64 + next int + full bool +} + +func newSampleRing(capacity int) *sampleRing { + return &sampleRing{buf: make([]float64, capacity)} +} + +func (r *sampleRing) add(v float64) { + r.mu.Lock() + defer r.mu.Unlock() + r.buf[r.next] = v + r.next++ + if r.next == len(r.buf) { + r.next = 0 + r.full = true + } +} + +// snapshot returns a copy of the samples currently held. Order doesn't +// matter — statutil.Percentile sorts its own copy. +func (r *sampleRing) snapshot() []float64 { + r.mu.Lock() + defer r.mu.Unlock() + if r.full { + out := make([]float64, len(r.buf)) + copy(out, r.buf) + return out + } + out := make([]float64, r.next) + copy(out, r.buf[:r.next]) + return out +} + +// writePercentiles emits p50/p90/p99 for r under the "name_pNN" series +// names, below percentileMinSamples samples emits nothing at all — a +// missing series is the honest signal, never a percentile computed from too +// few points to mean anything. Stored/read as a latest-snapshot gauge (like +// histogram's own min/max), not summed or averaged across a window — same +// invariant documented at internal/store/store.go's CompressorSavingsSampleRow. +func writePercentiles(write func(string, ...any), name string, r *sampleRing) { + vals := r.snapshot() + if len(vals) < percentileMinSamples { + return + } + write("%s_p50 %g\n", name, statutil.Percentile(vals, 50)) + write("%s_p90 %g\n", name, statutil.Percentile(vals, 90)) + write("%s_p99 %g\n", name, statutil.Percentile(vals, 99)) +} + // labelCounter is a set of independent counters keyed by one label value // (e.g. provider name, model name). type labelCounter struct { diff --git a/go/cmd/forge-compress/server.go b/go/cmd/forge-compress/server.go index e1d9204..4d8660b 100644 --- a/go/cmd/forge-compress/server.go +++ b/go/cmd/forge-compress/server.go @@ -168,8 +168,10 @@ func (s *server) handleChatCompletions(w http.ResponseWriter, r *http.Request) { } compressStart := time.Now() budget := time.Duration(s.cfg.FailOpenBudgetMS) * time.Millisecond - originalTokens, compressedTokens, foTimeout, foError := compressMessages(s.engine, body, budget) - s.metrics.overhead.observe(msSince(compressStart)) + originalTokens, compressedTokens, foTimeout, foError, outcomeSize := compressMessages(s.engine, body, budget) + overheadMs := msSince(compressStart) + s.metrics.overhead.observe(overheadMs) + s.metrics.overheadRing.add(overheadMs) s.metrics.tokensInput.add(originalTokens) if originalTokens > compressedTokens { s.metrics.tokensSaved.add(originalTokens - compressedTokens) @@ -180,6 +182,9 @@ func (s *server) handleChatCompletions(w http.ResponseWriter, r *http.Request) { if foError > 0 { s.metrics.failOpenError.add(foError) } + for label, count := range outcomeSize { + s.metrics.messagesByOutcomeSize.add(label, count) + } if reencoded, err := json.Marshal(body); err == nil { mutatedBody = reencoded } diff --git a/go/cmd/forge-compress/server_test.go b/go/cmd/forge-compress/server_test.go index 22c79e6..65d98a4 100644 --- a/go/cmd/forge-compress/server_test.go +++ b/go/cmd/forge-compress/server_test.go @@ -56,6 +56,18 @@ func (fakeScorer) Score(inputIDs, _ []int64) ([]float32, error) { return scores, nil } +func (f fakeScorer) ScoreBatch(inputIDs, attentionMask [][]int64) ([][]float32, error) { + out := make([][]float32, len(inputIDs)) + for i := range inputIDs { + s, err := f.Score(inputIDs[i], attentionMask[i]) + if err != nil { + return nil, err + } + out[i] = s + } + return out, nil +} + func testEngine() *compress.Engine { return &compress.Engine{ Tokenizer: fakeTokenizer{}, diff --git a/go/internal/collector/compressor.go b/go/internal/collector/compressor.go index 54c22f4..90d6fdb 100644 --- a/go/internal/collector/compressor.go +++ b/go/internal/collector/compressor.go @@ -61,12 +61,29 @@ type CompressorSample struct { OverheadSumMsDelta float64 OverheadMinMsSinceStart *float64 OverheadMaxMsSinceStart *float64 + // OverheadP50/P90/P99MsRecent are percentiles of the proxy's own recent + // (bounded ring, not lifetime) overhead samples — nil below the 10-sample + // floor (cmd/forge-compress/metrics.go's percentileMinSamples), same + // null-not-zero convention as the Min/Max gauges above. Unlike those, + // "recent" here does NOT mean "since process start" — it's the last + // ~256 requests, since the mean alone was found (2026-09-11) to hide a + // bimodal shape: most messages barely pay the compression tax, a few + // huge ones pay a lot. + OverheadP50MsRecent *float64 + OverheadP90MsRecent *float64 + OverheadP99MsRecent *float64 // RequestsByProviderDelta / RequestsByModelDelta are request COUNTS per // label value, not token counts — Compressor's compressor_requests_by_{ // provider,model} metrics carry no token dimension. RequestsByProviderDelta map[string]int64 RequestsByModelDelta map[string]int64 + // MessagesByOutcomeSizeDelta is per-MESSAGE (not per-request) counts + // keyed by a composite "outcome:size_tier" label value (e.g. + // "compressed:huge") — see cmd/forge-compress/messages.go's + // messageOutcomeSize. Added 2026-09-11 to make compression's real + // value visible by content-size tier instead of only as a blended mean. + MessagesByOutcomeSizeDelta map[string]int64 // Provider cache metrics (scraped from compress_cache_read_tokens_total, // compress_uncached_input_tokens_total, etc. — labelled by provider). diff --git a/go/internal/collector/llama.go b/go/internal/collector/llama.go index 3044313..2ad6516 100644 --- a/go/internal/collector/llama.go +++ b/go/internal/collector/llama.go @@ -334,11 +334,20 @@ type compressorCounters struct { TTFBCount, TTFBSum, TTFBMin, TTFBMax float64 LatencyCount, LatencySum, LatencyMin, LatencyMax float64 OverheadCount, OverheadSum, OverheadMin, OverheadMax float64 + // OverheadP50/P90/P99 (forge-compress only, absent on legacy proxies — + // an honest 0/false via parsePromScalar's discarded ok below) are + // percentiles of a bounded recent-sample ring, not a lifetime gauge — + // see CompressorSample.OverheadP50MsRecent's doc comment. + OverheadP50, OverheadP90, OverheadP99 float64 // RequestsByProvider / RequestsByModel are request COUNTS keyed by // label value (compressor_requests_by_{provider,model}) — no token // dimension is exposed per label. RequestsByProvider map[string]float64 RequestsByModel map[string]float64 + // MessagesByOutcomeSize is per-MESSAGE counts keyed by a composite + // "outcome:size_tier" label (compress_messages_total{outcome_size}) — + // forge-compress only, absent on legacy proxies. + MessagesByOutcomeSize map[string]float64 // Provider cache metrics — compress_cache_read_tokens_total{provider}, // compress_uncached_input_tokens_total{provider}, etc. Available since // at least 0.30.0 but lazily registered (only appear after first @@ -393,8 +402,12 @@ func (l *LlamaClient) scrapeCompressorCounters(ctx context.Context, port int) (* c.OverheadSum, _ = parsePromScalar(text, "compress_overhead_ms_sum") c.OverheadMin, _ = parsePromScalar(text, "compress_overhead_ms_min") c.OverheadMax, _ = parsePromScalar(text, "compress_overhead_ms_max") + c.OverheadP50, _ = parsePromScalar(text, "compress_overhead_ms_p50") + c.OverheadP90, _ = parsePromScalar(text, "compress_overhead_ms_p90") + c.OverheadP99, _ = parsePromScalar(text, "compress_overhead_ms_p99") c.RequestsByProvider = parsePromByLabel(text, "compress_requests_by_provider", "provider") c.RequestsByModel = parsePromByLabel(text, "compress_requests_by_model", "model") + c.MessagesByOutcomeSize = parsePromByLabel(text, "compress_messages_total", "outcome_size") c.CacheReadTokens = parsePromByLabel(text, "compress_cache_read_tokens_total", "provider") c.CacheWriteTokens = parsePromByLabel(text, "compress_cache_write_tokens_total", "provider") c.UncachedTokens = parsePromByLabel(text, "compress_uncached_input_tokens_total", "provider") diff --git a/go/internal/collector/run.go b/go/internal/collector/run.go index d04b670..cd21ad8 100644 --- a/go/internal/collector/run.go +++ b/go/internal/collector/run.go @@ -1122,6 +1122,7 @@ func (c *Collector) recordCompressorSavings(ctx context.Context) { OverheadSumMsDelta: deltaF(cur.OverheadSum, prev.OverheadSum), RequestsByProviderDelta: diffLabelMap(cur.RequestsByProvider, prev.RequestsByProvider), RequestsByModelDelta: diffLabelMap(cur.RequestsByModel, prev.RequestsByModel), + MessagesByOutcomeSizeDelta: diffLabelMap(cur.MessagesByOutcomeSize, prev.MessagesByOutcomeSize), CacheReadTokensDelta: diffLabelMap(cur.CacheReadTokens, prev.CacheReadTokens), CacheWriteTokensDelta: diffLabelMap(cur.CacheWriteTokens, prev.CacheWriteTokens), UncachedTokensDelta: diffLabelMap(cur.UncachedTokens, prev.UncachedTokens), @@ -1144,6 +1145,10 @@ func (c *Collector) recordCompressorSavings(ctx context.Context) { min, max := cur.OverheadMin, cur.OverheadMax sample.OverheadMinMsSinceStart, sample.OverheadMaxMsSinceStart = &min, &max } + if cur.OverheadP50 > 0 { + p50, p90, p99 := cur.OverheadP50, cur.OverheadP90, cur.OverheadP99 + sample.OverheadP50MsRecent, sample.OverheadP90MsRecent, sample.OverheadP99MsRecent = &p50, &p90, &p99 + } if len(cur.TransformTimingMax) > 0 { sample.TransformTimingMaxSinceStart = copyFloatMap(cur.TransformTimingMax) } diff --git a/go/internal/compress/chunk.go b/go/internal/compress/chunk.go index a5e25d4..20619b6 100644 --- a/go/internal/compress/chunk.go +++ b/go/internal/compress/chunk.go @@ -2,6 +2,8 @@ package compress +import "fmt" + // maxChunkTokens matches the ONNX model's declared max_length used at // inference (512 — kompress_compressor.py's tokenizer(..., max_length=512) // call). @@ -102,3 +104,69 @@ func scoreChunk(scorer Scorer, enc Encoding) ([]float32, error) { } return scores, nil } + +// scoreEncodings scores multiple chunk encodings, grouping those that +// individually fit within maxChunkTokens into real batched Scorer.ScoreBatch +// calls (up to batchSize encodings per call — a batchSize <= 1 degrades to +// one call per chunk, same shape as v1, still correct) instead of v1's +// sequential one-call-per-chunk loop. Any chunk whose own encoding already +// exceeds maxChunkTokens (the pathological single-huge-word case chunkWords' +// doc comment describes) is never batched — it falls through to +// scoreChunk's existing single-item sub-batching, unchanged, since that path +// already bounds the Scorer's per-call input size for exactly this case +// (the 2026-08-20 OOM incident's shape). +// +// Returns one []float32 per input encoding, in the same order — order is +// preserved regardless of how encodings were grouped into batches. +func scoreEncodings(scorer Scorer, encs []Encoding, batchSize int) ([][]float32, error) { + if batchSize < 1 { + batchSize = 1 + } + out := make([][]float32, len(encs)) + var pendingIdx []int + var pendingIDs, pendingMask [][]int64 + + flush := func() error { + if len(pendingIdx) == 0 { + return nil + } + scores, err := scorer.ScoreBatch(pendingIDs, pendingMask) + if err != nil { + return err + } + if len(scores) != len(pendingIdx) { + return fmt.Errorf("compress: ScoreBatch returned %d results for %d inputs", len(scores), len(pendingIdx)) + } + for j, idx := range pendingIdx { + out[idx] = scores[j] + } + pendingIdx, pendingIDs, pendingMask = nil, nil, nil + return nil + } + + for i, enc := range encs { + if len(enc.IDs) > maxChunkTokens { + if err := flush(); err != nil { + return nil, err + } + scores, err := scoreChunk(scorer, enc) + if err != nil { + return nil, err + } + out[i] = scores + continue + } + pendingIdx = append(pendingIdx, i) + pendingIDs = append(pendingIDs, enc.IDs) + pendingMask = append(pendingMask, enc.AttentionMask) + if len(pendingIdx) >= batchSize { + if err := flush(); err != nil { + return nil, err + } + } + } + if err := flush(); err != nil { + return nil, err + } + return out, nil +} diff --git a/go/internal/compress/chunk_test.go b/go/internal/compress/chunk_test.go index c8599dd..1873ead 100644 --- a/go/internal/compress/chunk_test.go +++ b/go/internal/compress/chunk_test.go @@ -177,6 +177,173 @@ func TestScoreChunk_UnderBudgetIsOneCall(t *testing.T) { } } +// recordingBatchScorer records every ScoreBatch call's batch size and +// returns one all-zero-score slice per input sequence, matching each +// sequence's own length. Score is never expected to be called by these +// tests (scoreEncodings only falls back to it for an oversized single +// encoding) — it panics if it is, so a wiring mistake fails loudly instead +// of silently passing. +type recordingBatchScorer struct { + batchSizes []int +} + +func (r *recordingBatchScorer) Score(inputIDs, _ []int64) ([]float32, error) { + panic("recordingBatchScorer.Score called unexpectedly — scoreEncodings should batch this input") +} + +func (r *recordingBatchScorer) ScoreBatch(inputIDs, attentionMask [][]int64) ([][]float32, error) { + if len(inputIDs) != len(attentionMask) { + panic("ids/mask batch length mismatch") + } + r.batchSizes = append(r.batchSizes, len(inputIDs)) + out := make([][]float32, len(inputIDs)) + for i, ids := range inputIDs { + // Score is the sequence's own index+1 (never 0), so tests can + // verify per-sequence results land back at the right position. + s := make([]float32, len(ids)) + for j := range s { + s[j] = float32(i + 1) + } + out[i] = s + } + return out, nil +} + +func TestScoreEncodings_GroupsUpToBatchSize(t *testing.T) { + // 10 small encodings, batchSize 4 -> batches of 4, 4, 2. + encs := make([]Encoding, 10) + for i := range encs { + encs[i] = Encoding{IDs: []int64{1, 2, 3}, AttentionMask: []int64{1, 1, 1}} + } + sc := &recordingBatchScorer{} + if _, err := scoreEncodings(sc, encs, 4); err != nil { + t.Fatal(err) + } + want := []int{4, 4, 2} + if len(sc.batchSizes) != len(want) { + t.Fatalf("ScoreBatch called %d times with sizes %v, want %d calls sized %v", len(sc.batchSizes), sc.batchSizes, len(want), want) + } + for i := range want { + if sc.batchSizes[i] != want[i] { + t.Errorf("batch %d size = %d, want %d (sizes: %v)", i, sc.batchSizes[i], want[i], sc.batchSizes) + } + } +} + +func TestScoreEncodings_PreservesOrderAcrossBatches(t *testing.T) { + // 5 encodings, batchSize 2 -> 3 batches. Each result must land back at + // its ORIGINAL index regardless of which batch it was scored in. + encs := make([]Encoding, 5) + for i := range encs { + encs[i] = Encoding{IDs: []int64{1}, AttentionMask: []int64{1}} + } + sc := &recordingBatchScorer{} + scores, err := scoreEncodings(sc, encs, 2) + if err != nil { + t.Fatal(err) + } + if len(scores) != 5 { + t.Fatalf("got %d results, want 5", len(scores)) + } + // Batches are [0,1], [2,3], [4] — within-batch sequence index+1 gives + // scores 1,2 / 1,2 / 1 respectively, NOT a monotonically increasing + // global sequence — this pins down that scoreEncodings doesn't + // accidentally assume batch-local index == global index. + want := []float32{1, 2, 1, 2, 1} + for i, w := range want { + if len(scores[i]) != 1 || scores[i][0] != w { + t.Errorf("scores[%d] = %v, want [%v]", i, scores[i], w) + } + } +} + +func TestScoreEncodings_OversizedEncodingBypassesBatchingWithoutDisruptingNeighbors(t *testing.T) { + // A middle encoding exceeds maxChunkTokens (the pathological + // single-huge-word case) — it must go through scoreChunk's existing + // single-item sub-batch path (via Score, not ScoreBatch), and must not + // get silently folded into a surrounding ScoreBatch call, nor prevent + // the normal encodings before/after it from still being batched + // together with each other. + normal := Encoding{IDs: []int64{1, 2}, AttentionMask: []int64{1, 1}} + oversized := Encoding{ + IDs: make([]int64, maxChunkTokens+10), + AttentionMask: make([]int64, maxChunkTokens+10), + } + for i := range oversized.IDs { + oversized.IDs[i] = int64(i) + oversized.AttentionMask[i] = 1 + } + encs := []Encoding{normal, normal, oversized, normal, normal} + + var scoreCalls, scoreBatchCalls int + var maxScoreBatchLen int + sc := fakeScorerFunc(func(inputIDs, attentionMask []int64) ([]float32, error) { + scoreCalls++ + if len(inputIDs) > maxChunkTokens { + t.Fatalf("Score received %d tokens, want <= %d", len(inputIDs), maxChunkTokens) + } + return make([]float32, len(inputIDs)), nil + }) + // Wrap to also count/inspect ScoreBatch calls, since fakeScorerFunc's + // own ScoreBatch just loops Score — swap in a small local wrapper that + // delegates but records. + wrapped := recordingWrapper{inner: sc, onBatch: func(n int) { + scoreBatchCalls++ + if n > maxScoreBatchLen { + maxScoreBatchLen = n + } + }} + + scores, err := scoreEncodings(wrapped, encs, 4) + if err != nil { + t.Fatal(err) + } + if len(scores) != len(encs) { + t.Fatalf("got %d results, want %d", len(scores), len(encs)) + } + if len(scores[2]) != len(oversized.IDs) { + t.Errorf("oversized encoding's score length = %d, want %d (must cover every token)", len(scores[2]), len(oversized.IDs)) + } + // The oversized encoding forces a flush before/after it, so the 4 + // normal encodings split into two ScoreBatch calls of 2 (indices 0,1 + // then 3,4) rather than one call of 4 spanning across it. + if scoreBatchCalls != 2 { + t.Errorf("ScoreBatch called %d times, want 2 (flushed around the oversized encoding)", scoreBatchCalls) + } + if maxScoreBatchLen != 2 { + t.Errorf("largest ScoreBatch call = %d items, want 2 (never spans the oversized encoding)", maxScoreBatchLen) + } + // scoreChunk splits the oversized encoding into ceil((maxChunkTokens+10)/maxChunkTokens) = 2 Score calls. + if scoreCalls != 2 { + t.Errorf("Score called %d times, want 2 (scoreChunk's sub-batching of the oversized encoding)", scoreCalls) + } +} + +// recordingWrapper adapts a fakeScorerFunc into a Scorer whose ScoreBatch +// calls onBatch(len) before delegating per-item to the wrapped Score func — +// lets a test both assert ScoreBatch call shape AND reuse fakeScorerFunc's +// existing Score-call assertions. +type recordingWrapper struct { + inner fakeScorerFunc + onBatch func(n int) +} + +func (r recordingWrapper) Score(inputIDs, attentionMask []int64) ([]float32, error) { + return r.inner(inputIDs, attentionMask) +} + +func (r recordingWrapper) ScoreBatch(inputIDs, attentionMask [][]int64) ([][]float32, error) { + r.onBatch(len(inputIDs)) + // Deliberately does NOT delegate to r.inner (Score) — that's reserved + // for scoreChunk's oversized-encoding sub-batching, so a test can count + // Score calls and ScoreBatch calls as two independent signals. + out := make([][]float32, len(inputIDs)) + for i, ids := range inputIDs { + out[i] = make([]float32, len(ids)) + } + return out, nil +} + func TestChunkWords_Empty(t *testing.T) { if got := chunkWords(nil, nil); got != nil { t.Errorf("chunkWords(nil, ...) = %v, want nil", got) diff --git a/go/internal/compress/compress.go b/go/internal/compress/compress.go index 32b4322..be4af05 100644 --- a/go/internal/compress/compress.go +++ b/go/internal/compress/compress.go @@ -57,15 +57,46 @@ type Config struct { // untouched rather than paying compression latency for a token saving // smaller than the latency cost. Sprint 2's measured net-positive // crossover recommended 2KB as the default (docs/v5-headroom-replacement.md - // Sprint 2 result, Probe A). + // Sprint 2 result, Probe A). Found stale 2026-09-11: that crossover was + // measured against a small fixture ladder, nowhere near real DeepSeek + // traffic's actual message sizes (avg ~31K tokens/request this session's + // investigation measured, code elsewhere cites up to ~262K) — the byte + // threshold alone doesn't gate the real cost driver, which is scoring + // hundreds of chunks with no economy of scale (see BatchSize). ByteThreshold int + // BatchSize bounds how many chunk encodings are sent to Scorer.ScoreBatch + // in one native call. Added 2026-09-11: v1 scored one chunk per call, + // sequentially. <= 1 behaves as one call per chunk (degrades to the old + // shape, still correct, just not exploiting batching). + // + // Measured live on ForgeHost's real model/hardware, same session (100 + // sequential ~510-token calls vs. one 100-item ScoreBatch call): this is + // a real but MODEST win (~1.2x), not a transformative one — the + // bottleneck here is CPU FLOPs for the matmul itself, not per-call + // dispatch overhead, so batching doesn't unlock new parallelism the way + // it would for a workload with real per-call fixed costs. The same + // sweep also disproved this comment's original speculation that a + // small ~510-token chunk wouldn't need 16 intra-op threads: it very + // much does (COMPRESS_ONNX_INTRA_THREADS=16 measured 5x faster per call + // than 1 thread, and was the ONLY thread count where batching helped at + // all — every lower thread count made both the sequential AND the + // batched path slower). Don't retune that value down without + // re-measuring. The bigger remaining lever for real messages is + // avoiding redundant recompute of unchanged earlier conversation turns + // (a bounded content-hash cache) — deliberately not built here, since + // it reopens a decision (docs/v5-headroom-replacement.md decision 1, + // "stateless, no cache") made specifically to kill a past OOM class; + // see the compressor investigation's plan for why that needs an + // explicit go-ahead rather than being bundled into this change. + BatchSize int } // DefaultConfig returns the values this initiative's own research settled // on: ScoreThreshold 0.5 (Kompress's own default), MinWords 10 (Kompress's -// own early-return), ByteThreshold 2048 (Sprint 2 Probe A). +// own early-return), ByteThreshold 2048 (Sprint 2 Probe A), BatchSize 16 +// (an untuned starting point — see BatchSize's own doc comment). func DefaultConfig() Config { - return Config{ScoreThreshold: 0.5, MinWords: 10, ByteThreshold: 2048} + return Config{ScoreThreshold: 0.5, MinWords: 10, ByteThreshold: 2048, BatchSize: 16} } // Encoding is one Tokenizer result: token ids plus, for every entry, the @@ -100,13 +131,25 @@ type Tokenizer interface { EncodeWords(words []string) (Encoding, error) } -// Scorer runs the Kompress ONNX model for one chunk (batch size 1 — this -// engine never batches; "one forward pass per request" is a deliberate v1 -// scope decision, docs/v5-headroom-replacement.md decision 1), returning -// one score per input position. The real implementation (onnxscorer) wraps -// onnxruntime_go via cgo; tests use a fake. +// Scorer runs the Kompress ONNX model. The real implementation (onnxscorer) +// wraps onnxruntime_go via cgo; tests use a fake. type Scorer interface { + // Score runs one chunk (batch size 1). Retained for the pathological + // single-huge-word case chunkWords' doc comment describes (chunk.go's + // scoreChunk) — the normal multi-chunk path uses ScoreBatch instead + // (2026-09-11: v1's "one forward pass per request" scope decision, + // docs/v5-headroom-replacement.md decision 1, was found this session to + // be the root cause of real production latency/fail-open problems on + // large real-world messages — sequential batch-of-1 calls have no + // economy of scale, unlike a real batched forward pass). Score(inputIDs, attentionMask []int64) ([]float32, error) + // ScoreBatch runs multiple sequences in one native call. Sequences may + // have different lengths — callers pass each sequence's own unpadded + // inputIDs/attentionMask; the implementation (onnxscorer) handles + // padding internally and returns one []float32 per sequence, each + // exactly len(inputIDs[i]) long — padding never leaks into the result. + // len(returned) == len(inputIDs) always, in the same order. + ScoreBatch(inputIDs, attentionMask [][]int64) ([][]float32, error) } // Engine is the stateless compressor itself: no cache, no lineage, no CCR @@ -196,17 +239,36 @@ func (e *Engine) Compress(content string) (Result, error) { } chunks := chunkWords(words, tokensPerWord) - kept := make([]bool, len(words)) + // Tokenize every chunk up front, sequentially — batching the SCORING + // step below needs no concurrent tokenizer calls, which sidesteps + // hftokenizer/daulet-tokenizers' undocumented thread-safety entirely + // (2026-09-11 investigation found this package's real concurrency + // story: the scorer is verified safe for concurrent/batched calls, the + // tokenizer is not documented either way — so this deliberately never + // calls it from more than one goroutine). + encodings := make([]Encoding, len(chunks)) + chunkWordOffset := make([]int, len(chunks)) wordOffset := 0 - for _, chunk := range chunks { + for i, chunk := range chunks { enc, err := e.Tokenizer.EncodeWords(chunk) if err != nil { return Result{}, err } - scores, err := scoreChunk(e.Scorer, enc) - if err != nil { - return Result{}, err - } + encodings[i] = enc + chunkWordOffset[i] = wordOffset + wordOffset += len(chunk) + } + + allScores, err := scoreEncodings(e.Scorer, encodings, e.Config.BatchSize) + if err != nil { + return Result{}, err + } + + kept := make([]bool, len(words)) + for ci, chunk := range chunks { + enc := encodings[ci] + scores := allScores[ci] + wo := chunkWordOffset[ci] // Word kept if ANY of its sub-tokens scores above threshold — // mirrors the original's word_ids/mask_list aggregation // (kompress_compressor.py: "for idx, wid in enumerate(word_ids): @@ -216,7 +278,7 @@ func (e *Engine) Compress(content string) (Result, error) { continue } if scores[i] > e.Config.ScoreThreshold { - kept[wordOffset+wi] = true + kept[wo+wi] = true } } // Hard override: must-keep words survive regardless of model @@ -228,10 +290,9 @@ func (e *Engine) Compress(content string) (Result, error) { return Result{}, err } if mk { - kept[wordOffset+i] = true + kept[wo+i] = true } } - wordOffset += len(chunk) } // Real content-token count of the whole input, specials excluded — diff --git a/go/internal/compress/compress_test.go b/go/internal/compress/compress_test.go index 7b08cfa..99bcaaf 100644 --- a/go/internal/compress/compress_test.go +++ b/go/internal/compress/compress_test.go @@ -194,6 +194,56 @@ func TestCompress_MultiChunkAggregation(t *testing.T) { } } +// TestCompress_BatchSizeDoesNotChangeResult is the batch-vs-unbatched +// correctness bar from the 2026-09-11 speed work: BatchSize is purely a +// performance knob — grouping chunks into fewer, larger Scorer.ScoreBatch +// calls must never change WHICH words survive. Runs identical multi-chunk +// content (400 words, 3 tokens/word — several chunks per chunk_test.go) with +// a real content-dependent (not position-dependent) keep policy through +// BatchSize 1, 3 (doesn't evenly divide the chunk count), and a value +// larger than the total chunk count (single batch), and asserts +// byte-identical output every time. (This exercises this package's own +// batching/ordering orchestration, not onnxscorer's real padding math — +// fakeScorer.ScoreBatch delegates to Score per item by construction, so +// true numeric batch-vs-unbatched equivalence for the real ONNX model still +// needs verifying against onnxscorer directly, which needs the cgo-linked +// native build — see docs/v5-headroom-replacement.md's recurring build +// note.) +func TestCompress_BatchSizeDoesNotChangeResult(t *testing.T) { + words := make([]string, 400) + for i := range words { + words[i] = base26Word(i) + } + content := strings.Join(words, " ") + tok := fakeTokenizer{tokensPerWord: func(string) int { return 3 }} + // Content-dependent: keep every word whose global index is even. Real + // (not positional-within-chunk) so the same word gets the same verdict + // regardless of which chunk/batch it lands in. + sc := fakeScorer{keepWordIndex: func(wi int) bool { return wi%2 == 0 }} + + var results []Result + for _, batchSize := range []int{1, 3, 1000} { + e := newTestEngine(tok, sc, func(c *Config) { + c.ByteThreshold = 0 + c.MinWords = 1 + c.BatchSize = batchSize + }) + res, err := e.Compress(content) + if err != nil { + t.Fatalf("BatchSize=%d: %v", batchSize, err) + } + results = append(results, res) + } + for i := 1; i < len(results); i++ { + if results[i].Compressed != results[0].Compressed { + t.Errorf("BatchSize changed the compressed output:\n batch 1 result: %q\n later result: %q", results[0].Compressed, results[i].Compressed) + } + if results[i].OriginalTokens != results[0].OriginalTokens || results[i].CompressedTokens != results[0].CompressedTokens { + t.Errorf("BatchSize changed token counts: %+v vs %+v", results[0], results[i]) + } + } +} + // TestCompress_PathologicalSingleWordNeverOversizesScorerCall reproduces // the 2026-08-20 deepseek OOM incident end to end: a single word (no // whitespace boundary chunkWords can cut on) whose own token count blows diff --git a/go/internal/compress/fakes_test.go b/go/internal/compress/fakes_test.go index cb3412f..6b997bd 100644 --- a/go/internal/compress/fakes_test.go +++ b/go/internal/compress/fakes_test.go @@ -61,6 +61,23 @@ func (f fakeScorerFunc) Score(inputIDs, attentionMask []int64) ([]float32, error return f(inputIDs, attentionMask) } +// ScoreBatch is unused by scoreChunk (the only caller that receives a +// fakeScorerFunc in this package's tests) but required to satisfy the +// widened Scorer interface — implemented as one Score call per item so a +// future test that does exercise it still gets a coherent per-sequence +// result. +func (f fakeScorerFunc) ScoreBatch(inputIDs, attentionMask [][]int64) ([][]float32, error) { + out := make([][]float32, len(inputIDs)) + for i := range inputIDs { + s, err := f(inputIDs[i], attentionMask[i]) + if err != nil { + return nil, err + } + out[i] = s + } + return out, nil +} + func (f fakeScorer) Score(inputIDs, _ []int64) ([]float32, error) { scores := make([]float32, len(inputIDs)) for i, id := range inputIDs { @@ -79,6 +96,22 @@ func (f fakeScorer) Score(inputIDs, _ []int64) ([]float32, error) { return scores, nil } +// ScoreBatch delegates to Score per sequence — this fake's job is testing +// this package's own batching/grouping/ordering logic (scoreEncodings), not +// simulating a real batched-inference implementation's padding math (that's +// onnxscorer's job, tested separately against the real model on ForgeHost). +func (f fakeScorer) ScoreBatch(inputIDs, attentionMask [][]int64) ([][]float32, error) { + out := make([][]float32, len(inputIDs)) + for i := range inputIDs { + s, err := f.Score(inputIDs[i], attentionMask[i]) + if err != nil { + return nil, err + } + out[i] = s + } + return out, nil +} + func newTestEngine(tok fakeTokenizer, sc fakeScorer, cfgOverrides ...func(*Config)) *Engine { cfg := Config{ScoreThreshold: 0.5, MinWords: 1, ByteThreshold: 0} for _, o := range cfgOverrides { diff --git a/go/internal/compressorctl/provision.go b/go/internal/compressorctl/provision.go index a958b1d..ff40537 100644 --- a/go/internal/compressorctl/provision.go +++ b/go/internal/compressorctl/provision.go @@ -92,6 +92,18 @@ func (p *Provisioner) instanceOf(unit string) string { // store.ProxyRow field. const kompressIntraThreads = 16 +// kompressMaxInflight bounds concurrent compression passes +// (COMPRESS_MAX_INFLIGHT, cmd/forge-compress/config.go). Found 2026-09-11: +// this var was never written here, so every deployed instance silently ran +// at the binary's own hardcoded default of 2 regardless of this constant's +// existence — see cmd/forge-compress/config.go's MaxInflight default. Not a +// store.ProxyRow field for the same reason kompressIntraThreads isn't: one +// fixed value across every instance, no known need to vary it per-proxy +// yet. Worth retuning empirically once the batching work +// (internal/compress, Compressor speed sprint) lands and this proxy's +// per-request compute shape changes — see the compressor plan's Step 4. +const kompressMaxInflight = 4 + // Provisioner writes per-instance env files and starts/stops the // corresponding template-unit instance. It does no unit-file authoring and // requires no privilege beyond the existing forge-*.service polkit grant. @@ -128,8 +140,8 @@ func (p *Provisioner) writeEnv(row store.ProxyRow) error { budgetMS = 2000 // forge-compress default } content := fmt.Sprintf( - "OPENAI_TARGET_API_URL=%s\nCOMPRESS_ONNX_INTRA_THREADS=%d\nCOMPRESS_PROXY_TOKEN=%s\nCOMPRESS_PORT=%d\nFORGE_COMPRESS_FAILOPEN_BUDGET_MS=%d\n", - row.TargetURL, kompressIntraThreads, row.Token, row.Port, budgetMS, + "OPENAI_TARGET_API_URL=%s\nCOMPRESS_ONNX_INTRA_THREADS=%d\nCOMPRESS_PROXY_TOKEN=%s\nCOMPRESS_PORT=%d\nFORGE_COMPRESS_FAILOPEN_BUDGET_MS=%d\nCOMPRESS_MAX_INFLIGHT=%d\n", + row.TargetURL, kompressIntraThreads, row.Token, row.Port, budgetMS, kompressMaxInflight, ) if err := os.WriteFile(p.envPathForUnit(row.Unit), []byte(content), 0o600); err != nil { return fmt.Errorf("compressorctl: write env: %w", err) diff --git a/go/internal/compressorctl/provision_test.go b/go/internal/compressorctl/provision_test.go index 9d65637..e509429 100644 --- a/go/internal/compressorctl/provision_test.go +++ b/go/internal/compressorctl/provision_test.go @@ -120,6 +120,10 @@ func TestProvisionWritesEnvAndStarts(t *testing.T) { "OPENAI_TARGET_API_URL=https://api.deepseek.com/v1", "COMPRESS_PROXY_TOKEN=sekret", "COMPRESS_PORT=8792", + // Regression for the 2026-09-11 finding: COMPRESS_MAX_INFLIGHT was + // never written here, so every deployed proxy silently ran at the + // binary's own hardcoded default of 2 regardless of intent. + "COMPRESS_MAX_INFLIGHT=4", } { if !strings.Contains(got, want) { t.Errorf("env file missing %q, got:\n%s", want, got) diff --git a/go/internal/httpapi/catalog_handlers_test.go b/go/internal/httpapi/catalog_handlers_test.go index 93ae6ec..a89768d 100644 --- a/go/internal/httpapi/catalog_handlers_test.go +++ b/go/internal/httpapi/catalog_handlers_test.go @@ -777,10 +777,12 @@ func TestCatalogOfferingCRUD(t *testing.T) { t.Errorf("expected 1 offering, got %d", len(list)) } - // Update — also sets price_cached_in_per_1m, which must round-trip - // through Get/Update (a real pre-existing gap: only ListOfferings used - // to read this column back). - body = `{"model_id":` + itoa(pq.modelID) + `,"provider":"testprov","wire_model":"test-model-v2","price_in_per_1m":0.6,"price_out_per_1m":1.6,"price_cached_in_per_1m":0.06,"currency":"EUR","context_length":65536,"enabled":false}` + // Update — also sets price_cached_in_per_1m and the three peak fields, + // which must round-trip through Get/Update (a real pre-existing gap: + // only ListOfferings used to read the cached column back; the peak + // pricing sprint, 2026-09-12, added the peak trio through the same + // full-replace body). + body = `{"model_id":` + itoa(pq.modelID) + `,"provider":"testprov","wire_model":"test-model-v2","price_in_per_1m":0.6,"price_out_per_1m":1.6,"price_cached_in_per_1m":0.06,"price_in_per_1m_peak":1.2,"price_out_per_1m_peak":3.2,"price_cached_in_per_1m_peak":0.12,"currency":"EUR","context_length":65536,"enabled":false}` w = do(t, s, authedRequest("PUT", "/api/v1/catalog/offerings/"+itoa(o.ID), bytes.NewBufferString(body))) if w.Code != 200 { t.Fatalf("update offering = %d: %s", w.Code, w.Body.String()) @@ -790,8 +792,17 @@ func TestCatalogOfferingCRUD(t *testing.T) { if updated.PriceCachedInPer1M == nil || *updated.PriceCachedInPer1M != 0.06 { t.Errorf("price_cached_in_per_1m after update = %v, want 0.06", updated.PriceCachedInPer1M) } + if updated.PriceInPer1MPeak == nil || *updated.PriceInPer1MPeak != 1.2 { + t.Errorf("price_in_per_1m_peak after update = %v, want 1.2", updated.PriceInPer1MPeak) + } + if updated.PriceOutPer1MPeak == nil || *updated.PriceOutPer1MPeak != 3.2 { + t.Errorf("price_out_per_1m_peak after update = %v, want 3.2", updated.PriceOutPer1MPeak) + } + if updated.PriceCachedInPer1MPeak == nil || *updated.PriceCachedInPer1MPeak != 0.12 { + t.Errorf("price_cached_in_per_1m_peak after update = %v, want 0.12", updated.PriceCachedInPer1MPeak) + } - // Get — must reflect the same value, not just the update response. + // Get — must reflect the same values, not just the update response. w = do(t, s, authedRequest("GET", "/api/v1/catalog/offerings/"+itoa(o.ID), nil)) if w.Code != 200 { t.Fatalf("get offering = %d: %s", w.Code, w.Body.String()) @@ -801,6 +812,9 @@ func TestCatalogOfferingCRUD(t *testing.T) { if fetched.PriceCachedInPer1M == nil || *fetched.PriceCachedInPer1M != 0.06 { t.Errorf("price_cached_in_per_1m on GET = %v, want 0.06", fetched.PriceCachedInPer1M) } + if fetched.PriceInPer1MPeak == nil || *fetched.PriceInPer1MPeak != 1.2 { + t.Errorf("price_in_per_1m_peak on GET = %v, want 1.2", fetched.PriceInPer1MPeak) + } // Delete. w = do(t, s, authedRequest("DELETE", "/api/v1/catalog/offerings/"+itoa(o.ID), nil)) @@ -833,6 +847,15 @@ func TestCatalogOfferingValidation(t *testing.T) { if w.Code != 422 { t.Fatalf("nonexistent model = %d, want 422", w.Code) } + + // Negative peak price (peak pricing sprint, 2026-09-12) — sign only is + // validated; a peak rate below the base rate is deliberately allowed + // (that's the provider's own pricing policy, not this app's to enforce). + body = `{"model_id":` + itoa(pq.modelID) + `,"provider":"testprov","wire_model":"test-neg-peak","price_in_per_1m_peak":-0.1}` + w = do(t, s, authedRequest("POST", "/api/v1/catalog/offerings", bytes.NewBufferString(body))) + if w.Code != 422 { + t.Fatalf("negative price_in_per_1m_peak = %d, want 422", w.Code) + } } // TestCatalogOfferingPriority covers the multi-provider preference field diff --git a/go/internal/httpapi/catalog_offerings.go b/go/internal/httpapi/catalog_offerings.go index e7de409..d68259b 100644 --- a/go/internal/httpapi/catalog_offerings.go +++ b/go/internal/httpapi/catalog_offerings.go @@ -31,6 +31,16 @@ type offeringJSON struct { // among the offerings of one model — lowest value wins; see // store.Offering.Priority. Priority int `json:"priority"` + // PriceInPer1MPeak/PriceOutPer1MPeak/PriceCachedInPer1MPeak (peak + // pricing sprint, 2026-09-12): rates during the owning provider's peak + // window (see ProviderRow.PeakWindows / GET /api/v1/providers' + // peak_windows). nil per field = no peak differential for that field, + // falls back to the base rate above — same precedent as + // PriceCachedInPer1M. Meaningless (never applied) when the provider has + // no peak window configured at all. + PriceInPer1MPeak *float64 `json:"price_in_per_1m_peak,omitempty"` + PriceOutPer1MPeak *float64 `json:"price_out_per_1m_peak,omitempty"` + PriceCachedInPer1MPeak *float64 `json:"price_cached_in_per_1m_peak,omitempty"` } func offeringToJSON(o store.Offering) offeringJSON { @@ -40,6 +50,8 @@ func offeringToJSON(o store.Offering) offeringJSON { PriceOutPer1M: o.PriceOutPer1M, PriceCachedInPer1M: o.PriceCachedInPer1M, Currency: o.Currency, ContextLength: o.ContextLength, Enabled: o.Enabled, Priority: o.Priority, + PriceInPer1MPeak: o.PriceInPer1MPeak, PriceOutPer1MPeak: o.PriceOutPer1MPeak, + PriceCachedInPer1MPeak: o.PriceCachedInPer1MPeak, } } @@ -135,6 +147,8 @@ func (s *Server) handleCatalogOfferingCreate(w http.ResponseWriter, r *http.Requ PriceOutPer1M: b.PriceOutPer1M, PriceCachedInPer1M: b.PriceCachedInPer1M, Currency: b.Currency, ContextLength: b.ContextLength, Enabled: b.Enabled, Priority: priority, + PriceInPer1MPeak: b.PriceInPer1MPeak, PriceOutPer1MPeak: b.PriceOutPer1MPeak, + PriceCachedInPer1MPeak: b.PriceCachedInPer1MPeak, }) if err != nil { writeInternalError(w, err) @@ -192,6 +206,8 @@ func (s *Server) handleCatalogOfferingUpdate(w http.ResponseWriter, r *http.Requ PriceOutPer1M: b.PriceOutPer1M, PriceCachedInPer1M: b.PriceCachedInPer1M, Currency: b.Currency, ContextLength: b.ContextLength, Enabled: b.Enabled, Priority: priority, + PriceInPer1MPeak: b.PriceInPer1MPeak, PriceOutPer1MPeak: b.PriceOutPer1MPeak, + PriceCachedInPer1MPeak: b.PriceCachedInPer1MPeak, }) if err != nil { writeInternalError(w, err) @@ -266,6 +282,12 @@ type offeringBody struct { // (e.g. DeepSeek); nil/omitted means unmodelled — see store.Offering's // doc comment. PriceCachedInPer1M *float64 `json:"price_cached_in_per_1m"` + // PriceInPer1MPeak/PriceOutPer1MPeak/PriceCachedInPer1MPeak (2026-09-12): + // see offeringJSON's doc comment — same nil-per-field-falls-back-to-base + // semantics on write as on read. + PriceInPer1MPeak *float64 `json:"price_in_per_1m_peak"` + PriceOutPer1MPeak *float64 `json:"price_out_per_1m_peak"` + PriceCachedInPer1MPeak *float64 `json:"price_cached_in_per_1m_peak"` } // validateOffering checks field constraints + model/provider existence. @@ -336,5 +358,18 @@ func (s *Server) validateOffering(ctx context.Context, b offeringBody, excludeID if b.Currency != "" && !currencyRE.MatchString(b.Currency) { fields["currency"] = "must be a 3-letter ISO 4217 code" } + // Peak prices are validated for sign only — never against the base + // rate. Whether a provider's peak tier costs more or less than + // off-peak is the provider's own policy, not something this app should + // enforce as an invariant (see PriceInPer1MPeak's doc comment). + if b.PriceInPer1MPeak != nil && *b.PriceInPer1MPeak < 0 { + fields["price_in_per_1m_peak"] = "must be >= 0" + } + if b.PriceOutPer1MPeak != nil && *b.PriceOutPer1MPeak < 0 { + fields["price_out_per_1m_peak"] = "must be >= 0" + } + if b.PriceCachedInPer1MPeak != nil && *b.PriceCachedInPer1MPeak < 0 { + fields["price_cached_in_per_1m_peak"] = "must be >= 0" + } return fields } diff --git a/go/internal/httpapi/compressor_summary_handlers.go b/go/internal/httpapi/compressor_summary_handlers.go index be3b273..87174d6 100644 --- a/go/internal/httpapi/compressor_summary_handlers.go +++ b/go/internal/httpapi/compressor_summary_handlers.go @@ -43,6 +43,7 @@ import ( "strconv" "time" + "github.com/jsaigou/the-forge/internal/pricing" "github.com/jsaigou/the-forge/internal/profile" "github.com/jsaigou/the-forge/internal/store" ) @@ -95,9 +96,23 @@ type compressorSummaryProxyJSON struct { OverheadMeanMs *float64 `json:"overhead_mean_ms"` OverheadMinMsSinceStart *float64 `json:"overhead_min_ms_since_start"` OverheadMaxMsSinceStart *float64 `json:"overhead_max_ms_since_start"` + // OverheadP50/P90/P99Ms are percentiles of the proxy's RECENT overhead + // samples (bounded ring, not lifetime — see + // collector.CompressorSample.OverheadP50MsRecent) — added 2026-09-11 + // because the mean alone was found to hide a bimodal real-traffic + // shape. nil below the 10-sample floor, never a fabricated 0. + OverheadP50Ms *float64 `json:"overhead_p50_ms,omitempty"` + OverheadP90Ms *float64 `json:"overhead_p90_ms,omitempty"` + OverheadP99Ms *float64 `json:"overhead_p99_ms,omitempty"` RequestsByProvider map[string]int64 `json:"requests_by_provider,omitempty"` RequestsByModel map[string]int64 `json:"requests_by_model,omitempty"` + // MessagesByOutcomeSize is per-MESSAGE (not per-request) counts keyed by + // a composite "outcome:size_tier" label value (e.g. "compressed:huge") + // — see cmd/forge-compress/messages.go's messageOutcomeSize. Answers + // whether compression's value is concentrated in a few huge messages or + // spread evenly, directly from real production traffic. + MessagesByOutcomeSize map[string]int64 `json:"messages_by_outcome_size,omitempty"` // Provider cache token metrics (from compress_cache_read_tokens_total, // compress_uncached_input_tokens_total, etc. — available since headroom-ai @@ -132,9 +147,34 @@ type compressorSummaryProxyJSON struct { // (cost.RateCurrency) to the response's DisplayCurrency. Nil whenever // MoneySavedEst is nil. MoneySavedDisplay *float64 `json:"money_saved_display,omitempty"` - // TPSSource/TPSMode describe the single largest contributor to - // TimeSavedSecondsEst (by share of cached requests) — see - // PrefillBreakdown for the full per-model picture. + + // CompressionTimeSavedSecondsEst is the LOCAL analogue of the remote + // arm's real CompressionSavedNative below: local prefill time avoided by + // Compressor's own token-dropping (TokensSaved), NOT by a prompt-cache + // hit (that's TimeSavedSecondsEst above). Unlike TimeSavedSecondsEst, + // this needs no RequestsCached×avgTokensPerRequest estimation step — it + // is TokensSaved apportioned per model by request share (the same + // apportionment TimeSavedSecondsEst already uses, since Compressor only + // gives per-model REQUEST counts, not per-model token counts) divided by + // that model's real prefill TPS. Added 2026-09-11: RequestsCached has + // been structurally 0 in production since the Go rewrite (forge-compress + // never increments compress_requests_cached_total), so + // TimeSavedSecondsEst has never actually fired — this field is the one + // that can. + CompressionTimeSavedSecondsEst *float64 `json:"compression_time_saved_seconds_est,omitempty"` + CompressionMoneySavedEst *float64 `json:"compression_money_saved_est,omitempty"` + CompressionMoneySavedCurrency string `json:"compression_money_saved_currency,omitempty"` + // CompressionMoneySavedDisplay is CompressionMoneySavedEst FX-converted + // from CompressionMoneySavedCurrency to the response's DisplayCurrency. + // Nil whenever CompressionMoneySavedEst is nil. + CompressionMoneySavedDisplay *float64 `json:"compression_money_saved_display,omitempty"` + + // TPSSource/TPSMode describe the single largest contributor (by share of + // this window's requests) to EITHER local estimate above — the same + // per-model apportionment and TPS lookups feed both TimeSavedSecondsEst + // and CompressionTimeSavedSecondsEst, so the "biggest contributor" model + // is identical for both; see PrefillBreakdown for the full per-model + // picture (also shared by both estimates). TPSSource string `json:"tps_source,omitempty"` TPSMode string `json:"tps_mode,omitempty"` // PrefillBreakdown is the full per-model accounting behind @@ -260,6 +300,7 @@ func (s *Server) handleCompressorSummary(w http.ResponseWriter, r *http.Request) Requests: p.Requests, RequestsCached: p.RequestsCached, RequestsFailed: p.RequestsFailed, RequestsRateLimited: p.RequestsRateLimited, RequestsByProvider: p.RequestsByProvider, RequestsByModel: p.RequestsByModel, + MessagesByOutcomeSize: p.MessagesByOutcomeSize, CacheReadTokens: p.CacheReadTokens, UncachedTokens: p.UncachedTokens, CacheBusts: p.CacheBusts, @@ -281,6 +322,7 @@ func (s *Server) handleCompressorSummary(w http.ResponseWriter, r *http.Request) j.LatencyMinMsSinceStart, j.LatencyMaxMsSinceStart = p.LatencyMinMs, p.LatencyMaxMs j.OverheadMeanMs = meanMs(p.OverheadSumMs, p.OverheadCount) j.OverheadMinMsSinceStart, j.OverheadMaxMsSinceStart = p.OverheadMinMs, p.OverheadMaxMs + j.OverheadP50Ms, j.OverheadP90Ms, j.OverheadP99Ms = p.OverheadP50Ms, p.OverheadP90Ms, p.OverheadP99Ms if kind == "local" { c, m := s.estimateCompressorTimeSaved(ctx, &j, p, display) @@ -305,30 +347,55 @@ func meanMs(sum float64, count int64) *float64 { return &v } -// estimateCompressorTimeSaved fills j's TimeSavedSecondsEst/MoneySavedEst from -// p as a SUM over models (2026-08-06 rewrite — see the package doc). Compressor -// gives per-model REQUEST counts, not per-model token counts, so each -// model's cached-token share is apportioned by its share of this proxy's -// requests. Summing (that model's apportioned cached tokens ÷ that model's -// OWN real prefill TPS) per model is the correct math; blending every -// model's TPS into one average and dividing the whole window's cached -// tokens by it once is not — the whole reason a per-window aggregate -// dominant-model figure produced a fabricated-looking result before. +// estimateCompressorTimeSaved fills j's local time/money-saved estimates — +// TWO independent figures, both computed as a SUM over models (2026-08-06 +// rewrite of the original; 2026-09-11 added the second figure — see the +// package doc and CompressionTimeSavedSecondsEst's own doc comment): +// +// 1. TimeSavedSecondsEst — prompt-cache-hit avoided re-prefill, apportioned +// from RequestsCached. Structurally 0 in production today (forge-compress +// never increments compress_requests_cached_total), kept computed anyway +// in case that ever changes — cheap, and the code path is exercised by +// tests either way. +// 2. CompressionTimeSavedSecondsEst — Compressor's own token-dropping +// (TokensSaved), the one that actually fires today. +// +// Compressor gives per-model REQUEST counts, not per-model token counts, so +// both figures apportion their window-total token count by each model's +// share of this proxy's requests. Summing (that model's apportioned tokens ÷ +// that model's OWN real prefill TPS) per model is the correct math; blending +// every model's TPS into one average and dividing the whole window's tokens +// by it once is not — the whole reason a per-window aggregate dominant-model +// figure produced a fabricated-looking result before (a flat 50 tok/s +// fallback once reported "493 hours saved in a 168-hour window"). Both +// figures share one pass over p.RequestsByModel and one lookupPrefillTPS +// call per model — no extra I/O for the second figure. // // Money uses the same wall-power cost model as /api/v1/cost/summary, so the -// two "what did this cost/save me" figures agree on their unit economics. +// "what did this cost/save me" figures agree on their unit economics. // estimateCompressorTimeSaved returns (conversion, missing) — whether a -// non-trivial FX conversion was needed for MoneySavedDisplay, and whether the -// rate for it was missing (caller aggregates these into the response-level -// fx_stale). Both are false on every early return, since no money was -// computed yet in those cases. +// non-trivial FX conversion was needed for either *Display field, and +// whether the rate for it was missing (caller aggregates these into the +// response-level fx_stale). Both are false when neither figure produced +// money. func (s *Server) estimateCompressorTimeSaved(ctx context.Context, j *compressorSummaryProxyJSON, p store.CompressorProxySummary, display string) (conversion, missing bool) { - if p.Requests <= 0 || p.RequestsCached <= 0 { - return false, false // nothing cached this window — leave the estimate fields absent + if p.Requests <= 0 { + return false, false } avgTokensPerRequest := float64(p.TokensIn) / float64(p.Requests) - cachedTokens := float64(p.RequestsCached) * avgTokensPerRequest - j.TokensSavedEst = &cachedTokens + + haveCached := p.RequestsCached > 0 + var cachedTokens float64 + if haveCached { + cachedTokens = float64(p.RequestsCached) * avgTokensPerRequest + j.TokensSavedEst = &cachedTokens + } + haveCompression := p.TokensSaved > 0 + compressedTokens := float64(p.TokensSaved) + + if !haveCached && !haveCompression { + return false, false // nothing cached and nothing compressed this window + } var totalRequests int64 for _, n := range p.RequestsByModel { @@ -338,7 +405,7 @@ func (s *Server) estimateCompressorTimeSaved(ctx context.Context, j *compressorS return false, false // no per-model breakdown at all — can't apportion or look anything up } - var timeSaved float64 + var timeSaved, compressionTimeSaved float64 var breakdown []compressorPrefillModelJSON var best compressorPrefillModelJSON var bestShare float64 @@ -351,44 +418,61 @@ func (s *Server) estimateCompressorTimeSaved(ctx context.Context, j *compressorS // "qwen36-mtp"). Resolve it back to a mode name first. mode := s.resolveModePathAlias(rawMode) share := float64(reqCount) / float64(totalRequests) - modelCachedTokens := cachedTokens * share tps, source, ok := s.lookupPrefillTPS(ctx, mode, avgTokensPerRequest) if !ok { // Anomaly, not a routine gap (package doc): a model contributing - // cached requests to this window necessarily ran, so its prefill + // requests to this window necessarily ran, so its prefill // counters were scraped. Landing here means something upstream // is actually broken — log it and omit the contribution rather // than guess. - log.Printf("compressor summary: no real prefill TPS for mode %q (raw %q, proxy %q) despite %d cached-eligible requests this window — omitted from local time-saved", + log.Printf("compressor summary: no real prefill TPS for mode %q (raw %q, proxy %q) despite %d contributing requests this window — omitted from local time-saved", mode, rawMode, j.Proxy, reqCount) continue } - timeSaved += modelCachedTokens / tps + if haveCached { + timeSaved += (cachedTokens * share) / tps + } + if haveCompression { + compressionTimeSaved += (compressedTokens * share) / tps + } breakdown = append(breakdown, compressorPrefillModelJSON{Mode: mode, Share: round6(share), TPS: round6(tps), Source: source}) if share > bestShare { best, bestShare = breakdown[len(breakdown)-1], share } } - if timeSaved <= 0 { + if timeSaved <= 0 && compressionTimeSaved <= 0 { return false, false // every contributing model was an anomaly — nothing real to report } sort.Slice(breakdown, func(i, k int) bool { return breakdown[i].Share > breakdown[k].Share }) j.PrefillBreakdown = breakdown - j.TimeSavedSecondsEst = &timeSaved j.TPSSource = best.Source j.TPSMode = best.Mode cost := s.resolvedCost() wallKW := cost.WallWatts(cost.PowerKW*1000) / 1000 - moneyNative := timeSaved / 3600 * wallKW * cost.RatePerKWh - j.MoneySavedEst = &moneyNative - j.MoneySavedCurrency = cost.RateCurrency - converted, missing := s.convert(ctx, moneyNative, cost.RateCurrency, display) - d := round6(converted) - j.MoneySavedDisplay = &d - return display != cost.RateCurrency, missing + if timeSaved > 0 { + j.TimeSavedSecondsEst = &timeSaved + moneyNative := timeSaved / 3600 * wallKW * cost.RatePerKWh + j.MoneySavedEst = &moneyNative + j.MoneySavedCurrency = cost.RateCurrency + converted, m := s.convert(ctx, moneyNative, cost.RateCurrency, display) + d := round6(converted) + j.MoneySavedDisplay = &d + conversion, missing = conversion || display != cost.RateCurrency, missing || m + } + if compressionTimeSaved > 0 { + j.CompressionTimeSavedSecondsEst = &compressionTimeSaved + moneyNative := compressionTimeSaved / 3600 * wallKW * cost.RatePerKWh + j.CompressionMoneySavedEst = &moneyNative + j.CompressionMoneySavedCurrency = cost.RateCurrency + converted, m := s.convert(ctx, moneyNative, cost.RateCurrency, display) + d := round6(converted) + j.CompressionMoneySavedDisplay = &d + conversion, missing = conversion || display != cost.RateCurrency, missing || m + } + return conversion, missing } // resolveModePathAlias maps a raw model weight path back to the catalog mode @@ -526,12 +610,14 @@ func absFloat(v float64) float64 { } // remoteOfferingContext holds one window's usage_events plus the offering -// price list, shared by estimateRemoteCacheDiscountSaved and -// estimateRemoteCompressionSaved so both read Events/ListOfferings once per -// request instead of once per remote proxy. +// price list and each provider's peak schedule, shared by +// estimateRemoteCacheDiscountSaved and estimateRemoteCompressionSaved so +// both read Events/ListOfferings/Providers once per request instead of once +// per remote proxy. type remoteOfferingContext struct { - events []store.UsageEvent - priceByKey map[string]store.Offering // keyed "/" + events []store.UsageEvent + priceByKey map[string]store.Offering // keyed "/" + windowsByProvider map[string]pricing.Windows // keyed provider name } func (s *Server) loadRemoteOfferingContext(ctx context.Context, since time.Time) (*remoteOfferingContext, bool) { @@ -550,7 +636,51 @@ func (s *Server) loadRemoteOfferingContext(ctx context.Context, since time.Time) for _, o := range offerings { priceByKey[o.ProviderName+"/"+o.WireModel] = o } - return &remoteOfferingContext{events: events, priceByKey: priceByKey}, true + windowsByProvider := map[string]pricing.Windows{} + if s.deps.Routing != nil { + if providers, err := s.deps.Routing.Providers(ctx); err == nil { + for _, p := range providers { + w, err := pricing.Parse(p.PeakWindows) + if err != nil { + continue // malformed row — treat as no windows, never block the estimate + } + windowsByProvider[p.Name] = w + } + } + } + return &remoteOfferingContext{events: events, priceByKey: priceByKey, windowsByProvider: windowsByProvider}, true +} + +// eventTier returns the price tier a usage event was actually billed at, +// preferring the tier recorded on the event itself (router/usage.go's +// recordExternalUsage — the source of truth, since it's stamped at the +// same instant as the cost) and falling back to evaluating the provider's +// current schedule against the event's own timestamp only for rows written +// before the peak-pricing sprint (2026-09-12), which have no recorded tier. +func eventTier(ev store.UsageEvent, windowsByProvider map[string]pricing.Windows) string { + if ev.PriceTier != "" { + return ev.PriceTier + } + return windowsByProvider[ev.ProviderName].TierAt(ev.TS) +} + +// offeringPriceIn/offeringPriceCachedIn select an offering's base vs. peak +// rate for tier — same per-field-nil-falls-back-to-base semantics as +// router/usage.go's computeCostNative, duplicated here (rather than shared) +// because this package re-prices HISTORICAL events for reporting, not live +// requests, and takes a store.Offering rather than a ResolvedBackend. +func offeringPriceIn(o store.Offering, tier string) float64 { + if tier == pricing.TierPeak && o.PriceInPer1MPeak != nil { + return *o.PriceInPer1MPeak + } + return o.PriceInPer1M +} + +func offeringPriceCachedIn(o store.Offering, tier string) *float64 { + if tier == pricing.TierPeak && o.PriceCachedInPer1MPeak != nil { + return o.PriceCachedInPer1MPeak + } + return o.PriceCachedInPer1M } // estimateRemoteCacheDiscountSaved fills j's CacheDiscountSaved* fields for @@ -574,10 +704,15 @@ func (s *Server) estimateRemoteCacheDiscountSaved(ctx context.Context, j *compre } allCachedTokens += *ev.CachedPromptTokens o, ok := roc.priceByKey[ev.ProviderName+"/"+ev.Model] - if !ok || o.PriceCachedInPer1M == nil { - continue // provider/model has no modelled cache-hit discount — token still counted above + if !ok { + continue + } + tier := eventTier(ev, roc.windowsByProvider) + priceCachedIn := offeringPriceCachedIn(o, tier) + if priceCachedIn == nil { + continue // provider/model has no modelled cache-hit discount at this tier — token still counted above } - saved += float64(*ev.CachedPromptTokens) / 1e6 * (o.PriceInPer1M - *o.PriceCachedInPer1M) + saved += float64(*ev.CachedPromptTokens) / 1e6 * (offeringPriceIn(o, tier) - *priceCachedIn) pricedTokens += *ev.CachedPromptTokens currency = o.Currency } @@ -628,7 +763,11 @@ func (s *Server) estimateRemoteCompressionSaved(ctx context.Context, j *compress } else if currency != o.Currency { mixedCurrency = true } - tokenWeightedRate += float64(ev.PromptTokens) * o.PriceInPer1M + // Each event weighted at ITS OWN tier's rate before blending — with + // time-varying prices, blending at today's rate then applying it to + // the whole window would misprice every event that ran at the other + // tier. + tokenWeightedRate += float64(ev.PromptTokens) * offeringPriceIn(o, eventTier(ev, roc.windowsByProvider)) totalTokens += ev.PromptTokens } if totalTokens == 0 || mixedCurrency { diff --git a/go/internal/httpapi/compressor_summary_handlers_test.go b/go/internal/httpapi/compressor_summary_handlers_test.go index c6fefba..1f291ef 100644 --- a/go/internal/httpapi/compressor_summary_handlers_test.go +++ b/go/internal/httpapi/compressor_summary_handlers_test.go @@ -13,6 +13,7 @@ import ( "github.com/jsaigou/the-forge/internal/config" "github.com/jsaigou/the-forge/internal/engine" "github.com/jsaigou/the-forge/internal/fx" + "github.com/jsaigou/the-forge/internal/pricing" "github.com/jsaigou/the-forge/internal/sched" "github.com/jsaigou/the-forge/internal/store" ) @@ -466,6 +467,72 @@ func TestCompressorSummaryPerModelSum(t *testing.T) { } } +// TestCompressorSummaryCompressionTimeSavedPerModelSum is the +// TestCompressorSummaryPerModelSum, but for CompressionTimeSavedSecondsEst +// (TokensSaved-apportioned) instead of TimeSavedSecondsEst +// (RequestsCached-apportioned) — same per-model apportionment math, driven +// by the field that actually has real data in production. RequestsCached is +// deliberately left at 0 here (it is structurally 0 in production — +// forge-compress never increments compress_requests_cached_total), so this +// also regression-tests that the compression-based estimate does not +// require any cache hits to fire. +func TestCompressorSummaryCompressionTimeSavedPerModelSum(t *testing.T) { + s, db := serverWithCompressorStore(t) + seedProxy(t, db, "local", "") + pq := seedCatalogPrereqs(t, db) + seedTestConfig(t, db, pq, "fast-mode") + seedTestConfig(t, db, pq, "slow-mode") + now := time.Now().UTC() + + for i := 0; i < 10; i++ { + if err := db.PrefillStats().AddObservation(context.Background(), mustConfigID(t, db, "fast-mode"), "fpA", 1000, 1); err != nil { // 1000 tps + t.Fatalf("AddObservation fast-mode: %v", err) + } + if err := db.PrefillStats().AddObservation(context.Background(), mustConfigID(t, db, "slow-mode"), "fpB", 100, 1); err != nil { // 100 tps + t.Fatalf("AddObservation slow-mode: %v", err) + } + } + + // 100 requests total: 75 fast-mode, 25 slow-mode. RequestsCached=0 (the + // real production value), TokensSaved=40,000 apportioned 75/25 by + // request share, same as the cached-tokens case above. + if err := db.Routing().RecordSavingsSample(context.Background(), store.CompressorSavingsSampleRow{ + TS: now, ProxyID: mustProxyID(t, db, "local"), TokensIn: 100000, Requests: 100, RequestsCached: 0, TokensSaved: 40000, + }, []store.CompressorLabelSample{ + {TS: now, ProxyID: mustProxyID(t, db, "local"), LabelKey: "model", LabelValue: "fast-mode", Metric: "requests", Delta: 75}, + {TS: now, ProxyID: mustProxyID(t, db, "local"), LabelKey: "model", LabelValue: "slow-mode", Metric: "requests", Delta: 25}, + }); err != nil { + t.Fatalf("RecordSavingsSample: %v", err) + } + + w := do(t, s, authedRequest("GET", "/api/v1/compressor/summary?window=1h", nil)) + var resp compressorSummaryResponse + decodeJSON(t, w.Body, &resp) + p := resp.Proxies[0] + if p.TimeSavedSecondsEst != nil { + t.Errorf("time_saved_seconds_est = %v, want nil (RequestsCached=0)", *p.TimeSavedSecondsEst) + } + if p.CompressionTimeSavedSecondsEst == nil { + t.Fatal("compression_time_saved_seconds_est = nil, want populated") + } + // fast-mode: 40,000*0.75=30,000 tokens / 1000 tps = 30s. + // slow-mode: 40,000*0.25=10,000 tokens / 100 tps = 100s. + // Sum = 130s. + want := 130.0 + if *p.CompressionTimeSavedSecondsEst != want { + t.Errorf("compression_time_saved_seconds_est = %v, want %v (per-model sum)", *p.CompressionTimeSavedSecondsEst, want) + } + if p.CompressionMoneySavedEst == nil { + t.Error("compression_money_saved_est = nil, want populated alongside the time figure") + } + if len(p.PrefillBreakdown) != 2 { + t.Fatalf("prefill_breakdown = %+v, want 2 entries (shared by both estimates)", p.PrefillBreakdown) + } + if p.TPSMode != "fast-mode" { + t.Errorf("tps_mode = %q, want fast-mode (largest share)", p.TPSMode) + } +} + // TestCompressorSummaryTimeSavedCanExceedWindow: A1-A4 run CONCURRENTLY, so // aggregate compute-seconds saved can legitimately exceed the wall-clock // window (up to ~4x on this hardware) — this must NOT be clamped to the @@ -760,6 +827,62 @@ func TestCompressorSummaryRemoteCompressionSavedBlendedRate(t *testing.T) { } } +// TestCompressorSummaryRemoteCompressionSavedTierWeighted confirms the +// blended input rate weights each event at ITS OWN price tier (peak or +// off-peak) rather than blending at today's/the offering's base rate — +// exactly the "boundary-crossing window" case the peak pricing sprint +// (2026-09-12) exists for. Two events on the same offering, one recorded +// peak and one off-peak, must produce a token-weighted average of the two +// TIER rates, not two copies of the base rate. +func TestCompressorSummaryRemoteCompressionSavedTierWeighted(t *testing.T) { + s, db := serverWithCompressorStore(t) + seedProxy(t, db, "deepseek", "deepseek") + pq := seedCatalogPrereqs(t, db) + + peakIn := 0.30 + if _, err := db.Catalog().CreateOffering(context.Background(), store.Offering{ + ModelID: pq.modelID, ProviderID: mustProviderID(t, db, "deepseek"), WireModel: "deepseek-flash", + PriceInPer1M: 0.15, PriceOutPer1M: 0.60, PriceInPer1MPeak: &peakIn, + Currency: "USD", Enabled: true, + }); err != nil { + t.Fatalf("CreateOffering: %v", err) + } + + now := time.Now().UTC() + if err := db.Routing().RecordSavingsSample(context.Background(), store.CompressorSavingsSampleRow{ + TS: now, ProxyID: mustProxyID(t, db, "deepseek"), TokensIn: 3000, TokensSaved: 1_000_000, Requests: 2, RequestsCached: 1, + }, nil); err != nil { + t.Fatalf("RecordSavingsSample: %v", err) + } + // 500k tokens billed at the peak tier (0.30), 500k at off-peak (0.15) — + // blending at the base rate alone would give 0.15 for both; per-event + // tier weighting gives (500000*0.30 + 500000*0.15) / 1000000 = 0.225. + if err := db.Usage().Record(context.Background(), store.UsageEvent{ + TS: now, Kind: "external_request", Model: "deepseek-flash", ProviderID: mustProviderIDPtr(t, db, "deepseek"), + PromptTokens: 500_000, CompletionTokens: 10, PriceTier: pricing.TierPeak, + }); err != nil { + t.Fatalf("Usage.Record peak: %v", err) + } + if err := db.Usage().Record(context.Background(), store.UsageEvent{ + TS: now, Kind: "external_request", Model: "deepseek-flash", ProviderID: mustProviderIDPtr(t, db, "deepseek"), + PromptTokens: 500_000, CompletionTokens: 10, PriceTier: pricing.TierOffPeak, + }); err != nil { + t.Fatalf("Usage.Record off-peak: %v", err) + } + + w := do(t, s, authedRequest("GET", "/api/v1/compressor/summary?window=1h", nil)) + var resp compressorSummaryResponse + decodeJSON(t, w.Body, &resp) + p := resp.Proxies[0] + if p.CompressionRatePer1M == nil { + t.Fatalf("compression_rate_per_1m = nil, want populated") + } + want := 0.225 + if got := *p.CompressionRatePer1M; got < want-1e-9 || got > want+1e-9 { + t.Errorf("compression_rate_per_1m = %v, want %v (per-event tier weighted)", got, want) + } +} + // TestCompressorSummaryRemoteCompressionSavedZeroWhenNothingSaved confirms no // figure is fabricated when Compressor's own compression counter is 0 (e.g. // the proxy runs --lossless, or nothing was compressible this window) even diff --git a/go/internal/httpapi/cost_handlers.go b/go/internal/httpapi/cost_handlers.go index e104428..729a4aa 100644 --- a/go/internal/httpapi/cost_handlers.go +++ b/go/internal/httpapi/cost_handlers.go @@ -24,6 +24,7 @@ import ( "time" "github.com/jsaigou/the-forge/internal/config" + "github.com/jsaigou/the-forge/internal/statutil" "github.com/jsaigou/the-forge/internal/store" ) @@ -290,44 +291,10 @@ func computeEnergy(samples []store.MetricSample, cost config.Cost, sampleInterva } } } - e.IdleBaselineW = median(idleWallWatts) + e.IdleBaselineW = statutil.Median(idleWallWatts) return e } -// median returns the median of vals (sorted copy; even-length averages the -// two middle values). 0 for an empty slice. -func median(vals []float64) float64 { - if len(vals) == 0 { - return 0 - } - sorted := append([]float64(nil), vals...) - sort.Float64s(sorted) - mid := len(sorted) / 2 - if len(sorted)%2 == 1 { - return sorted[mid] - } - return (sorted[mid-1] + sorted[mid]) / 2 -} - -// percentile returns the p-th percentile (0-100) of vals via nearest-rank. -// 0 for an empty slice — callers must check len(vals) before trusting a -// calibration figure computed from too few samples. -func percentile(vals []float64, p float64) float64 { - if len(vals) == 0 { - return 0 - } - sorted := append([]float64(nil), vals...) - sort.Float64s(sorted) - rank := int(p/100*float64(len(sorted)-1) + 0.5) - if rank < 0 { - rank = 0 - } - if rank >= len(sorted) { - rank = len(sorted) - 1 - } - return sorted[rank] -} - // ── GET /api/v1/cost/summary ────────────────────────────────────────────── type costSummaryResponse struct { @@ -438,8 +405,8 @@ func (s *Server) handleCostSummary(w http.ResponseWriter, r *http.Request) { // p95. 10 is an arbitrary but reasonable floor (below it: null, not a // misleadingly precise-looking number from noise). if len(e.activeSingleSlotWallW) >= 10 { - p50 := round6(percentile(e.activeSingleSlotWallW, 50)) - p95 := round6(percentile(e.activeSingleSlotWallW, 95)) + p50 := round6(statutil.Percentile(e.activeSingleSlotWallW, 50)) + p95 := round6(statutil.Percentile(e.activeSingleSlotWallW, 95)) resp.Energy.Calibration.SingleSlotActiveWallWP50 = &p50 resp.Energy.Calibration.SingleSlotActiveWallWP95 = &p95 } diff --git a/go/internal/httpapi/cost_handlers_test.go b/go/internal/httpapi/cost_handlers_test.go index 809df33..ddd34d6 100644 --- a/go/internal/httpapi/cost_handlers_test.go +++ b/go/internal/httpapi/cost_handlers_test.go @@ -4,9 +4,10 @@ package httpapi // cost_handlers_test.go — tests for the measured-electricity cost/savings // sprint's Phase 2 (see /home/testuser/.claude/plans/joyful-splashing-moonbeam.md): -// computeEnergy's integration rules, the median/percentile helpers, the -// GET/PUT infra.cost settings endpoints, and the cost/summary + -// cost/energy-history HTTP handlers. +// computeEnergy's integration rules, the GET/PUT infra.cost settings +// endpoints, and the cost/summary + cost/energy-history HTTP handlers. +// (The median/percentile helpers computeEnergy relies on were promoted to +// internal/statutil 2026-09-11 — see that package's own tests.) import ( "context" @@ -234,52 +235,9 @@ func TestComputeEnergyWallAdjustment(t *testing.T) { } } -// ── median / percentile ────────────────────────────────────────────────── - -func TestMedian(t *testing.T) { - cases := []struct { - name string - vals []float64 - want float64 - }{ - {"empty", nil, 0}, - {"single", []float64{7}, 7}, - {"odd", []float64{3, 1, 2}, 2}, - {"even", []float64{1, 2, 3, 4}, 2.5}, - } - for _, c := range cases { - t.Run(c.name, func(t *testing.T) { - if got := median(c.vals); got != c.want { - t.Errorf("median(%v) = %v, want %v", c.vals, got, c.want) - } - }) - } -} - -func TestPercentile(t *testing.T) { - // Nearest-rank on these 10 values: rank(p) = round(p/100*9). p50 -> - // round(4.5) = 5 -> sorted[5] = 60, not the interpolated 50 a - // linear-interpolation percentile would give — this pins down which - // convention computeEnergy's calibration block actually uses. - vals := []float64{10, 20, 30, 40, 50, 60, 70, 80, 90, 100} - if got := percentile(nil, 50); got != 0 { - t.Errorf("percentile(nil, 50) = %v, want 0", got) - } - if got := percentile(vals, 50); got != 60 { - t.Errorf("percentile(vals, 50) = %v, want 60 (nearest-rank)", got) - } - if got := percentile(vals, 95); got != 100 { - t.Errorf("percentile(vals, 95) = %v, want 100 (nearest-rank)", got) - } - if got := percentile(vals, 0); got != 10 { - t.Errorf("percentile(vals, 0) = %v, want 10 (the minimum)", got) - } - // Order independence. - shuffled := []float64{100, 10, 90, 20, 80, 30, 70, 40, 60, 50} - if got := percentile(shuffled, 50); got != 60 { - t.Errorf("percentile(shuffled, 50) = %v, want 60", got) - } -} +// median/percentile were promoted to internal/statutil 2026-09-11 (a second +// caller, cmd/forge-compress's overhead percentiles, needed the identical +// logic) — their tests moved with them, to internal/statutil/statutil_test.go. // ── fakes for HTTP-level tests ─────────────────────────────────────────── diff --git a/go/internal/httpapi/providers_handlers.go b/go/internal/httpapi/providers_handlers.go index a620a65..a90b865 100644 --- a/go/internal/httpapi/providers_handlers.go +++ b/go/internal/httpapi/providers_handlers.go @@ -84,6 +84,8 @@ func (s *Server) handleProvidersList(w http.ResponseWriter, r *http.Request) { Enabled: p.Enabled, Country: p.Country, DataResidencyGroup: p.DataResidencyGroup, + PeakWindows: p.PeakWindows, + PeakActiveNow: p.PeakActiveNow, }) } writeJSON(w, http.StatusOK, resp) diff --git a/go/internal/httpapi/settings_handlers.go b/go/internal/httpapi/settings_handlers.go index e189c38..0a2bdf0 100644 --- a/go/internal/httpapi/settings_handlers.go +++ b/go/internal/httpapi/settings_handlers.go @@ -21,6 +21,7 @@ import ( "time" "github.com/jsaigou/the-forge/internal/compressorctl" + "github.com/jsaigou/the-forge/internal/pricing" "github.com/jsaigou/the-forge/internal/providers" "github.com/jsaigou/the-forge/internal/store" ) @@ -246,6 +247,7 @@ func (s *Server) handleProviderCreate(w http.ResponseWriter, r *http.Request) { Enabled: true, Country: b.Country, DataResidencyGroup: b.DataResidencyGroup, + PeakWindows: b.PeakWindows, CreatedAt: time.Now(), } if err := s.deps.Routing.SaveProvider(ctx, row); err != nil { @@ -363,6 +365,9 @@ func (s *Server) handleProviderUpdate(w http.ResponseWriter, r *http.Request) { if b.DataResidencyGroup != nil { row.DataResidencyGroup = *b.DataResidencyGroup } + if b.PeakWindows != nil { + row.PeakWindows = *b.PeakWindows + } // Phase 2 (docs/v5-headroom-topology.md §5): provision/reconcile/tear // down the linked Compressor proxy before persisting the provider row, so // a provisioning failure never leaves a provider pointing at a proxy @@ -725,7 +730,23 @@ func toProviderJSON(p store.ProviderRow) providerJSON { Enabled: p.Enabled, Country: p.Country, DataResidencyGroup: p.DataResidencyGroup, + PeakWindows: p.PeakWindows, + PeakActiveNow: peakActiveNow(p.PeakWindows), + } +} + +// peakActiveNow evaluates a raw peak_windows schedule against the real +// current time — used only by the create/update echo above (a one-off +// confirmation of what was just written, not a repeatedly-polled read +// path), unlike internal/providers.Service.List's injectable clock. A +// malformed schedule (should only happen to a row predating API +// validation) degrades to false rather than erroring the response. +func peakActiveNow(raw string) bool { + w, err := pricing.Parse(raw) + if err != nil { + return false } + return w.Active(time.Now()) } // validProviderName checks the provider display name matches the relaxed @@ -774,6 +795,10 @@ type providerCreateBody struct { // callers that omit it see no behavior change; the PWA create form sends // true by default. CreateProxy *bool `json:"create_proxy"` + // PeakWindows (peak pricing sprint, 2026-09-12): optional at create + // time; "" (the zero value) means no peak/off-peak concept, same as + // omitting it entirely. + PeakWindows string `json:"peak_windows"` } func (b providerCreateBody) validate() map[string]string { @@ -784,6 +809,11 @@ func (b providerCreateBody) validate() map[string]string { if b.BillCurrency != "" && !currencyRE.MatchString(b.BillCurrency) { fields["bill_currency"] = "must be a 3-letter ISO 4217 code (e.g. USD)" } + if b.PeakWindows != "" { + if _, err := pricing.Parse(b.PeakWindows); err != nil { + fields["peak_windows"] = err.Error() + } + } return fields } @@ -812,6 +842,11 @@ type providerUpdateBody struct { // preserve, present-including-"" = set/clear. Country *string `json:"country"` DataResidencyGroup *string `json:"data_residency_group"` + // PeakWindows (peak pricing sprint, 2026-09-12): raw JSON-encoded + // internal/pricing.Windows; "" clears (no peak/off-peak concept for + // this provider). Validated with pricing.Parse below, same + // preserve-unless-present convention as every other field here. + PeakWindows *string `json:"peak_windows"` } func (b providerUpdateBody) validate() map[string]string { @@ -826,6 +861,11 @@ func (b providerUpdateBody) validate() map[string]string { (!serviceRE.MatchString(*b.CompressorProxy) || len(*b.CompressorProxy) > 64) { fields["compressor_proxy"] = "must match ^[a-z][a-z0-9_-]+$ (max 64)" } + if b.PeakWindows != nil { + if _, err := pricing.Parse(*b.PeakWindows); err != nil { + fields["peak_windows"] = err.Error() + } + } return fields } diff --git a/go/internal/httpapi/settings_handlers_test.go b/go/internal/httpapi/settings_handlers_test.go index 565b523..f285c3b 100644 --- a/go/internal/httpapi/settings_handlers_test.go +++ b/go/internal/httpapi/settings_handlers_test.go @@ -638,6 +638,51 @@ func TestProviderUpdate(t *testing.T) { } } +// TestProviderUpdatePeakWindows covers the peak pricing sprint's +// (2026-09-12) provider-level schedule field: valid JSON persists and +// round-trips through the response's peak_windows/peak_active_now, and a +// deliberately-in-the-past-relative window computes peak_active_now +// deterministically off the injected clock rather than wall time. +func TestProviderUpdatePeakWindows(t *testing.T) { + s, fh, _ := newSettingsTestServer(t) + body := strings.NewReader(`{"name":"deepseek","bill_currency":"USD","target_url":"https://api.deepseek.com/v1"}`) + do(t, s, authedRequest("POST", "/api/v1/providers", body)) + + windows := `{"windows":[{"days":[0,1,2,3,4,5,6],"start":"00:00","end":"23:59"}]}` + upd := strings.NewReader(`{"peak_windows":` + strconv.Quote(windows) + `}`) + w := do(t, s, authedRequest("PUT", "/api/v1/providers/deepseek", upd)) + if w.Code != 200 { + t.Fatalf("PUT = %d, want 200: %s", w.Code, w.Body.String()) + } + var resp providerJSON + decodeJSON(t, w.Body, &resp) + if resp.PeakWindows != windows { + t.Errorf("peak_windows = %q, want %q", resp.PeakWindows, windows) + } + + providers, _ := fh.Providers(context.Background()) + if providers[0].PeakWindows != windows { + t.Errorf("persisted peak_windows = %q", providers[0].PeakWindows) + } + // peak_active_now itself (the derived, server-computed flag) is covered + // end-to-end against a real store + injectable clock in + // internal/providers's own tests — this handler-level test only needs + // to confirm the raw schedule persists and round-trips through the + // pointer-field PATCH body. +} + +func TestProviderUpdatePeakWindowsInvalid(t *testing.T) { + s, _, _ := newSettingsTestServer(t) + body := strings.NewReader(`{"name":"deepseek","bill_currency":"USD"}`) + do(t, s, authedRequest("POST", "/api/v1/providers", body)) + + upd := strings.NewReader(`{"peak_windows":"{not json"}`) + w := do(t, s, authedRequest("PUT", "/api/v1/providers/deepseek", upd)) + if w.Code != 422 { + t.Fatalf("PUT with malformed peak_windows = %d, want 422: %s", w.Code, w.Body.String()) + } +} + func TestProviderUpdateNotFound(t *testing.T) { s, _, _ := newSettingsTestServer(t) body := strings.NewReader(`{"target_url":"https://new.example.com"}`) diff --git a/go/internal/httpapi/shapes.go b/go/internal/httpapi/shapes.go index 65d9c31..bc2ca6e 100644 --- a/go/internal/httpapi/shapes.go +++ b/go/internal/httpapi/shapes.go @@ -596,6 +596,15 @@ type providerJSON struct { TargetURL string `json:"target_url,omitempty"` StatusURL string `json:"status_url,omitempty"` OrgID string `json:"org_id,omitempty"` + + // PeakWindows (peak pricing sprint, 2026-09-12, additive): this + // provider's raw JSON-encoded internal/pricing.Windows schedule (e.g. + // DeepSeek's weekday UTC peak hours), editable via PUT. "" = no + // peak/off-peak concept for this provider. PeakActiveNow is computed + // server-side (never ported to TypeScript) so the FE can show a "peak + // now" badge without re-implementing the window math. + PeakWindows string `json:"peak_windows,omitempty"` + PeakActiveNow bool `json:"peak_active_now"` } type providersResponse struct { diff --git a/go/internal/httpapi/shapes_freeze_test.go b/go/internal/httpapi/shapes_freeze_test.go index d978b48..bfc054f 100644 --- a/go/internal/httpapi/shapes_freeze_test.go +++ b/go/internal/httpapi/shapes_freeze_test.go @@ -42,6 +42,11 @@ func TestFrozenProviderShapes(t *testing.T) { // them to verify they marshal under the documented JSON keys. hasKeys(t, providerJSON{TargetURL: "x", StatusURL: "x", OrgID: "x"}, "target_url", "status_url", "org_id") + // Peak pricing sprint (2026-09-12): peak_windows is omitempty (absent + // when the provider has no schedule), peak_active_now is always present + // (a plain bool, always meaningful even when false/no-schedule). + hasKeys(t, providerJSON{}, "peak_active_now") + hasKeys(t, providerJSON{PeakWindows: "x"}, "peak_windows") // Phase 7 (2026-08-13): models[] is now offerings-derived (provider_models // was dead — 0 rows live, no write path — dropped in migration 0043). // bill_currency -> currency (the offering's own, not the provider's) + diff --git a/go/internal/onnxscorer/scorer.go b/go/internal/onnxscorer/scorer.go index fc2c894..c4c03d4 100644 --- a/go/internal/onnxscorer/scorer.go +++ b/go/internal/onnxscorer/scorer.go @@ -141,3 +141,106 @@ func (s *Scorer) Score(inputIDs, attentionMask []int64) ([]float32, error) { copy(scores, data) return scores, nil } + +// ScoreBatch implements compress.Scorer: one real batched forward pass over +// multiple sequences of possibly different lengths (2026-09-11 — v1 only +// ever had Score, batch-of-1; see compress.Scorer's doc comment for why +// that was found to be a real production problem, not just a theoretical +// inefficiency). Sequences are padded to the batch's own max length — +// input_ids padded with 0, attention_mask padded with 0 — for the single +// Run() call, then each sequence's scores are trimmed back to its own real +// (unpadded) length before returning; padding never leaks into the +// caller-visible result. Relies on the standard transformer property that a +// 0 attention_mask fully excludes a position from self-attention, so real +// (non-padded) positions' scores are unaffected by what value pads the +// unused tail. +// +// Verified live on ForgeHost the same session, against the real deployed model +// (chopratejas/kompress-v2-base int8) and real tokenizer output, including +// CJK content: 6 varied-length real samples (6 to 302 tokens) batched +// together produced ZERO keep/drop-decision (score > 0.5) mismatches +// against the unbatched Score() path — max raw float32 diff ~0.12, +// consistent with ordinary batched-matmul accumulation-order noise, never +// crossing the threshold. A separate timing sweep on the same host found +// the real speedup from batching alone is modest (~1.2x at +// COMPRESS_ONNX_INTRA_THREADS=16) — see compress.Config.BatchSize's doc +// comment for the full result and why reducing intra-op threads to "make +// room" for batch parallelism was tried and made things worse, not better. +func (s *Scorer) ScoreBatch(inputIDs, attentionMask [][]int64) ([][]float32, error) { + if len(inputIDs) != len(attentionMask) { + return nil, fmt.Errorf("onnxscorer: batch input_ids length %d != attention_mask length %d", len(inputIDs), len(attentionMask)) + } + if len(inputIDs) == 0 { + return nil, nil + } + if len(inputIDs) == 1 { + // No padding needed for a batch of one — reuse Score directly + // rather than duplicating its single-sequence tensor-building path. + scores, err := s.Score(inputIDs[0], attentionMask[0]) + if err != nil { + return nil, err + } + return [][]float32{scores}, nil + } + + maxLen := 0 + for i := range inputIDs { + if len(inputIDs[i]) != len(attentionMask[i]) { + return nil, fmt.Errorf("onnxscorer: batch item %d: input_ids length %d != attention_mask length %d", i, len(inputIDs[i]), len(attentionMask[i])) + } + if len(inputIDs[i]) > maxLen { + maxLen = len(inputIDs[i]) + } + } + if maxLen == 0 { + return make([][]float32, len(inputIDs)), nil + } + + batch := len(inputIDs) + flatIDs := make([]int64, batch*maxLen) + flatMask := make([]int64, batch*maxLen) + for i := range inputIDs { + copy(flatIDs[i*maxLen:], inputIDs[i]) + copy(flatMask[i*maxLen:], attentionMask[i]) + // The remainder of each row stays zero-valued (Go's zero-init) — + // 0 input_ids id, 0 attention_mask — see the doc comment above. + } + + shape := ort.NewShape(int64(batch), int64(maxLen)) + idsTensor, err := ort.NewTensor(shape, flatIDs) + if err != nil { + return nil, fmt.Errorf("onnxscorer: build batch input_ids tensor: %w", err) + } + defer idsTensor.Destroy() + + maskTensor, err := ort.NewTensor(shape, flatMask) + if err != nil { + return nil, fmt.Errorf("onnxscorer: build batch attention_mask tensor: %w", err) + } + defer maskTensor.Destroy() + + outputs := []ort.Value{nil} + if err := s.session.Run([]ort.Value{idsTensor, maskTensor}, outputs); err != nil { + return nil, fmt.Errorf("onnxscorer: batch run: %w", err) + } + out, ok := outputs[0].(*ort.Tensor[float32]) + if !ok { + outputs[0].Destroy() + return nil, fmt.Errorf("onnxscorer: final_scores output is %T, want *Tensor[float32]", outputs[0]) + } + defer out.Destroy() + + data := out.GetData() + if len(data) != batch*maxLen { + return nil, fmt.Errorf("onnxscorer: batch output length %d, want %d (batch=%d, maxLen=%d)", len(data), batch*maxLen, batch, maxLen) + } + + results := make([][]float32, batch) + for i := range inputIDs { + seqLen := len(inputIDs[i]) + scores := make([]float32, seqLen) + copy(scores, data[i*maxLen:i*maxLen+seqLen]) + results[i] = scores + } + return results, nil +} diff --git a/go/internal/pricing/windows.go b/go/internal/pricing/windows.go new file mode 100644 index 0000000..7e946a0 --- /dev/null +++ b/go/internal/pricing/windows.go @@ -0,0 +1,208 @@ +// Package pricing evaluates provider time-of-day pricing windows (e.g. +// DeepSeek's UTC peak/off-peak schedule) against a point in time. It knows +// nothing about money — callers pick a price tier by name ("peak"/ +// "off_peak"/"flat") and apply their own rates. +package pricing + +import ( + "encoding/json" + "fmt" + "sync" + "time" +) + +// TierPeak/TierOffPeak/TierFlat are the tier names Active/TierAt return. +// "flat" means the provider has no configured windows at all — every +// request is priced at the base (off-peak) rate unconditionally. +const ( + TierPeak = "peak" + TierOffPeak = "off_peak" + TierFlat = "flat" +) + +// Window is one recurring time-of-day window (in the owning Windows' +// timezone — see Windows.TZ), active on the listed weekdays between Start +// (inclusive) and End (exclusive) — a half-open [Start, End) range, so a +// request landing exactly on End is NOT in the window. Start/End are +// "HH:MM", 24h; a window may wrap midnight (Start > End), e.g. +// "22:00"-"02:00". +type Window struct { + Days []time.Weekday `json:"days"` + Start string `json:"start"` + End string `json:"end"` +} + +// Windows is a provider's full peak schedule. The zero value has no +// windows, so every time is off-peak-with-no-peak-concept (Flat). +// +// TZ is an IANA zone name ("America/Los_Angeles", "Asia/Tokyo"); "" +// (the default, and the only value before the 2026-09-13 timezone-input +// sprint) means UTC. A provider's published peak hours are almost always +// quoted in one local timezone ("9am-5pm Pacific"), not UTC — requiring the +// operator to hand-convert to UTC was a real usability trap: it goes stale +// twice a year across a DST transition, since a fixed UTC offset entered in +// January is wrong by an hour in July. Evaluating in the configured zone +// via Go's IANA tzdata (time.Time.In) is correct across DST automatically +// — there is no stored UTC range to go stale, because none is stored; +// every evaluation re-derives the real offset for that exact date. +type Windows struct { + TZ string `json:"tz,omitempty"` + Windows []Window `json:"windows"` +} + +// locCache memoizes time.LoadLocation — Active runs on the remote-request +// hot path (router/usage.go's computeCostNative, once per response), and +// the stdlib re-parses zoneinfo from disk/embedded data on every call with +// no cache of its own. IANA zone data is static for the life of the +// process, so caching by name is safe. +var locCache sync.Map // string -> *time.Location + +func loadLocation(name string) (*time.Location, error) { + if name == "" { + return time.UTC, nil + } + if v, ok := locCache.Load(name); ok { + return v.(*time.Location), nil + } + loc, err := time.LoadLocation(name) + if err != nil { + return nil, fmt.Errorf("pricing: unknown timezone %q: %w", name, err) + } + locCache.Store(name, loc) + return loc, nil +} + +// minutesOfDay parses "HH:MM" into minutes since midnight. +func minutesOfDay(s string) (int, error) { + var h, m int + if _, err := fmt.Sscanf(s, "%d:%d", &h, &m); err != nil { + return 0, fmt.Errorf("pricing: invalid time %q: %w", s, err) + } + if h < 0 || h > 23 || m < 0 || m > 59 { + return 0, fmt.Errorf("pricing: time %q out of range", s) + } + return h*60 + m, nil +} + +// Parse decodes a Windows schedule from its JSON-encoded form (as stored in +// router_providers.peak_windows). An empty string is valid and yields the +// zero value (no windows configured, UTC). TZ (if present) is validated +// against the real IANA database — an unrecognized name (a typo, or an +// abbreviation like "PST" that IANA deliberately doesn't resolve — always +// use the full zone name, e.g. "America/Los_Angeles") is rejected here +// rather than silently degrading to UTC at evaluation time. Every window is +// validated: weekdays in 0-6, well-formed "HH:MM" times, and Start != End +// (a zero-width window can never be intentional and would otherwise +// silently match nothing). +func Parse(raw string) (Windows, error) { + var w Windows + if raw == "" { + return w, nil + } + if err := json.Unmarshal([]byte(raw), &w); err != nil { + return Windows{}, fmt.Errorf("pricing: parse windows: %w", err) + } + if _, err := loadLocation(w.TZ); err != nil { + return Windows{}, err + } + for i, win := range w.Windows { + for _, d := range win.Days { + if d < time.Sunday || d > time.Saturday { + return Windows{}, fmt.Errorf("pricing: window %d: day %d out of range 0-6", i, d) + } + } + if len(win.Days) == 0 { + return Windows{}, fmt.Errorf("pricing: window %d: no days set", i) + } + start, err := minutesOfDay(win.Start) + if err != nil { + return Windows{}, fmt.Errorf("pricing: window %d: start: %w", i, err) + } + end, err := minutesOfDay(win.End) + if err != nil { + return Windows{}, fmt.Errorf("pricing: window %d: end: %w", i, err) + } + if start == end { + return Windows{}, fmt.Errorf("pricing: window %d: start == end (%s), zero-width window", i, win.Start) + } + } + return w, nil +} + +// dayIn reports whether d appears in days. +func dayIn(days []time.Weekday, d time.Weekday) bool { + for _, x := range days { + if x == d { + return true + } + } + return false +} + +// activeOn reports whether lt — already converted into the schedule's +// configured zone by the caller (Active) — falls inside win. A window that +// wraps midnight (Start > End) is active either on the day it starts (from +// Start through 24:00) or on the following day (from 00:00 up to End), so +// the weekday check is applied to whichever side of the wrap lt's own local +// weekday matches. +func (win Window) activeOn(lt time.Time) bool { + start, err := minutesOfDay(win.Start) + if err != nil { + return false + } + end, err := minutesOfDay(win.End) + if err != nil { + return false + } + nowMin := lt.Hour()*60 + lt.Minute() + weekday := lt.Weekday() + + if start < end { + return dayIn(win.Days, weekday) && nowMin >= start && nowMin < end + } + // Wraps midnight: the "start day" segment runs [start, 24:00) on the + // listed weekday; the "end day" segment runs [0:00, end) on the + // following weekday. + if dayIn(win.Days, weekday) && nowMin >= start { + return true + } + prevDay := (weekday + 6) % 7 // yesterday, wrapping Sunday->Saturday + return dayIn(win.Days, prevDay) && nowMin < end +} + +// Active reports whether t falls inside any configured window. t is +// converted into the schedule's configured zone (w.TZ, "" = UTC) before +// any day-of-week or time-of-day comparison — a real IANA conversion via +// Go's tzdata, not a fixed offset, so it stays correct across a DST +// transition and correctly shifts which calendar day/hour it is in zones +// far from UTC (e.g. a window quoted "Monday 00:00-04:00 Asia/Tokyo" is +// evaluated as Monday in Tokyo, which is still Sunday afternoon UTC). +func (w Windows) Active(t time.Time) bool { + if len(w.Windows) == 0 { + return false + } + loc, err := loadLocation(w.TZ) + if err != nil { + return false // invalid zone (shouldn't happen post-Parse) — degrade safely, never panic + } + lt := t.In(loc) + for _, win := range w.Windows { + if win.activeOn(lt) { + return true + } + } + return false +} + +// TierAt returns TierFlat when w has no windows configured at all (the +// provider has no peak/off-peak concept), otherwise TierPeak or +// TierOffPeak depending on whether t falls inside a window. +func (w Windows) TierAt(t time.Time) string { + if len(w.Windows) == 0 { + return TierFlat + } + if w.Active(t) { + return TierPeak + } + return TierOffPeak +} diff --git a/go/internal/pricing/windows_test.go b/go/internal/pricing/windows_test.go new file mode 100644 index 0000000..625f9c6 --- /dev/null +++ b/go/internal/pricing/windows_test.go @@ -0,0 +1,206 @@ +package pricing + +import ( + "testing" + "time" +) + +// deepseekSchedule mirrors the real published DeepSeek peak schedule: +// 01:00-04:00 and 06:00-10:00 UTC, Monday-Friday. +const deepseekSchedule = `{"windows":[ + {"days":[1,2,3,4,5],"start":"01:00","end":"04:00"}, + {"days":[1,2,3,4,5],"start":"06:00","end":"10:00"} +]}` + +func mustParse(t *testing.T, raw string) Windows { + t.Helper() + w, err := Parse(raw) + if err != nil { + t.Fatalf("Parse(%q): %v", raw, err) + } + return w +} + +func utc(y, m, d, h, min int) time.Time { + return time.Date(y, time.Month(m), d, h, min, 0, 0, time.UTC) +} + +func TestActive_DeepSeekSchedule(t *testing.T) { + w := mustParse(t, deepseekSchedule) + + cases := []struct { + name string + t time.Time + want bool + }{ + {"Tue 02:00 in first window", utc(2026, 9, 15, 2, 0), true}, + {"Tue 04:00 exact end is off-peak (half-open)", utc(2026, 9, 15, 4, 0), false}, + {"Tue 03:59 still in window", utc(2026, 9, 15, 3, 59), true}, + {"Tue 01:00 exact start is peak (half-open)", utc(2026, 9, 15, 1, 0), true}, + {"Tue 05:00 gap between windows is off-peak", utc(2026, 9, 15, 5, 0), false}, + {"Tue 06:00 in second window", utc(2026, 9, 15, 6, 0), true}, + {"Tue 09:59 still in second window", utc(2026, 9, 15, 9, 59), true}, + {"Tue 10:00 exact end is off-peak", utc(2026, 9, 15, 10, 0), false}, + {"Tue 12:00 off-peak", utc(2026, 9, 15, 12, 0), false}, + {"Sat 02:00 weekday excluded", utc(2026, 9, 19, 2, 0), false}, + {"Sun 07:00 weekday excluded", utc(2026, 9, 20, 7, 0), false}, + } + for _, c := range cases { + t.Run(c.name, func(t *testing.T) { + got := w.Active(c.t) + if got != c.want { + t.Errorf("Active(%s) = %v, want %v", c.t.Format(time.RFC3339), got, c.want) + } + }) + } +} + +func TestActive_UTCConversion(t *testing.T) { + w := mustParse(t, deepseekSchedule) + // 02:00 UTC on a Tuesday, expressed in a non-UTC location, must still + // evaluate as UTC (Active must not use the input's own wall-clock + // fields). + loc := time.FixedZone("UTC-5", -5*60*60) + inPST := time.Date(2026, 9, 14, 21, 0, 0, 0, loc) // == Tue 2026-09-15 02:00 UTC + if !w.Active(inPST) { + t.Errorf("Active did not convert non-UTC input to UTC before evaluating") + } +} + +func TestActive_MidnightWrap(t *testing.T) { + // A window from 22:00 Mon through 02:00 (next day) Mon-listed. + w := mustParse(t, `{"windows":[{"days":[1],"start":"22:00","end":"02:00"}]}`) + + cases := []struct { + name string + t time.Time + want bool + }{ + {"Mon 23:00 in start-day segment", utc(2026, 9, 14, 23, 0), true}, + {"Tue 01:00 in end-day segment (wrap)", utc(2026, 9, 15, 1, 0), true}, + {"Tue 02:00 exact end is off (half-open)", utc(2026, 9, 15, 2, 0), false}, + {"Mon 21:59 before start", utc(2026, 9, 14, 21, 59), false}, + {"Wed 01:00 wrong day entirely", utc(2026, 9, 16, 1, 0), false}, + } + for _, c := range cases { + t.Run(c.name, func(t *testing.T) { + got := w.Active(c.t) + if got != c.want { + t.Errorf("Active(%s) = %v, want %v", c.t.Format(time.RFC3339), got, c.want) + } + }) + } +} + +func TestTierAt(t *testing.T) { + w := mustParse(t, deepseekSchedule) + if got := w.TierAt(utc(2026, 9, 15, 2, 0)); got != TierPeak { + t.Errorf("TierAt(peak time) = %q, want %q", got, TierPeak) + } + if got := w.TierAt(utc(2026, 9, 15, 12, 0)); got != TierOffPeak { + t.Errorf("TierAt(off-peak time) = %q, want %q", got, TierOffPeak) + } + + flat := Windows{} + if got := flat.TierAt(utc(2026, 9, 15, 2, 0)); got != TierFlat { + t.Errorf("TierAt(no windows) = %q, want %q", got, TierFlat) + } +} + +func TestParse_Empty(t *testing.T) { + w, err := Parse("") + if err != nil { + t.Fatalf("Parse(\"\") error: %v", err) + } + if len(w.Windows) != 0 { + t.Errorf("Parse(\"\") = %+v, want zero value", w) + } +} + +func TestParse_Rejects(t *testing.T) { + cases := []struct { + name string + raw string + }{ + {"malformed json", `{not json`}, + {"start equals end", `{"windows":[{"days":[1],"start":"04:00","end":"04:00"}]}`}, + {"day out of range", `{"windows":[{"days":[7],"start":"01:00","end":"02:00"}]}`}, + {"negative day", `{"windows":[{"days":[-1],"start":"01:00","end":"02:00"}]}`}, + {"no days", `{"windows":[{"days":[],"start":"01:00","end":"02:00"}]}`}, + {"bad start format", `{"windows":[{"days":[1],"start":"1am","end":"02:00"}]}`}, + {"hour out of range", `{"windows":[{"days":[1],"start":"25:00","end":"02:00"}]}`}, + {"minute out of range", `{"windows":[{"days":[1],"start":"01:60","end":"02:00"}]}`}, + {"unknown timezone", `{"tz":"Not/AZone","windows":[{"days":[1],"start":"01:00","end":"02:00"}]}`}, + {"tz abbreviation not a real IANA name", `{"tz":"PST","windows":[{"days":[1],"start":"01:00","end":"02:00"}]}`}, + } + for _, c := range cases { + t.Run(c.name, func(t *testing.T) { + if _, err := Parse(c.raw); err == nil { + t.Errorf("Parse(%q) succeeded, want error", c.raw) + } + }) + } +} + +// ── Timezone support (2026-09-13) ─────────────────────────────────────────── + +// TestActive_TimezoneDST proves the schedule is evaluated in the +// configured zone's REAL local wall-clock time — via Go's IANA tzdata, not +// a fixed offset frozen at entry time — so it stays correct across a DST +// transition with no stored UTC range to go stale. A "09:00-17:00 +// America/Los_Angeles" window is active at 16:30 UTC in July (PDT, UTC-7 -> +// 09:30 local) but NOT at the same 16:30 UTC clock time in January (PST, +// UTC-8 -> 08:30 local, before the window opens). +func TestActive_TimezoneDST(t *testing.T) { + w := mustParse(t, `{"tz":"America/Los_Angeles","windows":[{"days":[1,2,3,4,5],"start":"09:00","end":"17:00"}]}`) + + // 2026-01-15 is a Thursday; 2026-07-15 is a Wednesday. Both weekdays. + winterInactive := utc(2026, 1, 15, 16, 30) // 08:30 PST -- before the window + winterActive := utc(2026, 1, 15, 17, 30) // 09:30 PST -- inside the window + summerActive := utc(2026, 7, 15, 16, 30) // 09:30 PDT -- inside the window (DST) + + if w.Active(winterInactive) { + t.Error("Active(16:30 UTC in January / 08:30 PST) = true, want false") + } + if !w.Active(winterActive) { + t.Error("Active(17:30 UTC in January / 09:30 PST) = false, want true") + } + if !w.Active(summerActive) { + t.Error("Active(16:30 UTC in July / 09:30 PDT, DST) = false, want true — same UTC clock time as the January-inactive case, but DST shifts the local hour into the window") + } +} + +// TestActive_TimezoneShiftsWeekday proves the WEEKDAY the schedule matches +// against is the configured zone's local day, not UTC's — a window quoted +// "Monday 00:00-04:00 Asia/Tokyo" (UTC+9, no DST) is active during Sunday +// afternoon UTC, because it's already Monday in Tokyo. +func TestActive_TimezoneShiftsWeekday(t *testing.T) { + w := mustParse(t, `{"tz":"Asia/Tokyo","windows":[{"days":[1],"start":"00:00","end":"04:00"}]}`) + + // 2026-01-11 is a Sunday (UTC). + sundayNoon := utc(2026, 1, 11, 12, 0) // 21:00 JST Sunday -- still Sunday locally + sundayAfternoon := utc(2026, 1, 11, 16, 0) // 01:00 JST Monday -- already Monday locally, inside the window + + if w.Active(sundayNoon) { + t.Error("Active(12:00 UTC Sun / 21:00 JST Sun) = true, want false (still Sunday in Tokyo)") + } + if !w.Active(sundayAfternoon) { + t.Error("Active(16:00 UTC Sun / 01:00 JST Mon) = false, want true (already Monday in Tokyo, inside the window)") + } +} + +// TestActive_EmptyTZDefaultsUTC confirms an unset tz behaves exactly as +// before the timezone-input sprint (the DeepSeek schedule stored live has +// no "tz" key at all). +func TestActive_EmptyTZDefaultsUTC(t *testing.T) { + withTZ := mustParse(t, `{"tz":"UTC","windows":[{"days":[1,2,3,4,5],"start":"01:00","end":"04:00"}]}`) + withoutTZ := mustParse(t, deepseekSchedule) + tPeak := utc(2026, 9, 15, 2, 0) + tOff := utc(2026, 9, 15, 12, 0) + if withTZ.Active(tPeak) != withoutTZ.Active(tPeak) { + t.Error("explicit tz:UTC and omitted tz disagree at a peak time") + } + if withTZ.Active(tOff) != withoutTZ.Active(tOff) { + t.Error("explicit tz:UTC and omitted tz disagree at an off-peak time") + } +} diff --git a/go/internal/providers/providers.go b/go/internal/providers/providers.go index bc961b8..a16f839 100644 --- a/go/internal/providers/providers.go +++ b/go/internal/providers/providers.go @@ -28,6 +28,7 @@ import ( "sync" "time" + "github.com/jsaigou/the-forge/internal/pricing" "github.com/jsaigou/the-forge/internal/store" ) @@ -83,6 +84,13 @@ type Provider struct { // from router_providers (0008 columns). "" when unknown. Country string DataResidencyGroup string + + // PeakWindows/PeakActiveNow (peak pricing sprint, 2026-09-12): raw + // schedule plus a server-evaluated "is peak in force right now" flag, + // computed once here against s.deps.now() rather than left for the FE + // to re-derive from the raw schedule. + PeakWindows string + PeakActiveNow bool } // Model is one row under a provider — Phase 7 (2026-08-13): sourced from @@ -340,6 +348,7 @@ func (s *service) List(ctx context.Context) ([]Provider, error) { if models == nil { models = []Model{} } + windows, _ := pricing.Parse(row.PeakWindows) // malformed row -> zero value, PeakActiveNow false out = append(out, Provider{ ID: row.ID, Name: row.Name, @@ -357,6 +366,8 @@ func (s *service) List(ctx context.Context) ([]Provider, error) { Enabled: row.Enabled, Country: row.Country, DataResidencyGroup: row.DataResidencyGroup, + PeakWindows: row.PeakWindows, + PeakActiveNow: windows.Active(s.deps.now()), }) } return out, nil diff --git a/go/internal/providers/providers_test.go b/go/internal/providers/providers_test.go index 5346adb..715e3de 100644 --- a/go/internal/providers/providers_test.go +++ b/go/internal/providers/providers_test.go @@ -113,6 +113,52 @@ func TestListNoProvidersReturnsEmpty(t *testing.T) { } } +// TestListPeakWindows covers the peak pricing sprint's (2026-09-12) +// PeakWindows/PeakActiveNow fields: the raw schedule passes through +// unchanged, PeakActiveNow is computed against the injected clock (not wall +// time), and a provider with no schedule reports false rather than erroring. +func TestListPeakWindows(t *testing.T) { + const windows = `{"windows":[{"days":[1,2,3,4,5],"start":"01:00","end":"04:00"}]}` + cat := newFakeCatalog([]store.ProviderRow{ + {Name: "deepseek", APIKey: "sk-d", Enabled: true, PeakWindows: windows}, + {Name: "aiand", APIKey: "sk-a", Enabled: true}, + }) + // Tuesday 02:00 UTC — inside deepseek's window. + inWindow := time.Date(2026, 9, 15, 2, 0, 0, 0, time.UTC) + svc := New(Deps{Catalog: cat, now: func() time.Time { return inWindow }}) + + out, err := svc.List(context.Background()) + if err != nil { + t.Fatalf("List: %v", err) + } + byName := map[string]Provider{} + for _, p := range out { + byName[p.Name] = p + } + if byName["deepseek"].PeakWindows != windows { + t.Errorf("deepseek PeakWindows = %q, want %q", byName["deepseek"].PeakWindows, windows) + } + if !byName["deepseek"].PeakActiveNow { + t.Error("deepseek PeakActiveNow = false, want true (injected clock is inside the window)") + } + if byName["aiand"].PeakActiveNow { + t.Error("aiand PeakActiveNow = true, want false (no schedule configured)") + } + + // Same schedule, clock moved outside the window. + outsideWindow := time.Date(2026, 9, 15, 12, 0, 0, 0, time.UTC) + svc2 := New(Deps{Catalog: cat, now: func() time.Time { return outsideWindow }}) + out2, err := svc2.List(context.Background()) + if err != nil { + t.Fatalf("List (outside window): %v", err) + } + for _, p := range out2 { + if p.Name == "deepseek" && p.PeakActiveNow { + t.Error("deepseek PeakActiveNow = true at noon, want false") + } + } +} + // ── Test: DeepSeek + AI& with live probes ─────────────────────────────────── // // Mirrors ForgeHost's real config (health live-verified 2026-07-22): DeepSeek diff --git a/go/internal/router/config.go b/go/internal/router/config.go index 071367d..85233f9 100644 --- a/go/internal/router/config.go +++ b/go/internal/router/config.go @@ -10,6 +10,7 @@ import ( "net/url" "time" + "github.com/jsaigou/the-forge/internal/pricing" "github.com/jsaigou/the-forge/internal/store" ) @@ -80,6 +81,19 @@ type Backend struct { PriceOutPer1M float64 PriceCachedInPer1M *float64 // nil = provider's cache-hit discount unmodelled PriceCurrency string + + // Peak pricing (2026-09-12): PriceInPer1MPeak/PriceOutPer1MPeak/ + // PriceCachedInPer1MPeak carry the offering's peak-tier rates (nil per + // field = no peak differential, falls back to the base rate above) and + // PeakWindows is the owning provider's parsed schedule — both copied + // alongside the base prices at offeringChain's single copy point, so + // computeCostNative can pick a tier per-request with no extra DB read. + // PeakWindows.Windows is empty for a foundry_slot backend and for any + // remote provider with no configured schedule. + PriceInPer1MPeak *float64 + PriceOutPer1MPeak *float64 + PriceCachedInPer1MPeak *float64 + PeakWindows pricing.Windows } // Route is one [[router.routes]] entry: a logical model name maps to an diff --git a/go/internal/router/routing.go b/go/internal/router/routing.go index 83bd5d3..7908354 100644 --- a/go/internal/router/routing.go +++ b/go/internal/router/routing.go @@ -7,6 +7,7 @@ import ( "encoding/json" "errors" "fmt" + "log" "net/http" "strconv" "strings" @@ -14,6 +15,7 @@ import ( "github.com/jsaigou/the-forge/internal/activity" "github.com/jsaigou/the-forge/internal/authz" + "github.com/jsaigou/the-forge/internal/pricing" "github.com/jsaigou/the-forge/internal/sched" "github.com/jsaigou/the-forge/internal/store" ) @@ -43,6 +45,14 @@ type ResolvedBackend struct { PriceOutPer1M float64 PriceCachedInPer1M *float64 PriceCurrency string + // Peak pricing (2026-09-12) — see Backend's matching fields. Carried + // through resolveBackend's remote case alongside the base prices; + // computeCostNative picks base vs. peak per field, per request, using + // the response-completion timestamp (recordExternalUsage's now). + PriceInPer1MPeak *float64 + PriceOutPer1MPeak *float64 + PriceCachedInPer1MPeak *float64 + PeakWindows pricing.Windows // UpstreamOverride, when non-empty, is sent as the x-compress-base-url // request header (docs/v5-headroom-topology.md §3/§4): the shared local // Compressor proxy (BaseURL) honors it to route this specific request to @@ -221,19 +231,23 @@ func (s *Server) resolveBackend(ctx context.Context, b *Backend) (ResolvedBacken directUpstreamURL = provider.TargetURL } return ResolvedBackend{ - Name: b.Name, - BaseURL: baseURL, - APIKey: provider.APIKey, - WireModel: b.WireModel, - Provider: provider.Name, - ProviderID: provider.ID, - PriceInPer1M: b.PriceInPer1M, - PriceOutPer1M: b.PriceOutPer1M, - PriceCachedInPer1M: b.PriceCachedInPer1M, - PriceCurrency: b.PriceCurrency, - UpstreamOverride: upstreamOverride, - CompressorFronted: compressorFronted, - DirectUpstreamURL: directUpstreamURL, + Name: b.Name, + BaseURL: baseURL, + APIKey: provider.APIKey, + WireModel: b.WireModel, + Provider: provider.Name, + ProviderID: provider.ID, + PriceInPer1M: b.PriceInPer1M, + PriceOutPer1M: b.PriceOutPer1M, + PriceCachedInPer1M: b.PriceCachedInPer1M, + PriceCurrency: b.PriceCurrency, + PriceInPer1MPeak: b.PriceInPer1MPeak, + PriceOutPer1MPeak: b.PriceOutPer1MPeak, + PriceCachedInPer1MPeak: b.PriceCachedInPer1MPeak, + PeakWindows: b.PeakWindows, + UpstreamOverride: upstreamOverride, + CompressorFronted: compressorFronted, + DirectUpstreamURL: directUpstreamURL, }, nil default: @@ -674,11 +688,23 @@ func (s *Server) offeringChain(ctx context.Context, model string) (chain []*Back if !hasProxy { baseURL = provider.TargetURL // no Compressor proxy (dedicated or shared external) — straight passthrough } + // Peak pricing (2026-09-12): parse once here, at the single copy + // point, rather than per-request in computeCostNative. A malformed + // schedule (should only happen to a row written before API + // validation existed) degrades to "no windows" — never blocks + // routing over a pricing concern. + windows, err := pricing.Parse(provider.PeakWindows) + if err != nil { + log.Printf("router: offering_chain: provider %q: invalid peak_windows, treating as none: %v", provider.Name, err) + windows = pricing.Windows{} + } chain = append(chain, &Backend{ Name: o.ProviderName, Kind: "remote", BaseURL: baseURL, WireModel: o.WireModel, Credential: o.ProviderName, PriceInPer1M: o.PriceInPer1M, PriceOutPer1M: o.PriceOutPer1M, PriceCachedInPer1M: o.PriceCachedInPer1M, PriceCurrency: o.Currency, + PriceInPer1MPeak: o.PriceInPer1MPeak, PriceOutPer1MPeak: o.PriceOutPer1MPeak, + PriceCachedInPer1MPeak: o.PriceCachedInPer1MPeak, PeakWindows: windows, }) } if len(chain) == 0 { diff --git a/go/internal/router/usage.go b/go/internal/router/usage.go index e606516..647d6bf 100644 --- a/go/internal/router/usage.go +++ b/go/internal/router/usage.go @@ -22,6 +22,7 @@ import ( "log" "time" + "github.com/jsaigou/the-forge/internal/pricing" "github.com/jsaigou/the-forge/internal/store" ) @@ -210,26 +211,54 @@ func parseStreamingUsage(tail []byte) (prompt, completion, cached int64, ok bool } // computeCostNative prices promptTokens/completionTokens/cachedTokens -// against resolved's per-1M rates. ok=false ("" PriceCurrency) means no -// offering matched this request — the caller must not fabricate a cost. -// When the provider discounts cache hits but PriceCachedInPer1M is nil -// (unmodelled), cached tokens are priced at the full input rate — a -// documented upper bound, never an under-estimate. -func computeCostNative(resolved ResolvedBackend, promptTokens, completionTokens, cachedTokens int64) (cost float64, ok bool) { +// against resolved's per-1M rates, picking the price tier in force at t +// (the response-completion time — see recordExternalUsage). ok=false ("" +// PriceCurrency) means no offering matched this request — the caller must +// not fabricate a cost. When the provider discounts cache hits but the +// chosen tier's cached rate is nil (unmodelled), cached tokens are priced +// at the full input rate — a documented upper bound, never an under-estimate. +// +// Tier resolution is deliberately at record time, not request-start time: +// nothing upstream of this function ever uses price for routing decisions +// (select.go sorts by priority only), and pinning the tier here means the +// stored cost and the tier derivable from the stored event timestamp can +// never disagree — the one property compressor_summary_handlers.go's +// historical re-pricing (estimateRemoteCacheDiscountSaved / +// estimateRemoteCompressionSaved) depends on. A request whose response +// completes just after a tier boundary is therefore billed entirely at the +// new tier, even though most of the work happened in the old one — a +// reproducible, auditable rule, recorded on the event as tier, rather than +// a guess at DeepSeek's own internal accounting for boundary-crossers. +func computeCostNative(resolved ResolvedBackend, promptTokens, completionTokens, cachedTokens int64, t time.Time) (cost float64, tier string, ok bool) { if resolved.PriceCurrency == "" { - return 0, false + return 0, "", false } + tier = resolved.PeakWindows.TierAt(t) + + priceIn, priceOut, priceCachedIn := resolved.PriceInPer1M, resolved.PriceOutPer1M, resolved.PriceCachedInPer1M + if tier == pricing.TierPeak { + if resolved.PriceInPer1MPeak != nil { + priceIn = *resolved.PriceInPer1MPeak + } + if resolved.PriceOutPer1MPeak != nil { + priceOut = *resolved.PriceOutPer1MPeak + } + if resolved.PriceCachedInPer1MPeak != nil { + priceCachedIn = resolved.PriceCachedInPer1MPeak + } + } + billableIn := promptTokens - if cachedTokens > 0 && resolved.PriceCachedInPer1M != nil { + if cachedTokens > 0 && priceCachedIn != nil { billableIn = promptTokens - cachedTokens if billableIn < 0 { billableIn = 0 } - cost += float64(cachedTokens) / 1e6 * *resolved.PriceCachedInPer1M + cost += float64(cachedTokens) / 1e6 * *priceCachedIn } - cost += float64(billableIn) / 1e6 * resolved.PriceInPer1M - cost += float64(completionTokens) / 1e6 * resolved.PriceOutPer1M - return cost, true + cost += float64(billableIn) / 1e6 * priceIn + cost += float64(completionTokens) / 1e6 * priceOut + return cost, tier, true } // recordExternalUsage builds and persists one kind="external_request" @@ -258,8 +287,9 @@ func (s *Server) recordExternalUsage(resolved ResolvedBackend, model string, buf prompt, completion, cached, ok = parseNonStreamingUsage(buf) } + now := time.Now() ev := store.UsageEvent{ - TS: time.Now(), Kind: "external_request", + TS: now, Kind: "external_request", Model: resolved.WireModel, ProviderID: nonZeroInt64Ptr(resolved.ProviderID), } if !ok { @@ -272,9 +302,13 @@ func (s *Server) recordExternalUsage(resolved ResolvedBackend, model string, buf if cached > 0 { ev.CachedPromptTokens = &cached } - if cost, ok := computeCostNative(resolved, prompt, completion, cached); ok { + // now (not a separately-captured time) both stamps ev.TS above and + // selects the price tier below, by construction — see + // computeCostNative's doc comment for why that must be one value. + if cost, tier, ok := computeCostNative(resolved, prompt, completion, cached, now); ok { ev.CostNative = &cost ev.CostCurrency = resolved.PriceCurrency + ev.PriceTier = tier } } diff --git a/go/internal/router/usage_test.go b/go/internal/router/usage_test.go index 1ffa040..78cff70 100644 --- a/go/internal/router/usage_test.go +++ b/go/internal/router/usage_test.go @@ -13,6 +13,7 @@ import ( "testing" "time" + "github.com/jsaigou/the-forge/internal/pricing" "github.com/jsaigou/the-forge/internal/store" ) @@ -133,8 +134,13 @@ func TestParseStreamingUsagePicksLastUsageFrame(t *testing.T) { // ── computeCostNative ────────────────────────────────────────────────────── +// offPeakT is an arbitrary time with no configured windows in play for the +// pre-existing (pre-peak-pricing) test cases below — any time works since +// those ResolvedBackends carry a zero-value PeakWindows (TierAt -> "flat"). +var offPeakT = time.Date(2026, 9, 15, 12, 0, 0, 0, time.UTC) + func TestComputeCostNativeNoOffering(t *testing.T) { - _, ok := computeCostNative(ResolvedBackend{}, 1000, 500, 0) + _, _, ok := computeCostNative(ResolvedBackend{}, 1000, 500, 0, offPeakT) if ok { t.Error("ok = true with no PriceCurrency (no offering matched)") } @@ -142,7 +148,7 @@ func TestComputeCostNativeNoOffering(t *testing.T) { func TestComputeCostNativeBasic(t *testing.T) { rb := ResolvedBackend{PriceInPer1M: 1.0, PriceOutPer1M: 2.0, PriceCurrency: "USD"} - cost, ok := computeCostNative(rb, 1_000_000, 500_000, 0) + cost, tier, ok := computeCostNative(rb, 1_000_000, 500_000, 0, offPeakT) if !ok { t.Fatal("ok = false, want true") } @@ -150,11 +156,14 @@ func TestComputeCostNativeBasic(t *testing.T) { if cost != want { t.Errorf("cost = %v, want %v", cost, want) } + if tier != pricing.TierFlat { + t.Errorf("tier = %q, want %q (no windows configured)", tier, pricing.TierFlat) + } } func TestComputeCostNativeCachedTokensUpperBoundWhenUnmodelled(t *testing.T) { rb := ResolvedBackend{PriceInPer1M: 1.0, PriceOutPer1M: 2.0, PriceCurrency: "USD"} - cost, ok := computeCostNative(rb, 1_000_000, 0, 500_000) // half the input was cached + cost, _, ok := computeCostNative(rb, 1_000_000, 0, 500_000, offPeakT) // half the input was cached if !ok { t.Fatal("ok = false") } @@ -169,7 +178,7 @@ func TestComputeCostNativeCachedTokensUpperBoundWhenUnmodelled(t *testing.T) { func TestComputeCostNativeCachedTokensDiscounted(t *testing.T) { cachedRate := 0.1 rb := ResolvedBackend{PriceInPer1M: 1.0, PriceOutPer1M: 2.0, PriceCurrency: "USD", PriceCachedInPer1M: &cachedRate} - cost, ok := computeCostNative(rb, 1_000_000, 0, 500_000) + cost, _, ok := computeCostNative(rb, 1_000_000, 0, 500_000, offPeakT) if !ok { t.Fatal("ok = false") } @@ -180,6 +189,124 @@ func TestComputeCostNativeCachedTokensDiscounted(t *testing.T) { } } +// ── computeCostNative: peak pricing ───────────────────────────────────────── + +func deepSeekLikeWindows(t *testing.T) pricing.Windows { + t.Helper() + w, err := pricing.Parse(`{"windows":[ + {"days":[1,2,3,4,5],"start":"01:00","end":"04:00"}, + {"days":[1,2,3,4,5],"start":"06:00","end":"10:00"} + ]}`) + if err != nil { + t.Fatalf("pricing.Parse: %v", err) + } + return w +} + +func TestComputeCostNativePeakTier(t *testing.T) { + inPeak, outPeak, cachedPeak := 0.3, 1.2, 0.006 + rb := ResolvedBackend{ + PriceInPer1M: 0.15, PriceOutPer1M: 0.6, PriceCurrency: "USD", + PriceCachedInPer1M: floatPtr(0.003), + PriceInPer1MPeak: &inPeak, + PriceOutPer1MPeak: &outPeak, + PriceCachedInPer1MPeak: &cachedPeak, + PeakWindows: deepSeekLikeWindows(t), + } + peakT := time.Date(2026, 9, 15, 2, 0, 0, 0, time.UTC) // Tue 02:00 UTC — in the first window + cost, tier, ok := computeCostNative(rb, 1_000_000, 500_000, 500_000, peakT) + if !ok { + t.Fatal("ok = false") + } + if tier != pricing.TierPeak { + t.Errorf("tier = %q, want %q", tier, pricing.TierPeak) + } + // 500k billable @ $0.3/1M + 500k cached @ $0.006/1M + 500k out @ $1.2/1M + want := 0.15 + 0.003 + 0.6 + if cost != want { + t.Errorf("cost = %v, want %v", cost, want) + } +} + +func TestComputeCostNativeOffPeakTier(t *testing.T) { + inPeak := 0.3 + rb := ResolvedBackend{ + PriceInPer1M: 0.15, PriceOutPer1M: 0.6, PriceCurrency: "USD", + PriceInPer1MPeak: &inPeak, + PeakWindows: deepSeekLikeWindows(t), + } + offT := time.Date(2026, 9, 15, 12, 0, 0, 0, time.UTC) // Tue noon — outside both windows + cost, tier, ok := computeCostNative(rb, 1_000_000, 0, 0, offT) + if !ok { + t.Fatal("ok = false") + } + if tier != pricing.TierOffPeak { + t.Errorf("tier = %q, want %q", tier, pricing.TierOffPeak) + } + if cost != 0.15 { + t.Errorf("cost = %v, want 0.15 (base rate)", cost) + } +} + +func TestComputeCostNativeExactBoundary(t *testing.T) { + inPeak := 0.3 + rb := ResolvedBackend{ + PriceInPer1M: 0.15, PriceInPer1MPeak: &inPeak, + PriceCurrency: "USD", PeakWindows: deepSeekLikeWindows(t), + } + // Exactly 04:00:00 UTC is the window's exclusive end -> off-peak + // (half-open [start, end), pinned by internal/pricing's own tests). + boundary := time.Date(2026, 9, 15, 4, 0, 0, 0, time.UTC) + cost, tier, ok := computeCostNative(rb, 1_000_000, 0, 0, boundary) + if !ok { + t.Fatal("ok = false") + } + if tier != pricing.TierOffPeak { + t.Errorf("tier at exact boundary = %q, want %q", tier, pricing.TierOffPeak) + } + if cost != 0.15 { + t.Errorf("cost at exact boundary = %v, want 0.15 (base rate)", cost) + } +} + +func TestComputeCostNativePeakWindowsButNoPeakPricesFallsBackToBase(t *testing.T) { + // Windows configured (so the tier genuinely is "peak" right now) but no + // peak price columns set -> tier is still reported honestly as "peak" + // (the objective fact about the clock), while cost reflects base rates + // (per-field nil fallback). + rb := ResolvedBackend{ + PriceInPer1M: 0.15, PriceOutPer1M: 0.6, PriceCurrency: "USD", + PeakWindows: deepSeekLikeWindows(t), + } + peakT := time.Date(2026, 9, 15, 2, 0, 0, 0, time.UTC) + cost, tier, ok := computeCostNative(rb, 1_000_000, 0, 0, peakT) + if !ok { + t.Fatal("ok = false") + } + if tier != pricing.TierPeak { + t.Errorf("tier = %q, want %q", tier, pricing.TierPeak) + } + if cost != 0.15 { + t.Errorf("cost = %v, want 0.15 (base rate, no peak price modelled)", cost) + } +} + +func TestComputeCostNativeNoWindowsIsFlat(t *testing.T) { + rb := ResolvedBackend{PriceInPer1M: 1.0, PriceCurrency: "USD"} + cost, tier, ok := computeCostNative(rb, 1_000_000, 0, 0, time.Date(2026, 9, 15, 2, 0, 0, 0, time.UTC)) + if !ok { + t.Fatal("ok = false") + } + if tier != pricing.TierFlat { + t.Errorf("tier = %q, want %q", tier, pricing.TierFlat) + } + if cost != 1.0 { + t.Errorf("cost = %v, want 1.0", cost) + } +} + +func floatPtr(f float64) *float64 { return &f } + // ── usageTap ──────────────────────────────────────────────────────────────── func TestUsageTapNonStreamingPassesThroughAndCapturesAll(t *testing.T) { diff --git a/go/internal/smith/tools.go b/go/internal/smith/tools.go index de9480c..27fb4ef 100644 --- a/go/internal/smith/tools.go +++ b/go/internal/smith/tools.go @@ -543,6 +543,13 @@ type toolOfferingView struct { PriceInPer1M float64 `json:"price_in_per_1m"` PriceOutPer1M float64 `json:"price_out_per_1m"` Currency string `json:"currency"` + // PriceNote (peak pricing sprint, 2026-09-12) discloses that + // price_in_per_1m/price_out_per_1m above are the OFF-PEAK/base rate + // only, when the offering has any peak-tier price configured — without + // this the tool would silently state half the truth for a + // time-of-day-priced provider (e.g. DeepSeek). Empty when the offering + // has no peak pricing at all. + PriceNote string `json:"price_note,omitempty"` } func catalogLookupTool(ctx context.Context, env *ToolEnv, args json.RawMessage) (any, error) { @@ -587,10 +594,14 @@ func catalogLookupTool(ctx context.Context, env *ToolEnv, args json.RawMessage) if a.Name != "" && o.WireModel != a.Name { continue } - out = append(out, toolOfferingView{ + view := toolOfferingView{ Provider: o.ProviderName, WireModel: o.WireModel, ContextLength: o.ContextLength, Enabled: o.Enabled, PriceInPer1M: o.PriceInPer1M, PriceOutPer1M: o.PriceOutPer1M, Currency: o.Currency, - }) + } + if o.PriceInPer1MPeak != nil || o.PriceOutPer1MPeak != nil || o.PriceCachedInPer1MPeak != nil { + view.PriceNote = "off-peak/base rate shown — this provider also has a higher peak-hours rate" + } + out = append(out, view) } return map[string]any{"offerings": out}, nil default: diff --git a/go/internal/statutil/statutil.go b/go/internal/statutil/statutil.go new file mode 100644 index 0000000..913694c --- /dev/null +++ b/go/internal/statutil/statutil.go @@ -0,0 +1,52 @@ +// SPDX-License-Identifier: Apache-2.0 + +// Package statutil holds small, dependency-free statistics helpers shared +// across packages that compute a distribution from a bounded set of raw +// samples (rather than a streaming histogram — this repo has no +// bucketed-histogram library anywhere, see internal/httpapi/cost_handlers.go +// and cmd/forge-compress/metrics.go's doc comments). Promoted 2026-09-11 +// from internal/httpapi/cost_handlers.go's unexported percentile/median +// (originally written for the cost/summary energy-calibration figures) so a +// second caller (cmd/forge-compress's overhead percentiles) doesn't +// reimplement the same ~20 lines with its own drift risk. +package statutil + +import "sort" + +// Median returns the median of vals (sorted copy; even-length averages the +// two middle values). 0 for an empty slice. +func Median(vals []float64) float64 { + if len(vals) == 0 { + return 0 + } + sorted := append([]float64(nil), vals...) + sort.Float64s(sorted) + mid := len(sorted) / 2 + if len(sorted)%2 == 1 { + return sorted[mid] + } + return (sorted[mid-1] + sorted[mid]) / 2 +} + +// Percentile returns the p-th percentile (0-100) of vals via nearest-rank. +// 0 for an empty slice — callers must check len(vals) against their own +// minimum-sample floor before trusting a figure computed from too few +// samples (this repo's established convention is 10 — see +// cost_handlers.go's activeSingleSlotWallW gate and +// compressor_summary_handlers.go's prefillObservedMinSamples, both of which +// cite this same floor by name). +func Percentile(vals []float64, p float64) float64 { + if len(vals) == 0 { + return 0 + } + sorted := append([]float64(nil), vals...) + sort.Float64s(sorted) + rank := int(p/100*float64(len(sorted)-1) + 0.5) + if rank < 0 { + rank = 0 + } + if rank >= len(sorted) { + rank = len(sorted) - 1 + } + return sorted[rank] +} diff --git a/go/internal/statutil/statutil_test.go b/go/internal/statutil/statutil_test.go new file mode 100644 index 0000000..813459f --- /dev/null +++ b/go/internal/statutil/statutil_test.go @@ -0,0 +1,51 @@ +// SPDX-License-Identifier: Apache-2.0 + +package statutil + +import "testing" + +func TestMedian(t *testing.T) { + cases := []struct { + name string + vals []float64 + want float64 + }{ + {"empty", nil, 0}, + {"single", []float64{7}, 7}, + {"odd", []float64{3, 1, 2}, 2}, + {"even", []float64{1, 2, 3, 4}, 2.5}, + } + for _, c := range cases { + t.Run(c.name, func(t *testing.T) { + if got := Median(c.vals); got != c.want { + t.Errorf("Median(%v) = %v, want %v", c.vals, got, c.want) + } + }) + } +} + +func TestPercentile(t *testing.T) { + // Nearest-rank on these 10 values: rank(p) = round(p/100*9). p50 -> + // round(4.5) = 5 -> sorted[5] = 60, not the interpolated 50 a + // linear-interpolation percentile would give — this pins down which + // convention every caller (computeEnergy's calibration block, + // forge-compress's overhead percentiles) actually gets. + vals := []float64{10, 20, 30, 40, 50, 60, 70, 80, 90, 100} + if got := Percentile(nil, 50); got != 0 { + t.Errorf("Percentile(nil, 50) = %v, want 0", got) + } + if got := Percentile(vals, 50); got != 60 { + t.Errorf("Percentile(vals, 50) = %v, want 60 (nearest-rank)", got) + } + if got := Percentile(vals, 95); got != 100 { + t.Errorf("Percentile(vals, 95) = %v, want 100 (nearest-rank)", got) + } + if got := Percentile(vals, 0); got != 10 { + t.Errorf("Percentile(vals, 0) = %v, want 10 (the minimum)", got) + } + // Order independence. + shuffled := []float64{100, 10, 90, 20, 80, 30, 70, 40, 60, 50} + if got := Percentile(shuffled, 50); got != 60 { + t.Errorf("Percentile(shuffled, 50) = %v, want 60", got) + } +} diff --git a/go/internal/store/catalog.go b/go/internal/store/catalog.go index 39cbf2f..284678e 100644 --- a/go/internal/store/catalog.go +++ b/go/internal/store/catalog.go @@ -267,6 +267,20 @@ type Offering struct { // computation must then price cached tokens at the full PriceInPer1M // rate (a documented upper bound, never an under-estimate). PriceCachedInPer1M *float64 + + // PriceInPer1MPeak/PriceOutPer1MPeak/PriceCachedInPer1MPeak (peak + // pricing sprint, 2026-09-12) are the rates in force during the + // provider's peak window (ProviderRow.PeakWindows) — e.g. DeepSeek's + // weekday UTC peak hours, currently 2x off-peak. nil on any of the + // three means "no peak differential for this field" and falls back to + // its base (PriceInPer1M/PriceOutPer1M/PriceCachedInPer1M) rate — same + // per-field-nil convention as PriceCachedInPer1M itself, deliberately + // not a single "has peak pricing" flag (a flag can disagree with the + // data). Meaningless when the owning provider has no PeakWindows + // configured; see internal/pricing.Windows. + PriceInPer1MPeak *float64 + PriceOutPer1MPeak *float64 + PriceCachedInPer1MPeak *float64 } // ── Annotations ────────────────────────────────────────────────────────────── @@ -1106,11 +1120,14 @@ func (v catalogView) CreateOffering(ctx context.Context, o Offering) (int64, err res, err := v.d.sql.ExecContext(ctx, `INSERT INTO offerings (model_id, variant_id, provider_id, wire_model, price_in_per_1m, price_out_per_1m, currency, context_length, enabled, - price_cached_in_per_1m, priority) - VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)`, + price_cached_in_per_1m, priority, + price_in_per_1m_peak, price_out_per_1m_peak, price_cached_in_per_1m_peak) + VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)`, o.ModelID, nullInt64(o.VariantID), o.ProviderID, o.WireModel, o.PriceInPer1M, o.PriceOutPer1M, o.Currency, o.ContextLength, - boolInt(o.Enabled), floatPtrArg(o.PriceCachedInPer1M), o.Priority) + boolInt(o.Enabled), floatPtrArg(o.PriceCachedInPer1M), o.Priority, + floatPtrArg(o.PriceInPer1MPeak), floatPtrArg(o.PriceOutPer1MPeak), + floatPtrArg(o.PriceCachedInPer1MPeak)) if err != nil { return 0, fmt.Errorf("store: catalog.create_offering: %w", err) } @@ -1120,22 +1137,27 @@ func (v catalogView) CreateOffering(ctx context.Context, o Offering) (int64, err const offeringSelectCols = `o.id, o.model_id, o.variant_id, o.provider_id, rp.name, o.wire_model, o.price_in_per_1m, o.price_out_per_1m, o.currency, o.context_length, o.enabled, - o.price_cached_in_per_1m, o.priority + o.price_cached_in_per_1m, o.priority, + o.price_in_per_1m_peak, o.price_out_per_1m_peak, o.price_cached_in_per_1m_peak FROM offerings o JOIN router_providers rp ON rp.id = o.provider_id` func scanOffering(s scanner) (Offering, error) { var o Offering var varID sql.NullInt64 var enabled int64 - var priceCachedIn sql.NullFloat64 + var priceCachedIn, priceInPeak, priceOutPeak, priceCachedInPeak sql.NullFloat64 if err := s.Scan(&o.ID, &o.ModelID, &varID, &o.ProviderID, &o.ProviderName, &o.WireModel, &o.PriceInPer1M, &o.PriceOutPer1M, &o.Currency, &o.ContextLength, - &enabled, &priceCachedIn, &o.Priority); err != nil { + &enabled, &priceCachedIn, &o.Priority, + &priceInPeak, &priceOutPeak, &priceCachedInPeak); err != nil { return Offering{}, err } o.VariantID = intOf(varID) o.Enabled = enabled != 0 o.PriceCachedInPer1M = nullFloat64Ptr(priceCachedIn) + o.PriceInPer1MPeak = nullFloat64Ptr(priceInPeak) + o.PriceOutPer1MPeak = nullFloat64Ptr(priceOutPeak) + o.PriceCachedInPer1MPeak = nullFloat64Ptr(priceCachedInPeak) return o, nil } @@ -1775,11 +1797,14 @@ func (v catalogView) UpdateOffering(ctx context.Context, o Offering) error { res, err := v.d.sql.ExecContext(ctx, `UPDATE offerings SET model_id=?, variant_id=?, provider_id=?, wire_model=?, price_in_per_1m=?, price_out_per_1m=?, currency=?, context_length=?, - enabled=?, price_cached_in_per_1m=?, priority=? + enabled=?, price_cached_in_per_1m=?, priority=?, + price_in_per_1m_peak=?, price_out_per_1m_peak=?, price_cached_in_per_1m_peak=? WHERE id=?`, o.ModelID, nullInt64(o.VariantID), o.ProviderID, o.WireModel, o.PriceInPer1M, o.PriceOutPer1M, o.Currency, o.ContextLength, - boolInt(o.Enabled), floatPtrArg(o.PriceCachedInPer1M), o.Priority, o.ID) + boolInt(o.Enabled), floatPtrArg(o.PriceCachedInPer1M), o.Priority, + floatPtrArg(o.PriceInPer1MPeak), floatPtrArg(o.PriceOutPer1MPeak), + floatPtrArg(o.PriceCachedInPer1MPeak), o.ID) if err != nil { return fmt.Errorf("store: catalog.update_offering: %w", err) } diff --git a/go/internal/store/catalog_test.go b/go/internal/store/catalog_test.go index 6ca993a..1e97b16 100644 --- a/go/internal/store/catalog_test.go +++ b/go/internal/store/catalog_test.go @@ -675,6 +675,89 @@ func TestListOfferingsPriceCachedInPer1M(t *testing.T) { } } +// TestOfferingPeakPricingRoundTrip guards offeringSelectCols/scanOffering +// (the single shared scan behind ListOfferings/ListOfferingsForModel/ +// GetOffering) against a missed column for the three peak price fields — +// exactly the class of bug that only fails at runtime, not compile time. +func TestOfferingPeakPricingRoundTrip(t *testing.T) { + db := openTest(t) + ctx := context.Background() + cat := db.Catalog() + + mdlID, err := cat.CreateModel(ctx, Model{Name: "TestModel"}) + if err != nil { + t.Fatalf("CreateModel: %v", err) + } + if _, err := db.SQL().ExecContext(ctx, + `INSERT INTO router_providers (name, api_key, created_at) VALUES ('deepseek', 'key', 0)`); err != nil { + t.Fatalf("seed provider: %v", err) + } + providerID := testProviderID(t, db, "deepseek") + + peakIn, peakOut, peakCached := 0.3, 1.2, 0.006 + offID, err := cat.CreateOffering(ctx, Offering{ + ModelID: mdlID, ProviderID: providerID, WireModel: "deepseek-flash", + PriceInPer1M: 0.15, PriceOutPer1M: 0.6, Currency: "USD", Enabled: true, + PriceInPer1MPeak: &peakIn, PriceOutPer1MPeak: &peakOut, PriceCachedInPer1MPeak: &peakCached, + }) + if err != nil { + t.Fatalf("CreateOffering: %v", err) + } + + assertPeak := func(t *testing.T, o Offering, label string) { + t.Helper() + if o.PriceInPer1MPeak == nil || *o.PriceInPer1MPeak != peakIn { + t.Errorf("%s: PriceInPer1MPeak = %v, want %v", label, o.PriceInPer1MPeak, peakIn) + } + if o.PriceOutPer1MPeak == nil || *o.PriceOutPer1MPeak != peakOut { + t.Errorf("%s: PriceOutPer1MPeak = %v, want %v", label, o.PriceOutPer1MPeak, peakOut) + } + if o.PriceCachedInPer1MPeak == nil || *o.PriceCachedInPer1MPeak != peakCached { + t.Errorf("%s: PriceCachedInPer1MPeak = %v, want %v", label, o.PriceCachedInPer1MPeak, peakCached) + } + } + + got, err := cat.GetOffering(ctx, offID) + if err != nil { + t.Fatalf("GetOffering: %v", err) + } + assertPeak(t, got, "GetOffering") + + list, err := cat.ListOfferings(ctx) + if err != nil { + t.Fatalf("ListOfferings: %v", err) + } + if len(list) != 1 { + t.Fatalf("ListOfferings: got %d rows, want 1", len(list)) + } + assertPeak(t, list[0], "ListOfferings") + + forModel, err := cat.ListOfferingsForModel(ctx, mdlID) + if err != nil { + t.Fatalf("ListOfferingsForModel: %v", err) + } + if len(forModel) != 1 { + t.Fatalf("ListOfferingsForModel: got %d rows, want 1", len(forModel)) + } + assertPeak(t, forModel[0], "ListOfferingsForModel") + + // A NULL peak field must round-trip as nil, not a zero value — Update + // clears all three. + got.PriceInPer1MPeak = nil + got.PriceOutPer1MPeak = nil + got.PriceCachedInPer1MPeak = nil + if err := cat.UpdateOffering(ctx, got); err != nil { + t.Fatalf("UpdateOffering: %v", err) + } + cleared, err := cat.GetOffering(ctx, offID) + if err != nil { + t.Fatalf("GetOffering after clear: %v", err) + } + if cleared.PriceInPer1MPeak != nil || cleared.PriceOutPer1MPeak != nil || cleared.PriceCachedInPer1MPeak != nil { + t.Errorf("peak fields not cleared after UpdateOffering(nil): %+v", cleared) + } +} + // TestOfferingPriorityRoundTripAndOrdering covers the 0032 priority column: // it round-trips through Create/Get/Update, and ListOfferings orders by // (priority, provider, wire_model) — the exact order the router's group diff --git a/go/internal/store/db_test.go b/go/internal/store/db_test.go index 58b886a..01abb18 100644 --- a/go/internal/store/db_test.go +++ b/go/internal/store/db_test.go @@ -19,8 +19,8 @@ func TestMigrateFresh(t *testing.T) { ).Scan(&version); err != nil { t.Fatalf("read version: %v", err) } - if version != 76 { - t.Fatalf("schema version = %d, want 76", version) + if version != 78 { + t.Fatalf("schema version = %d, want 78", version) } // Every Contract 3 table (0001) plus the Sprint 0 §0.11 polish tables diff --git a/go/internal/store/migration_0078_test.go b/go/internal/store/migration_0078_test.go new file mode 100644 index 0000000..04e6f63 --- /dev/null +++ b/go/internal/store/migration_0078_test.go @@ -0,0 +1,94 @@ +// SPDX-License-Identifier: Apache-2.0 + +package store + +import "testing" + +// TestMigration0078PeakPricingIsNullableAdditive seeds a pre-0078 offering +// + provider + usage event, applies 0078, and confirms: the new columns +// exist and read back NULL/"" (no backfill — every existing row keeps +// meaning exactly what it meant before), and the pre-existing base price +// and event data survive untouched. +func TestMigration0078PeakPricingIsNullableAdditive(t *testing.T) { + sqlDB := openThrough(t, 77) + defer sqlDB.Close() + + if _, err := sqlDB.Exec( + `INSERT INTO router_providers (name, api_key, created_at) VALUES ('deepseek', 'sk-x', 0)`, + ); err != nil { + t.Fatalf("seed router_providers: %v", err) + } + var providerID int64 + if err := sqlDB.QueryRow(`SELECT id FROM router_providers WHERE name = 'deepseek'`).Scan(&providerID); err != nil { + t.Fatalf("resolve provider id: %v", err) + } + + if _, err := sqlDB.Exec( + `INSERT INTO models (family_id, name) VALUES (NULL, 'deepseek-v4-flash')`, + ); err != nil { + t.Fatalf("seed models: %v", err) + } + var modelID int64 + if err := sqlDB.QueryRow(`SELECT id FROM models WHERE name = 'deepseek-v4-flash'`).Scan(&modelID); err != nil { + t.Fatalf("resolve model id: %v", err) + } + + if _, err := sqlDB.Exec( + `INSERT INTO offerings (model_id, provider_id, wire_model, price_in_per_1m, price_out_per_1m, currency, enabled, priority) + VALUES (?, ?, 'deepseek-v4-flash', 0.14, 0.28, 'USD', 1, 100)`, + modelID, providerID, + ); err != nil { + t.Fatalf("seed offerings: %v", err) + } + + if _, err := sqlDB.Exec( + `INSERT INTO usage_events (ts, kind, model, provider_id, prompt_tokens, completion_tokens, cost_usd) + VALUES (0, 'external_request', 'deepseek-v4-flash', ?, 1000, 500, 0.001)`, + providerID, + ); err != nil { + t.Fatalf("seed usage_events: %v", err) + } + + body, err := migrationsFS.ReadFile("migrations/0078_peak_pricing.sql") + if err != nil { + t.Fatalf("read 0078: %v", err) + } + if _, err := sqlDB.Exec(string(body)); err != nil { + t.Fatalf("apply 0078: %v", err) + } + + var priceIn, priceOut float64 + var peakIn, peakOut, peakCached any + if err := sqlDB.QueryRow( + `SELECT price_in_per_1m, price_out_per_1m, price_in_per_1m_peak, price_out_per_1m_peak, price_cached_in_per_1m_peak + FROM offerings WHERE model_id = ?`, modelID, + ).Scan(&priceIn, &priceOut, &peakIn, &peakOut, &peakCached); err != nil { + t.Fatalf("read offering after 0078: %v", err) + } + if priceIn != 0.14 || priceOut != 0.28 { + t.Errorf("base prices changed: in=%v out=%v, want 0.14/0.28 (no backfill)", priceIn, priceOut) + } + if peakIn != nil || peakOut != nil || peakCached != nil { + t.Errorf("peak columns not NULL after migration: in=%v out=%v cached=%v", peakIn, peakOut, peakCached) + } + + var peakWindows any + if err := sqlDB.QueryRow(`SELECT peak_windows FROM router_providers WHERE id = ?`, providerID).Scan(&peakWindows); err != nil { + t.Fatalf("read provider after 0078: %v", err) + } + if peakWindows != nil { + t.Errorf("peak_windows not NULL after migration: %v", peakWindows) + } + + var promptTokens int64 + var priceTier any + if err := sqlDB.QueryRow(`SELECT prompt_tokens, price_tier FROM usage_events LIMIT 1`).Scan(&promptTokens, &priceTier); err != nil { + t.Fatalf("read usage_event after 0078: %v", err) + } + if promptTokens != 1000 { + t.Errorf("pre-existing event data changed: prompt_tokens=%v, want 1000", promptTokens) + } + if priceTier != nil { + t.Errorf("price_tier not NULL for pre-migration event: %v", priceTier) + } +} diff --git a/go/internal/store/migrations/0077_compressor_overhead_percentiles.sql b/go/internal/store/migrations/0077_compressor_overhead_percentiles.sql new file mode 100644 index 0000000..4b9a425 --- /dev/null +++ b/go/internal/store/migrations/0077_compressor_overhead_percentiles.sql @@ -0,0 +1,11 @@ +-- Compressor investigation (2026-09-11): the mean-only overhead figure was +-- found to hide a bimodal real-traffic shape — most messages barely pay the +-- compression tax, a few huge ones pay a lot. Add nullable percentile +-- columns (latest-window-sample gauges, same convention as the existing +-- overhead_min_ms/overhead_max_ms — never summed/averaged) computed by +-- cmd/forge-compress's new bounded-ring overhead sampler +-- (percentileMinSamples = 10; NULL below that floor, never a fabricated 0). + +ALTER TABLE compressor_savings_samples ADD COLUMN overhead_p50_ms REAL; +ALTER TABLE compressor_savings_samples ADD COLUMN overhead_p90_ms REAL; +ALTER TABLE compressor_savings_samples ADD COLUMN overhead_p99_ms REAL; diff --git a/go/internal/store/migrations/0078_peak_pricing.sql b/go/internal/store/migrations/0078_peak_pricing.sql new file mode 100644 index 0000000..4171ea8 --- /dev/null +++ b/go/internal/store/migrations/0078_peak_pricing.sql @@ -0,0 +1,31 @@ +-- DeepSeek switched to time-of-day pricing on 2026-09-10 (peak hours cost +-- 2x off-peak). offerings previously stored exactly one flat price triple, +-- so roughly half of all DeepSeek spend was mispriced regardless of what +-- number was entered. All new columns are nullable additive: the existing +-- price_in_per_1m/price_out_per_1m/price_cached_in_per_1m columns keep +-- meaning "the price when no peak window is active" (off-peak), and a NULL +-- peak column falls back per-field to its base value — every existing +-- offering/provider/event is unchanged and correct by construction, no +-- backfill needed. See internal/pricing for the window-evaluation logic and +-- go/internal/router/usage.go's computeCostNative for how these are priced. + +ALTER TABLE offerings ADD COLUMN price_in_per_1m_peak REAL; +ALTER TABLE offerings ADD COLUMN price_out_per_1m_peak REAL; +ALTER TABLE offerings ADD COLUMN price_cached_in_per_1m_peak REAL; + +-- peak_windows is a JSON-encoded internal/pricing.Windows — a provider-wide +-- schedule (DeepSeek's peak hours apply to every one of its models, not one +-- specific model), so it lives on the provider, not the offering. NULL/"" +-- means no peak/off-peak concept for this provider (internal/pricing.Parse +-- treats "" as the zero value); every request is priced at the base rate. +ALTER TABLE router_providers ADD COLUMN peak_windows TEXT; + +-- price_tier records which tier ("peak"/"off_peak"/"flat") was in force +-- when a usage event's cost was computed, so historical reporting +-- (compressor_summary_handlers.go's savings estimators) can price past +-- events accurately instead of re-deriving from today's prices, and so an +-- operator can audit a recorded cost against the provider's invoice. NULL +-- for every event recorded before this migration or for an event whose +-- pricing wasn't resolvable at all (mirrors CostNative's existing nil +-- convention). +ALTER TABLE usage_events ADD COLUMN price_tier TEXT; diff --git a/go/internal/store/providers.go b/go/internal/store/providers.go index d7de93e..1636e21 100644 --- a/go/internal/store/providers.go +++ b/go/internal/store/providers.go @@ -78,7 +78,7 @@ func (v providersView) List(ctx context.Context) ([]ProviderRow, error) { `SELECT id, name, api_key, target_url, model, model2, bill_currency, status_url, credits_url, org_id, billing_enabled, billing_console_url, enabled, country, - data_residency_group, created_at + data_residency_group, created_at, peak_windows FROM router_providers WHERE deleted_at IS NULL ORDER BY name`) if err != nil { return nil, fmt.Errorf("store: providers.list: %w", err) @@ -92,15 +92,17 @@ func (v providersView) List(ctx context.Context) ([]ProviderRow, error) { var p ProviderRow var created int64 var billingEnabled, enabled int64 + var peakWindows sql.NullString if err := rows.Scan(&p.ID, &p.Name, &p.APIKey, &p.TargetURL, &p.Model, &p.Model2, &p.BillCurrency, &p.StatusURL, &p.CreditsURL, &p.OrgID, &billingEnabled, &p.BillingConsoleURL, &enabled, - &p.Country, &p.DataResidencyGroup, &created); err != nil { + &p.Country, &p.DataResidencyGroup, &created, &peakWindows); err != nil { return nil, fmt.Errorf("store: providers.list: %w", err) } p.BillingEnabled = billingEnabled != 0 p.Enabled = enabled != 0 p.CreatedAt = timeOf(sql.NullInt64{Int64: created, Valid: true}) + p.PeakWindows = strOf(peakWindows) out = append(out, p) } if err := rows.Err(); err != nil { diff --git a/go/internal/store/providers_test.go b/go/internal/store/providers_test.go index 436fa56..e76a20c 100644 --- a/go/internal/store/providers_test.go +++ b/go/internal/store/providers_test.go @@ -370,3 +370,72 @@ func TestProviderEnabledCountryRoundTrip(t *testing.T) { t.Error("Compressor.Providers: enabled should be false after disable") } } + +// TestProviderPeakWindowsRoundTrip guards all four provider read paths +// against a missed peak_windows SELECT column — routing.go's Providers() +// and providerBy() (backing both ProviderByID and ProviderByName), plus +// providers.go's List(). The last of those feeds the Settings UI behind a +// full-replace PUT (RemoteOfferings.tsx's own doc comment on the same +// hazard for offerings), so a column missed there would load an empty +// editor and let an operator silently delete a real schedule on save. +func TestProviderPeakWindowsRoundTrip(t *testing.T) { + db := openTest(t) + ctx := context.Background() + + const windows = `{"windows":[{"days":[1,2,3,4,5],"start":"01:00","end":"04:00"}]}` + if err := db.Routing().SaveProvider(ctx, ProviderRow{ + Name: "deepseek", APIKey: "sk-d", TargetURL: "https://api.deepseek.com/v1", + Enabled: true, PeakWindows: windows, CreatedAt: ts(2000), + }); err != nil { + t.Fatalf("SaveProvider: %v", err) + } + + routingRows, err := db.Routing().Providers(ctx) + if err != nil { + t.Fatalf("Routing.Providers: %v", err) + } + if len(routingRows) != 1 || routingRows[0].PeakWindows != windows { + t.Errorf("Routing.Providers PeakWindows = %q, want %q", routingRows[0].PeakWindows, windows) + } + + byName, ok, err := db.Routing().ProviderByName(ctx, "deepseek") + if err != nil || !ok { + t.Fatalf("ProviderByName: ok=%v err=%v", ok, err) + } + if byName.PeakWindows != windows { + t.Errorf("ProviderByName PeakWindows = %q, want %q", byName.PeakWindows, windows) + } + + byID, ok, err := db.Routing().ProviderByID(ctx, byName.ID) + if err != nil || !ok { + t.Fatalf("ProviderByID: ok=%v err=%v", ok, err) + } + if byID.PeakWindows != windows { + t.Errorf("ProviderByID PeakWindows = %q, want %q", byID.PeakWindows, windows) + } + + listRows, err := db.Providers().List(ctx) + if err != nil { + t.Fatalf("Providers.List: %v", err) + } + if len(listRows) != 1 || listRows[0].PeakWindows != windows { + t.Errorf("Providers.List PeakWindows = %q, want %q", listRows[0].PeakWindows, windows) + } + + // A full-replace SaveProvider (the same idiom the Settings PUT handler + // uses) that omits PeakWindows must clear it, not silently leave the + // old value behind under some UPDATE-skips-empty-string quirk. + if err := db.Routing().SaveProvider(ctx, ProviderRow{ + ID: byName.ID, Name: "deepseek", APIKey: "sk-d", TargetURL: "https://api.deepseek.com/v1", + Enabled: true, CreatedAt: ts(2000), + }); err != nil { + t.Fatalf("SaveProvider (clear): %v", err) + } + cleared, ok, err := db.Routing().ProviderByName(ctx, "deepseek") + if err != nil || !ok { + t.Fatalf("ProviderByName after clear: ok=%v err=%v", ok, err) + } + if cleared.PeakWindows != "" { + t.Errorf("PeakWindows not cleared: %q", cleared.PeakWindows) + } +} diff --git a/go/internal/store/routing.go b/go/internal/store/routing.go index 9c28b64..993a2f5 100644 --- a/go/internal/store/routing.go +++ b/go/internal/store/routing.go @@ -140,13 +140,13 @@ func (v routingView) SaveProvider(ctx context.Context, p ProviderRow) error { `INSERT INTO router_providers (name, api_key, target_url, model, model2, bill_currency, status_url, credits_url, org_id, billing_enabled, billing_console_url, enabled, country, - data_residency_group, created_at) - VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)`, + data_residency_group, created_at, peak_windows) + VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)`, p.Name, p.APIKey, p.TargetURL, p.Model, p.Model2, billCcy, p.StatusURL, p.CreditsURL, p.OrgID, boolInt(p.BillingEnabled), p.BillingConsoleURL, boolInt(p.Enabled), p.Country, p.DataResidencyGroup, - unixOf(orNow(p.CreatedAt)), + unixOf(orNow(p.CreatedAt)), p.PeakWindows, ) if err != nil { return fmt.Errorf("store: routing.save_provider: %w", err) @@ -163,12 +163,12 @@ func (v routingView) SaveProvider(ctx context.Context, p ProviderRow) error { bill_currency = COALESCE(NULLIF(?, ''), 'USD'), status_url = ?, credits_url = ?, org_id = ?, billing_enabled = ?, billing_console_url = ?, - enabled = ?, country = ?, data_residency_group = ? + enabled = ?, country = ?, data_residency_group = ?, peak_windows = ? WHERE id = ?`, p.Name, p.APIKey, p.TargetURL, p.Model, p.Model2, billCcy, p.StatusURL, p.CreditsURL, p.OrgID, boolInt(p.BillingEnabled), p.BillingConsoleURL, boolInt(p.Enabled), - p.Country, p.DataResidencyGroup, p.ID, + p.Country, p.DataResidencyGroup, p.PeakWindows, p.ID, ) if err != nil { return fmt.Errorf("store: routing.save_provider: %w", err) @@ -200,7 +200,7 @@ func (v routingView) Providers(ctx context.Context) ([]ProviderRow, error) { `SELECT rp.id, rp.name, rp.api_key, rp.target_url, rp.model, rp.model2, rp.bill_currency, rp.status_url, rp.credits_url, rp.org_id, rp.billing_enabled, rp.billing_console_url, rp.enabled, rp.country, - rp.data_residency_group, rp.created_at, hp.service + rp.data_residency_group, rp.created_at, hp.service, rp.peak_windows FROM router_providers rp LEFT JOIN compressor_proxies hp ON hp.provider_id = rp.id AND hp.orphaned_at IS NULL @@ -215,17 +215,18 @@ func (v routingView) Providers(ctx context.Context) ([]ProviderRow, error) { var p ProviderRow var created int64 var billingEnabled, enabled int64 - var proxyService sql.NullString + var proxyService, peakWindows sql.NullString if err := rows.Scan(&p.ID, &p.Name, &p.APIKey, &p.TargetURL, &p.Model, &p.Model2, &p.BillCurrency, &p.StatusURL, &p.CreditsURL, &p.OrgID, &billingEnabled, &p.BillingConsoleURL, &enabled, - &p.Country, &p.DataResidencyGroup, &created, &proxyService); err != nil { + &p.Country, &p.DataResidencyGroup, &created, &proxyService, &peakWindows); err != nil { return nil, fmt.Errorf("store: routing.providers: %w", err) } p.BillingEnabled = billingEnabled != 0 p.Enabled = enabled != 0 p.CreatedAt = timeOf(sql.NullInt64{Int64: created, Valid: true}) p.CompressorProxyName = strOf(proxyService) + p.PeakWindows = strOf(peakWindows) out = append(out, p) } if err := rows.Err(); err != nil { @@ -252,12 +253,13 @@ func (v routingView) providerBy(ctx context.Context, where string, arg any) (Pro var created int64 var billingEnabled, enabled int64 var deletedAt sql.NullInt64 - var proxyService sql.NullString + var proxyService, peakWindows sql.NullString err := v.d.sql.QueryRowContext(ctx, `SELECT rp.id, rp.name, rp.api_key, rp.target_url, rp.model, rp.model2, rp.bill_currency, rp.status_url, rp.credits_url, rp.org_id, rp.billing_enabled, rp.billing_console_url, rp.enabled, rp.country, - rp.data_residency_group, rp.deleted_at, rp.created_at, hp.service + rp.data_residency_group, rp.deleted_at, rp.created_at, hp.service, + rp.peak_windows FROM router_providers rp LEFT JOIN compressor_proxies hp ON hp.provider_id = rp.id AND hp.orphaned_at IS NULL @@ -265,7 +267,7 @@ func (v routingView) providerBy(ctx context.Context, where string, arg any) (Pro arg).Scan(&p.ID, &p.Name, &p.APIKey, &p.TargetURL, &p.Model, &p.Model2, &p.BillCurrency, &p.StatusURL, &p.CreditsURL, &p.OrgID, &billingEnabled, &p.BillingConsoleURL, &enabled, &p.Country, - &p.DataResidencyGroup, &deletedAt, &created, &proxyService) + &p.DataResidencyGroup, &deletedAt, &created, &proxyService, &peakWindows) if err == sql.ErrNoRows { return ProviderRow{}, false, nil } @@ -277,6 +279,7 @@ func (v routingView) providerBy(ctx context.Context, where string, arg any) (Pro p.CreatedAt = timeOf(sql.NullInt64{Int64: created, Valid: true}) p.DeletedAt = timeOf(deletedAt) p.CompressorProxyName = strOf(proxyService) + p.PeakWindows = strOf(peakWindows) return p, true, nil } @@ -342,8 +345,9 @@ func (v routingView) RecordSavingsSample(ctx context.Context, s CompressorSaving cache_busts, cache_bust_tokens_lost, ttfb_count, ttfb_sum_ms, ttfb_min_ms, ttfb_max_ms, latency_count, latency_sum_ms, latency_min_ms, latency_max_ms, - overhead_count, overhead_sum_ms, overhead_min_ms, overhead_max_ms - ) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)`, + overhead_count, overhead_sum_ms, overhead_min_ms, overhead_max_ms, + overhead_p50_ms, overhead_p90_ms, overhead_p99_ms + ) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)`, ts, s.ProxyID, s.TokensIn, s.TokensOut, s.CacheReadTokens, s.UncachedTokens, s.TokensSaved, s.Requests, s.RequestsCached, s.RequestsFailed, s.RequestsRateLimited, s.RequestsTimeout, s.RequestsCanceled, s.FailOpenTotal, @@ -351,6 +355,7 @@ func (v routingView) RecordSavingsSample(ctx context.Context, s CompressorSaving s.TTFBCount, s.TTFBSumMs, floatPtrArg(s.TTFBMinMs), floatPtrArg(s.TTFBMaxMs), s.LatencyCount, s.LatencySumMs, floatPtrArg(s.LatencyMinMs), floatPtrArg(s.LatencyMaxMs), s.OverheadCount, s.OverheadSumMs, floatPtrArg(s.OverheadMinMs), floatPtrArg(s.OverheadMaxMs), + floatPtrArg(s.OverheadP50Ms), floatPtrArg(s.OverheadP90Ms), floatPtrArg(s.OverheadP99Ms), ) if err != nil { return fmt.Errorf("store: routing.record_sample: %w", err) @@ -384,7 +389,8 @@ func (v routingView) SavingsSummary(ctx context.Context, since time.Time) (map[s hs.cache_read_tokens, hs.uncached_tokens, hs.cache_busts, hs.cache_bust_tokens_lost, hs.ttfb_count, hs.ttfb_sum_ms, hs.ttfb_min_ms, hs.ttfb_max_ms, hs.latency_count, hs.latency_sum_ms, hs.latency_min_ms, hs.latency_max_ms, - hs.overhead_count, hs.overhead_sum_ms, hs.overhead_min_ms, hs.overhead_max_ms + hs.overhead_count, hs.overhead_sum_ms, hs.overhead_min_ms, hs.overhead_max_ms, + hs.overhead_p50_ms, hs.overhead_p90_ms, hs.overhead_p99_ms FROM compressor_savings_samples hs JOIN compressor_proxies hp ON hp.id = hs.proxy_id WHERE hs.ts >= ? ORDER BY hs.ts ASC`, @@ -405,6 +411,7 @@ func (v routingView) SavingsSummary(ctx context.Context, since time.Time) (map[s var ttfbCount, latencyCount, overheadCount int64 var ttfbSum, latencySum, overheadSum float64 var ttfbMin, ttfbMax, latencyMin, latencyMax, overheadMin, overheadMax sql.NullFloat64 + var overheadP50, overheadP90, overheadP99 sql.NullFloat64 if err := rows.Scan(&proxy, &tokensIn, &tokensOut, &tokensSaved, &requests, &requestsCached, &requestsFailed, &requestsRateLimited, &requestsTimeout, &requestsCanceled, &failOpenTotal, @@ -412,6 +419,7 @@ func (v routingView) SavingsSummary(ctx context.Context, since time.Time) (map[s &ttfbCount, &ttfbSum, &ttfbMin, &ttfbMax, &latencyCount, &latencySum, &latencyMin, &latencyMax, &overheadCount, &overheadSum, &overheadMin, &overheadMax, + &overheadP50, &overheadP90, &overheadP99, ); err != nil { return nil, fmt.Errorf("store: routing.summary: %w", err) } @@ -455,6 +463,15 @@ func (v routingView) SavingsSummary(ctx context.Context, since time.Time) (map[s if v := nullFloat64Ptr(overheadMax); v != nil { p.OverheadMaxMs = v } + if v := nullFloat64Ptr(overheadP50); v != nil { + p.OverheadP50Ms = v + } + if v := nullFloat64Ptr(overheadP90); v != nil { + p.OverheadP90Ms = v + } + if v := nullFloat64Ptr(overheadP99); v != nil { + p.OverheadP99Ms = v + } out[proxy] = p } if err := rows.Err(); err != nil { @@ -518,6 +535,15 @@ func (v routingView) SavingsSummary(ctx context.Context, since time.Time) (map[s p.RequestsByModel = map[string]int64{} } p.RequestsByModel[labelValue] = sum + case "outcome_size": + // labelValue is a composite "outcome:size_tier" (e.g. + // "compressed:huge") — see messagesByOutcomeSize's doc comment + // in cmd/forge-compress/metrics.go for why composite rather than + // two independent label dimensions. + if p.MessagesByOutcomeSize == nil { + p.MessagesByOutcomeSize = map[string]int64{} + } + p.MessagesByOutcomeSize[labelValue] = sum case "transform": switch metric { case "timing_ms_sum": diff --git a/go/internal/store/routing_test.go b/go/internal/store/routing_test.go index e1b9f75..ddfb270 100644 --- a/go/internal/store/routing_test.go +++ b/go/internal/store/routing_test.go @@ -56,9 +56,11 @@ func TestRecordCompressorSampleAndSummary(t *testing.T) { TTFBCount: 10, TTFBSumMs: 500, TTFBMinMs: f64ptr(20), TTFBMaxMs: f64ptr(80), LatencyCount: 10, LatencySumMs: 4000, LatencyMinMs: f64ptr(200), LatencyMaxMs: f64ptr(900), OverheadCount: 10, OverheadSumMs: 50, OverheadMinMs: f64ptr(2), OverheadMaxMs: f64ptr(9), + OverheadP50Ms: f64ptr(4), OverheadP90Ms: f64ptr(8), OverheadP99Ms: f64ptr(9), }, []CompressorLabelSample{ {TS: base, ProxyID: localID, LabelKey: "provider", LabelValue: "openai", Metric: "requests", Delta: 8}, {TS: base, ProxyID: localID, LabelKey: "provider", LabelValue: "anthropic", Metric: "requests", Delta: 2}, + {TS: base, ProxyID: localID, LabelKey: "outcome_size", LabelValue: "compressed:medium", Metric: "messages", Delta: 6}, }); err != nil { t.Fatalf("RecordSavingsSample 1: %v", err) } @@ -75,9 +77,12 @@ func TestRecordCompressorSampleAndSummary(t *testing.T) { TTFBCount: 5, TTFBSumMs: 300, TTFBMinMs: f64ptr(15), TTFBMaxMs: f64ptr(95), LatencyCount: 5, LatencySumMs: 1500, LatencyMinMs: f64ptr(180), LatencyMaxMs: f64ptr(300), OverheadCount: 5, OverheadSumMs: 30, OverheadMinMs: f64ptr(1), OverheadMaxMs: f64ptr(6), + OverheadP50Ms: f64ptr(3), OverheadP90Ms: f64ptr(5), OverheadP99Ms: f64ptr(6), }, []CompressorLabelSample{ {TS: next, ProxyID: localID, LabelKey: "provider", LabelValue: "openai", Metric: "requests", Delta: 4}, {TS: next, ProxyID: localID, LabelKey: "model", LabelValue: "gemma4-e2b", Metric: "requests", Delta: 5}, + {TS: next, ProxyID: localID, LabelKey: "outcome_size", LabelValue: "compressed:huge", Metric: "messages", Delta: 3}, + {TS: next, ProxyID: localID, LabelKey: "outcome_size", LabelValue: "gated_passthrough:small", Metric: "messages", Delta: 2}, }); err != nil { t.Fatalf("RecordSavingsSample 2: %v", err) } @@ -132,6 +137,23 @@ func TestRecordCompressorSampleAndSummary(t *testing.T) { if local.RequestsByModel["gemma4-e2b"] != 5 { t.Errorf("RequestsByModel[gemma4-e2b] = %d, want 5", local.RequestsByModel["gemma4-e2b"]) } + // Same latest-sample-wins gauge semantics as TTFBMax/LatencyMax above — + // the second sample's percentiles win, not an average of the two. + if local.OverheadP50Ms == nil || *local.OverheadP50Ms != 3 { + t.Errorf("OverheadP50Ms = %v, want 3 (the latest sample's value)", local.OverheadP50Ms) + } + if local.OverheadP99Ms == nil || *local.OverheadP99Ms != 6 { + t.Errorf("OverheadP99Ms = %v, want 6 (the latest sample's value, even though lower than the earlier sample's 9)", local.OverheadP99Ms) + } + if local.MessagesByOutcomeSize["compressed:medium"] != 6 { + t.Errorf("MessagesByOutcomeSize[compressed:medium] = %d, want 6", local.MessagesByOutcomeSize["compressed:medium"]) + } + if local.MessagesByOutcomeSize["compressed:huge"] != 3 { + t.Errorf("MessagesByOutcomeSize[compressed:huge] = %d, want 3", local.MessagesByOutcomeSize["compressed:huge"]) + } + if local.MessagesByOutcomeSize["gated_passthrough:small"] != 2 { + t.Errorf("MessagesByOutcomeSize[gated_passthrough:small] = %d, want 2", local.MessagesByOutcomeSize["gated_passthrough:small"]) + } deepseek, ok := summary["deepseek"] if !ok { @@ -143,6 +165,9 @@ func TestRecordCompressorSampleAndSummary(t *testing.T) { if deepseek.TTFBMinMs != nil { t.Errorf("deepseek TTFBMinMs = %v, want nil (never recorded)", deepseek.TTFBMinMs) } + if deepseek.OverheadP50Ms != nil { + t.Errorf("deepseek OverheadP50Ms = %v, want nil (never recorded)", deepseek.OverheadP50Ms) + } } func TestCompressorSummaryWindowExcludesOldSamples(t *testing.T) { diff --git a/go/internal/store/store.go b/go/internal/store/store.go index da55d87..59a0477 100644 --- a/go/internal/store/store.go +++ b/go/internal/store/store.go @@ -260,6 +260,15 @@ type UsageEvent struct { // pre-existing call site (OnTokenSample, load/unload events) is // automatically correct without being touched. Unmetered bool + + // PriceTier (peak pricing sprint, 2026-09-12) records which tier — + // "peak"/"off_peak"/"flat" (internal/pricing.Tier* constants) — was in + // force when CostNative was computed, so historical reporting can price + // past events accurately without re-deriving from today's rates, and an + // operator can audit a recorded cost against the provider's invoice. + // "" when CostNative is nil (pricing wasn't resolvable) or for any row + // written before this migration. + PriceTier string } type ModeHistoryEntry struct { @@ -469,6 +478,12 @@ type CompressorSavingsSampleRow struct { OverheadCount int64 OverheadSumMs float64 OverheadMinMs, OverheadMaxMs *float64 + // OverheadP50Ms/P90Ms/P99Ms are percentiles of the proxy's own recent + // (bounded-ring, NOT lifetime) overhead samples — a latest-snapshot + // gauge like Min/Max above, never summed/averaged across a window. nil + // below the 10-sample floor. See collector.CompressorSample's matching + // doc comment for why "recent" (not "since start") here. + OverheadP50Ms, OverheadP90Ms, OverheadP99Ms *float64 } // CompressorLabelSample is one interval's per-(proxy,label) delta for a @@ -512,9 +527,17 @@ type CompressorProxySummary struct { OverheadCount int64 OverheadSumMs float64 OverheadMinMs, OverheadMaxMs *float64 + // OverheadP50Ms/P90Ms/P99Ms: latest window sample's percentiles (never + // summed/averaged) — see CompressorSavingsSampleRow's matching field. + OverheadP50Ms, OverheadP90Ms, OverheadP99Ms *float64 RequestsByProvider map[string]int64 RequestsByModel map[string]int64 + // MessagesByOutcomeSize is per-MESSAGE counts (not per-request) keyed by + // a composite "outcome:size_tier" label value — from + // compress_messages_total{outcome_size}. See + // cmd/forge-compress/messages.go's messageOutcomeSize. + MessagesByOutcomeSize map[string]int64 // Per-provider cache token breakdowns (from compress_cache_read_tokens_total, // compress_uncached_input_tokens_total, etc.). @@ -638,6 +661,15 @@ type ProviderRow struct { Country string DataResidencyGroup string CreatedAt time.Time + // PeakWindows (peak pricing sprint, 2026-09-12, 0078_peak_pricing.sql) + // is a JSON-encoded internal/pricing.Windows: this provider's recurring + // UTC peak-hours schedule (e.g. DeepSeek's weekday 01:00-04:00 + + // 06:00-10:00). Lives on the provider, not the offering, because a + // peak schedule is a fact about the provider's billing policy shared + // by every offering it serves. ""/unset means no peak/off-peak concept + // — every request is priced at the offering's base rate + // unconditionally (internal/pricing.Windows.TierAt returns "flat"). + PeakWindows string } type SavingsTotal struct { diff --git a/go/internal/store/usage.go b/go/internal/store/usage.go index 9f24690..3cfec79 100644 --- a/go/internal/store/usage.go +++ b/go/internal/store/usage.go @@ -19,13 +19,13 @@ func (v usageView) Record(ctx context.Context, e UsageEvent) error { _, err := v.d.sql.ExecContext(ctx, `INSERT INTO usage_events (ts, kind, model, slot, provider_id, prompt_tokens, completion_tokens, cost_usd, detail, - cost_native, cost_currency, cached_prompt_tokens, unmetered) - VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)`, + cost_native, cost_currency, cached_prompt_tokens, unmetered, price_tier) + VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)`, unixOf(orNow(e.TS)), e.Kind, nullStr(e.Model), nullStr(e.Slot), intPtrArg(e.ProviderID), e.PromptTokens, e.CompletionTokens, e.CostUSD, nullStr(e.Detail), floatPtrArg(e.CostNative), nullStr(e.CostCurrency), intPtrArg(e.CachedPromptTokens), - boolInt(e.Unmetered), + boolInt(e.Unmetered), nullStr(e.PriceTier), ) if err != nil { return fmt.Errorf("store: usage.record: %w", err) @@ -43,7 +43,7 @@ func (v usageView) Events(ctx context.Context, since time.Time, limit int) ([]Us `SELECT ue.ts, ue.kind, ue.model, ue.slot, ue.provider_id, rp.name, ue.prompt_tokens, ue.completion_tokens, ue.cost_usd, ue.detail, ue.cost_native, ue.cost_currency, - ue.cached_prompt_tokens, ue.unmetered + ue.cached_prompt_tokens, ue.unmetered, ue.price_tier FROM usage_events ue LEFT JOIN router_providers rp ON rp.id = ue.provider_id WHERE ue.ts >= ? ORDER BY ue.ts DESC, ue.id DESC LIMIT ?`, @@ -56,13 +56,13 @@ func (v usageView) Events(ctx context.Context, since time.Time, limit int) ([]Us for rows.Next() { var e UsageEvent var ts int64 - var model, slot, providerName, detail, costCurrency sql.NullString + var model, slot, providerName, detail, costCurrency, priceTier sql.NullString var providerID, prompt, completion, cachedPromptTokens sql.NullInt64 var cost, costNative sql.NullFloat64 var unmetered int64 if err := rows.Scan(&ts, &e.Kind, &model, &slot, &providerID, &providerName, &prompt, &completion, &cost, &detail, - &costNative, &costCurrency, &cachedPromptTokens, &unmetered); err != nil { + &costNative, &costCurrency, &cachedPromptTokens, &unmetered, &priceTier); err != nil { return nil, fmt.Errorf("store: usage.events: %w", err) } e.TS = timeOf(sql.NullInt64{Int64: ts, Valid: true}) @@ -78,6 +78,7 @@ func (v usageView) Events(ctx context.Context, since time.Time, limit int) ([]Us e.CostCurrency = strOf(costCurrency) e.CachedPromptTokens = nullInt64Ptr(cachedPromptTokens) e.Unmetered = unmetered != 0 + e.PriceTier = strOf(priceTier) out = append(out, e) } if err := rows.Err(); err != nil { diff --git a/go/internal/store/usage_test.go b/go/internal/store/usage_test.go index 4ca2188..07d6596 100644 --- a/go/internal/store/usage_test.go +++ b/go/internal/store/usage_test.go @@ -29,6 +29,7 @@ func TestUsageRecordAndEventsRoundTripNewFields(t *testing.T) { TS: now, Kind: "external_request", Model: "wire-m", ProviderID: &deepseekID, PromptTokens: 1000, CompletionTokens: 200, CostNative: &cost, CostCurrency: "USD", CachedPromptTokens: &cached, + PriceTier: "peak", }); err != nil { t.Fatalf("Record metered: %v", err) } @@ -90,6 +91,12 @@ func TestUsageRecordAndEventsRoundTripNewFields(t *testing.T) { if metered.Unmetered { t.Error("metered row: Unmetered = true, want false") } + if metered.PriceTier != "peak" { + t.Errorf("metered row PriceTier = %q, want %q", metered.PriceTier, "peak") + } + if inference.PriceTier != "" { + t.Errorf("inference row PriceTier = %q, want empty", inference.PriceTier) + } } func TestTokenActivityFiltersKindsAndWindow(t *testing.T) { diff --git a/systemd/forge-compress@.service b/systemd/forge-compress@.service index 4e876c8..82de5e0 100644 --- a/systemd/forge-compress@.service +++ b/systemd/forge-compress@.service @@ -17,13 +17,15 @@ # # The three Environment= paths below are this binary's own artifacts -- # constant across every instance on this host, so they're baked in here -# rather than written per-instance by Provisioner (which only ever varies -# OPENAI_TARGET_API_URL/HEADROOM_KOMPRESS_ONNX_INTRA_THREADS/ -# HEADROOM_PROXY_TOKEN/HEADROOM_PORT per row, same as it always has for -# headroom@ instances -- see writeEnv). HEADROOM_PROXY_TOKEN is accepted -# from that EnvironmentFile but not enforced as an auth gate by this -# binary; see cmd/forge-compress/config.go's ProxyToken doc comment for -# why. +# rather than written per-instance by Provisioner (writeEnv, +# internal/compressorctl/provision.go -- the var names in that function's +# own comment predate the Foundry->Forge rename and are stale here too; +# trust provision.go's actual code over this paragraph). As of 2026-09-11 +# writeEnv also writes COMPRESS_MAX_INFLIGHT and +# FORGE_COMPRESS_FAILOPEN_BUDGET_MS per row -- COMPRESS_PROXY_TOKEN is +# accepted from that EnvironmentFile but not enforced as an auth gate by +# this binary; see cmd/forge-compress/config.go's ProxyToken doc comment +# for why. # # One-time manual root install (this file only -- not something to script # unattended against a live production host, same convention as @@ -56,6 +58,15 @@ EnvironmentFile=/var/lib/forge/compress/%i.env Environment=FORGE_COMPRESS_MODEL_PATH=/opt/forge/compress/kompress-int8-wo.onnx Environment=FORGE_COMPRESS_TOKENIZER_PATH=/opt/forge/compress/tokenizer.json Environment=FORGE_COMPRESS_ONNXRUNTIME_LIB=/opt/forge/compress/libonnxruntime.so +# GOMEMLIMIT, added after the 2026-09-09 OOM-kill (forge-compress@deepseek, +# 1d7h uptime, kernel OOM killer under MemoryMax below). Go's GC has no +# innate awareness of a cgroup ceiling -- without this it paces itself +# against GOGC alone, oblivious to MemoryMax, and native onnxruntime +# allocations (cgo, off the Go heap) are invisible to it entirely, so the +# process can sail past the ceiling with no graceful GC response before the +# kernel hard-kills it. Set comfortably under MemoryMax so the GC has room +# to react before that happens. +Environment=GOMEMLIMIT=3500MiB # Legacy aliases -- the currently-deployed binary predates the Foundry->Forge # rename and still reads FOUNDRY_COMPRESS_*; remove these three lines once # forge-compress is rebuilt natively on this host from the renamed source. diff --git a/web/src/components/CatalogPanel.tsx b/web/src/components/CatalogPanel.tsx index 2b5a63b..c0a78eb 100644 --- a/web/src/components/CatalogPanel.tsx +++ b/web/src/components/CatalogPanel.tsx @@ -8,7 +8,7 @@ // merged-config seam picks up changes immediately. import { useEffect, useState } from "react"; -import { countryFlag, formatGB } from "../lib/format"; +import { countryFlag, formatCurrencyPrecise, formatGB } from "../lib/format"; import { familyInheritedIcon, modelInheritedIcon } from "../lib/iconInheritance"; import { preferredOfferingIds } from "../lib/offeringPreference"; import { providerIconSlug } from "../lib/providerPresets"; @@ -705,9 +705,19 @@ function OfferingsSection({ canAdmin }: { canAdmin: boolean }) { {m?.name ?? `#${o.model_id}`} - {o.price_in_per_1m}/{o.price_out_per_1m} + {formatCurrencyPrecise(o.price_in_per_1m, o.currency)}/{formatCurrencyPrecise(o.price_out_per_1m, o.currency)} {o.price_cached_in_per_1m != null && ( - (cached {o.price_cached_in_per_1m}) + (cached {formatCurrencyPrecise(o.price_cached_in_per_1m, o.currency)}) + )} + {(o.price_in_per_1m_peak != null || o.price_out_per_1m_peak != null) && ( + + {" "}(peak {formatCurrencyPrecise(o.price_in_per_1m_peak ?? o.price_in_per_1m, o.currency)}/{formatCurrencyPrecise(o.price_out_per_1m_peak ?? o.price_out_per_1m, o.currency)}) + + )} + {prov?.peakActiveNow && ( + + peak now + )} {o.priority} diff --git a/web/src/components/CompressionSavingsChips.tsx b/web/src/components/CompressionSavingsChips.tsx index 0223776..9f7820c 100644 --- a/web/src/components/CompressionSavingsChips.tsx +++ b/web/src/components/CompressionSavingsChips.tsx @@ -47,26 +47,48 @@ import type { CompressorSummaryProxy } from "../lib/types"; // model with no resolvable TPS is omitted (logged server-side as an // anomaly), which can legitimately leave this chip blank even when tokens // were cached — see local.hasTime below. +// +// 2026-09-11: this chip has shown "no data" for local traffic since the Go +// rewrite, root-caused this session — requests_cached (what time_saved_seconds_est +// above is apportioned from) is structurally 0 in production; forge-compress +// never increments compress_requests_cached_total. The real local win (up to +// 10 of 15 minutes on some tasks, per the operator's own early testing) comes +// from Compressor's own token-dropping, priced by the new +// compression_time_saved_seconds_est (apportioned from the real tokens_saved +// count, same per-model TPS math). Both mechanisms are genuinely additive — +// a cache hit and a compressed prompt are different ways the same request +// can avoid prefill work — so this chip sums both rather than picking one. function sumLocal(proxies: CompressorSummaryProxy[]) { - let tokensSaved = 0; + let tokensCachedEst = 0; + let tokensCompressed = 0; let timeSavedS = 0; - let hasTokens = false; + let hasCachedTokens = false; + let hasCompressedTokens = false; let hasTime = false; const sources = new Set(); for (const p of proxies) { if (p.kind !== "local") continue; if (p.tokens_saved_est != null) { - tokensSaved += p.tokens_saved_est; - hasTokens = true; + tokensCachedEst += p.tokens_saved_est; + hasCachedTokens = true; + } + if (p.tokens_saved > 0) { + tokensCompressed += p.tokens_saved; + hasCompressedTokens = true; } if (p.time_saved_seconds_est != null) { timeSavedS += p.time_saved_seconds_est; hasTime = true; } + if (p.compression_time_saved_seconds_est != null) { + timeSavedS += p.compression_time_saved_seconds_est; + hasTime = true; + } if (p.tps_source) sources.add(p.tps_source); } - return { tokensSaved, hasTokens, timeSavedS, hasTime, sources }; + const hasTokens = hasCachedTokens || hasCompressedTokens; + return { tokensCachedEst, hasCachedTokens, tokensCompressed, hasCompressedTokens, hasTokens, timeSavedS, hasTime, sources }; } function sumExternal(proxies: CompressorSummaryProxy[]) { @@ -122,10 +144,15 @@ export function CompressorSavingsChips({ window_ }: { window_: string }) { const external = sumExternal(proxies); const localTitle = !local.hasTokens - ? "No cached local requests this window" + ? "No cached or compressed local requests this window" : local.hasTime - ? `${formatTokens(local.tokensSaved)} tokens not re-prefilled, estimated against real prefill TPS via ${Array.from(local.sources).join(", ")}` - : "No model with cached requests this window had a real measured prefill TPS — time estimate unavailable"; + ? [ + local.hasCompressedTokens ? `${formatTokens(local.tokensCompressed)} tokens dropped by compression` : null, + local.hasCachedTokens ? `${formatTokens(local.tokensCachedEst)} tokens not re-prefilled (cache hit, estimated)` : null, + ] + .filter(Boolean) + .join(" + ") + `, avoided-prefill time estimated against real prefill TPS via ${Array.from(local.sources).join(", ")}` + : "No model with cached/compressed requests this window had a real measured prefill TPS — time estimate unavailable"; const externalTitle = !external.hasTokens ? "No compressed external requests this window" diff --git a/web/src/components/RemoteOfferings.tsx b/web/src/components/RemoteOfferings.tsx index 4242a11..ef12ff7 100644 --- a/web/src/components/RemoteOfferings.tsx +++ b/web/src/components/RemoteOfferings.tsx @@ -21,7 +21,7 @@ // enabled-only view. import { useState } from "react"; import { OfferingForm, useCatalogProviders } from "./catalog/OfferingForm"; -import { countryFlag } from "../lib/format"; +import { countryFlag, formatCurrencyPrecise } from "../lib/format"; import { groupOfferingsByModel, preferredOfferingIds } from "../lib/offeringPreference"; import { presetFor, providerIconSlug } from "../lib/providerPresets"; import { @@ -103,24 +103,14 @@ export function RemoteOfferings() { // rather than a bare {enabled} patch, or it would silently blank the rest // of the row (the exact "Full-replace curl verification hazard" from the // 2026-08-05 incident, but reachable from this button too if skipped). + // Spread the existing record rather than hand-enumerating columns (was a + // latent hazard fixed in the peak pricing sprint, 2026-09-12: a + // hand-enumerated body silently drops any field added after it was + // written — Settings → Routing's submitPatch already used this pattern). function handleToggleEnabled(o: CatalogOffering) { clearError(); - update.mutate({ - id: o.id, - o: { - model_id: o.model_id, - variant_id: o.variant_id, - provider: o.provider, - wire_model: o.wire_model, - price_in_per_1m: o.price_in_per_1m, - price_out_per_1m: o.price_out_per_1m, - price_cached_in_per_1m: o.price_cached_in_per_1m, - currency: o.currency, - context_length: o.context_length, - enabled: !o.enabled, - priority: o.priority, - }, - }, { onError: showError }); + const { id, ...rest } = o; + update.mutate({ id, o: { ...rest, enabled: !o.enabled } }, { onError: showError }); } return ( @@ -239,9 +229,17 @@ function RemoteModelRow({ {o.enabled && providerDisabled && disabled — not routing}
- {o.price_in_per_1m} {o.currency}/M in · {o.price_out_per_1m} out + + {formatCurrencyPrecise(o.price_in_per_1m, o.currency)}/M in · {formatCurrencyPrecise(o.price_out_per_1m, o.currency)} out + {(o.price_in_per_1m_peak != null || o.price_out_per_1m_peak != null) && ( + <> (peak {formatCurrencyPrecise(o.price_in_per_1m_peak ?? o.price_in_per_1m, o.currency)}/{formatCurrencyPrecise(o.price_out_per_1m_peak ?? o.price_out_per_1m, o.currency)}) + )} + {o.context_length > 0 && · {o.context_length.toLocaleString()} ctx} priority {o.priority} + {row?.peak_active_now && ( + peak now + )}
{canAdmin && ( diff --git a/web/src/components/catalog/OfferingForm.tsx b/web/src/components/catalog/OfferingForm.tsx index 57f7976..db92eae 100644 --- a/web/src/components/catalog/OfferingForm.tsx +++ b/web/src/components/catalog/OfferingForm.tsx @@ -19,6 +19,13 @@ export interface CatalogProviderRef { country: string; data_residency_group: string; enabled: boolean; + // Peak pricing sprint (2026-09-12): whether this provider has a peak + // schedule configured at all — drives OfferingForm's "no peak window + // configured" hint, since a peak price set without one never applies. + hasPeakWindows: boolean; + // Server-computed "peak in force right now" — never re-derived from the + // raw schedule client-side. + peakActiveNow: boolean; } export function useCatalogProviders(): CatalogProviderRef[] { @@ -28,6 +35,8 @@ export function useCatalogProviders(): CatalogProviderRef[] { country: p.country ?? "", data_residency_group: p.data_residency_group ?? "", enabled: p.enabled, + hasPeakWindows: !!p.peak_windows, + peakActiveNow: p.peak_active_now, })); } @@ -57,6 +66,9 @@ export function OfferingForm({ const [priceIn, setPriceIn] = useState(existing?.price_in_per_1m ?? 0); const [priceOut, setPriceOut] = useState(existing?.price_out_per_1m ?? 0); const [priceCachedIn, setPriceCachedIn] = useState(existing?.price_cached_in_per_1m ?? null); + const [priceInPeak, setPriceInPeak] = useState(existing?.price_in_per_1m_peak ?? null); + const [priceOutPeak, setPriceOutPeak] = useState(existing?.price_out_per_1m_peak ?? null); + const [priceCachedInPeak, setPriceCachedInPeak] = useState(existing?.price_cached_in_per_1m_peak ?? null); const [currency, setCurrency] = useState(existing?.currency ?? "USD"); const [contextLength, setContextLength] = useState(existing?.context_length ?? 0); const [enabled, setEnabled] = useState(existing?.enabled ?? true); @@ -75,6 +87,9 @@ export function OfferingForm({ price_in_per_1m: priceIn, price_out_per_1m: priceOut, price_cached_in_per_1m: priceCachedIn, + price_in_per_1m_peak: priceInPeak, + price_out_per_1m_peak: priceOutPeak, + price_cached_in_per_1m_peak: priceCachedInPeak, currency, context_length: contextLength, enabled, @@ -84,6 +99,12 @@ export function OfferingForm({ ); } + function fillPeakDouble() { + setPriceInPeak(priceIn * 2); + setPriceOutPeak(priceOut * 2); + if (priceCachedIn != null) setPriceCachedInPeak(priceCachedIn * 2); + } + return (
+
+
+ Time-of-day pricing (optional) + +
+ {selectedProvider && !selectedProvider.hasPeakWindows && (priceInPeak != null || priceOutPeak != null || priceCachedInPeak != null) && ( +
+ {selectedProvider.name} has no peak schedule configured (Settings → Providers) — these peak rates will never apply. +
+ )} + {((priceInPeak != null && priceInPeak < priceIn) || (priceOutPeak != null && priceOutPeak < priceOut)) && ( +
+ A peak rate is lower than its off-peak rate — unusual, but the provider's own pricing policy isn't enforced here. +
+ )} +
+ + + diff --git a/web/src/lib/queries.ts b/web/src/lib/queries.ts index 4d6bf8f..e277fd0 100644 --- a/web/src/lib/queries.ts +++ b/web/src/lib/queries.ts @@ -1232,6 +1232,12 @@ function useCatalogMutations() { qc.invalidateQueries({ queryKey: qk.status }); qc.invalidateQueries({ queryKey: qk.modelCards("7d") }); qc.invalidateQueries({ queryKey: qk.configCards("7d") }); + // Peak pricing sprint (2026-09-12): an offering/provider price edit + // previously left the Dashboard Cost tab and the provider price line + // stale until their own unrelated refetch — a catalog edit changes what + // those surfaces should show right now. + qc.invalidateQueries({ queryKey: ["cost"] }); + qc.invalidateQueries({ queryKey: qk.providers }); }; return { invalidateAll }; } diff --git a/web/src/lib/types.ts b/web/src/lib/types.ts index a349d64..376bb1e 100644 --- a/web/src/lib/types.ts +++ b/web/src/lib/types.ts @@ -893,6 +893,13 @@ export interface Provider { // Provider-level data residency ("" / undefined = unknown). country?: string; data_residency_group?: string; + // Peak pricing sprint (2026-09-12): peak_windows is the raw JSON-encoded + // schedule (edited as text — see PeakWindowsEditor), undefined/"" = no + // peak/off-peak concept for this provider. peak_active_now is computed + // server-side against the real clock — never re-derive the window math + // in the frontend. + peak_windows?: string; + peak_active_now: boolean; } export interface ProvidersResponse { @@ -955,6 +962,10 @@ export interface ProviderUpdateRequest { enabled?: boolean; country?: string; data_residency_group?: string; + // Peak pricing sprint (2026-09-12): raw JSON-encoded internal/pricing. + // Windows schedule; "" clears it. Validated server-side (PeakWindowsEditor + // surfaces the parse error rather than re-implementing the grammar here). + peak_windows?: string; } // POST /api/v1/providers/{name}/discover-billing response (product/QA @@ -1486,6 +1497,14 @@ export interface CatalogOffering { // routing sprint, 2026-08-06): LOWEST value wins; default 100 = "no // preference" (ties break by provider name). priority: number; + // Peak-tier rates (peak pricing sprint, 2026-09-12): undefined/null on + // any of the three = no peak differential for that field, falls back to + // the base rate above — same precedent as price_cached_in_per_1m. + // Meaningless (never applied) unless the owning provider has a + // peak_windows schedule configured — see CatalogProviderRef. + price_in_per_1m_peak?: number | null; + price_out_per_1m_peak?: number | null; + price_cached_in_per_1m_peak?: number | null; } export interface CatalogBenchmark { @@ -1667,16 +1686,31 @@ export interface CompressorSummaryProxy { // money_saved_est FX-converted to the response's display_currency // (2026-07-31 fix — this endpoint used to never convert). money_saved_display?: number; - // tps_source/tps_mode describe the SINGLE LARGEST contributor to - // time_saved_seconds_est (by share of cached requests this window) — see - // prefill_breakdown for the full per-model accounting. Every source here - // is a real measurement; "fallback" no longer exists. + + // compression_time_saved_seconds_est is the LOCAL analogue driven by + // Compressor's own token-dropping (tokens_saved) rather than a + // prompt-cache hit (requests_cached, above). Added 2026-09-11: + // requests_cached has been structurally 0 in production since the Go + // rewrite (forge-compress never increments + // compress_requests_cached_total), so time_saved_seconds_est has never + // actually fired — this is the field that does. Same per-model + // apportionment/TPS lookups as time_saved_seconds_est (they share + // prefill_breakdown/tps_source/tps_mode below). + compression_time_saved_seconds_est?: number; + compression_money_saved_est?: number; + compression_money_saved_currency?: string; + compression_money_saved_display?: number; + + // tps_source/tps_mode describe the SINGLE LARGEST contributor (by share + // of this window's requests) to EITHER estimate above — the apportionment + // is identical, so the biggest-contributor model is the same for both. + // Every source here is a real measurement; "fallback" no longer exists. tps_source?: "profile_depth_curve" | "observed" | "profile_scalar" | "catalog" | "live"; tps_mode?: string; - // prefill_breakdown is the full per-model accounting behind - // time_saved_seconds_est: every model that contributed cached requests - // this window AND had a resolvable real prefill TPS, sorted by share - // descending. Local-only. + // prefill_breakdown is the full per-model accounting shared by both + // time_saved_seconds_est and compression_time_saved_seconds_est: every + // model that contributed requests this window AND had a resolvable real + // prefill TPS, sorted by share descending. Local-only. prefill_breakdown?: { mode: string; share: number; // 0-1, this model's share of the proxy's requests this window diff --git a/web/src/settings/panels/ProviderKeys.tsx b/web/src/settings/panels/ProviderKeys.tsx index a92edbc..1fcf875 100644 --- a/web/src/settings/panels/ProviderKeys.tsx +++ b/web/src/settings/panels/ProviderKeys.tsx @@ -352,6 +352,11 @@ export function ProviderKeys({ canAdmin }: { canAdmin: boolean }) { /> Billing API enabled + setEditDraft({ ...editDraft, peak_windows: v })} + />
submitEdit(p.id)} /> @@ -456,3 +461,178 @@ export function ProviderKeys({ canAdmin }: { canAdmin: boolean }) { ); } + +const PEAK_WINDOWS_JSON_PLACEHOLDER = `[{"days":[1,2,3,4,5],"start":"01:00","end":"04:00"}]`; + +const WEEKDAY_SHORT = ["Sun", "Mon", "Tue", "Wed", "Thu", "Fri", "Sat"]; + +// A short curated list, not the full ~600-name IANA database — common +// zones a provider is actually likely to publish hours in, covering every +// populated continent, plus UTC itself. "Custom…" is the escape hatch for +// anything else (the free-text input below still accepts any valid IANA +// name typed directly — this select is a convenience, not a restriction). +const COMMON_TIMEZONES = [ + "UTC", + "America/Los_Angeles", + "America/Denver", + "America/Chicago", + "America/New_York", + "America/Sao_Paulo", + "Europe/London", + "Europe/Paris", + "Europe/Moscow", + "Africa/Cairo", + "Africa/Johannesburg", + "Asia/Dubai", + "Asia/Kolkata", + "Asia/Shanghai", + "Asia/Tokyo", + "Asia/Seoul", + "Australia/Sydney", + "Pacific/Auckland", +]; + +type PeakWindowRow = { days: number[]; start: string; end: string }; +type PeakSchedule = { tz?: string; windows?: PeakWindowRow[] }; + +// splitStored parses the full stored peak_windows JSON (as PUT to the API) +// into its two independently-edited pieces: the timezone name and the +// windows array (kept as its own JSON text so the textarea round-trips +// exactly what the operator typed, including in-progress edits). Falls +// back to (no tz, raw text in the windows box) when the stored value isn't +// valid JSON — should only happen for hand-typed data mid-edit, since +// every value that reaches here from the server already parsed once. +function splitStored(raw: string): { tz: string; windowsJson: string } { + if (raw.trim() === "") return { tz: "", windowsJson: "" }; + try { + const parsed = JSON.parse(raw) as PeakSchedule; + return { tz: parsed.tz ?? "", windowsJson: JSON.stringify(parsed.windows ?? []) }; + } catch { + return { tz: "", windowsJson: raw }; + } +} + +// composeStored is splitStored's inverse — called on every keystroke in +// either sub-field to rebuild the single JSON string the API/store expects. +// Omits "tz" entirely when empty (equivalent to UTC, matches the backend's +// omitempty convention) rather than writing {"tz":""}. +function composeStored(tz: string, windowsJson: string): string { + if (windowsJson.trim() === "" && tz.trim() === "") return ""; + let windows: unknown = []; + try { + windows = JSON.parse(windowsJson || "[]"); + } catch { + // Malformed mid-edit — still compose so `tz` isn't lost, but embed the + // raw (invalid) text as a placeholder the server will reject with a + // real parse error surfaced through the existing error banner. + return `{${tz ? `"tz":${JSON.stringify(tz)},` : ""}"windows":${windowsJson || "[]"}}`; + } + const schedule: PeakSchedule = { windows: windows as PeakWindowRow[] }; + if (tz.trim() !== "") schedule.tz = tz.trim(); + return JSON.stringify(schedule); +} + +// PeakWindowsEditor (peak pricing sprint, 2026-09-12; timezone input added +// 2026-09-13) — a labeled Timezone picker plus a validated JSON textarea +// for a provider's recurring peak-hours windows (e.g. DeepSeek's weekday +// 01:00-04:00 + 06:00-10:00 UTC). The windows LIST stays JSON — there's no +// precedent anywhere in this codebase for a bespoke hour-range picker (the +// closest is SchedulerJobs' raw cron text input) and it's edited rarely — +// but the timezone is a single flat value a provider almost always quotes +// in one local zone ("9am-5pm Pacific"), and typing "tz":"..." correctly +// inside hand-written JSON is exactly the kind of error a labeled field +// with a curated dropdown avoids. The server converts using the real IANA +// database (internal/pricing), so entering local hours + a zone here is +// enough — no manual UTC math, and it stays correct across DST automatically. +// The summary below only reformats the parsed JSON for readability — it +// never evaluates whether a window is currently active; that's +// peakActiveNow, computed server-side and passed in, never re-derived here. +function PeakWindowsEditor({ + value, + peakActiveNow, + onChange, +}: { + value: string; + peakActiveNow: boolean; + onChange: (v: string) => void; +}) { + const [tz, setTz] = useState(() => splitStored(value).tz); + const [windowsJson, setWindowsJson] = useState(() => splitStored(value).windowsJson); + const [tzMode, setTzMode] = useState<"select" | "custom">(() => + COMMON_TIMEZONES.includes(splitStored(value).tz) || splitStored(value).tz === "" ? "select" : "custom", + ); + + function update(nextTz: string, nextWindowsJson: string) { + setTz(nextTz); + setWindowsJson(nextWindowsJson); + onChange(composeStored(nextTz, nextWindowsJson)); + } + + let summary: string | null = null; + let localParseError: string | null = null; + if (windowsJson.trim() !== "") { + try { + const parsed = JSON.parse(windowsJson) as PeakWindowRow[]; + const tzLabel = tz.trim() || "UTC"; + summary = parsed + .map((w) => `${w.days.map((d) => WEEKDAY_SHORT[d] ?? `?${d}`).join("/")} ${w.start}–${w.end} ${tzLabel}`) + .join(", ") || "(no windows)"; + } catch { + localParseError = "Not valid JSON — the server will reject this until it parses."; + } + } + + return ( +
+ +