Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
85 changes: 85 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,91 @@ and this project uses [Semantic Versioning](https://semver.org/spec/v2.0.0.html)

### Fixed

- The cost ledger no longer fabricates a $0.00 price for an unpriced
provider/model. `PriceBook.compute_cost()` now returns a
`(cost_amount, currency_code, price_known)` 3-tuple; `UsageRecord` gains a
`price_known: bool` field (persisted through a new `usage_price_knowledge`
satellite table, joined the same way `usage_measurements` already is) so a
measured-but-unpriced request is distinguishable from a genuinely free one.
`CostLedger.rollup()`/`.report()`/`.total()` add an additive
`cost_amount_by_status`/`record_count_by_status` breakdown
(measured/estimated/unavailable) alongside the existing flat `cost_amount`
total, so measured, estimated, and unavailable-priced spend are no longer
opaquely blended into one authoritative-looking number.
Comment thread
seonghobae marked this conversation as resolved.
- `CostLedger.rollup()`/`.report()` add the same treatment one level further:
an additive `cost_amount_by_price_status`/`record_count_by_price_status`
breakdown (`known`/`unknown`) alongside the measured/estimated/unavailable
one, so an unpriced request's spend is visible even after rolling many
records up into one bucket. `cheapest_upstream()` no longer treats an
unpriced candidate as free (cost `0`) when selecting the lowest-cost
upstream — an unknown price is excluded from the comparison entirely
rather than winning it by default; `None` is returned when no candidate
has a known price.
- (Devin review on #956) `SqlLedgerStore._append_locked()`'s satellite
writes (`usage_measurements`, `usage_price_knowledge`, attribution) now
only run when the parent `llm_usage_records` insert is actually accepted.
A retried `usage_record_id` whose parent insert is correctly rejected as
a duplicate previously still ran these satellite inserts unconditionally
— for a parent row that predates one of these tables (e.g. an
upgrade-migrated row with no `usage_price_knowledge` child, intentionally
read as price-unknown) that silently backfilled its provenance from the
retry's current price/measurement state instead of what actually priced
the original spend. `append()` is now a true no-op on a rejected
duplicate.
- `price_known` now propagates from the ledger into every downstream usage
surface: `CostRoutingCoordinator.complete()`'s sync/provider-request cost
dicts and their per-currency components, `record_stream_usage()`,
`retrieve_batch()`'s per-item results, and `embeddings_batch_document()`.
An unpriced request's `cost_amount` is `null` rather than a silent `0`
wherever it surfaces, not just in the ledger's own rollups.
`PriceBook.compute_cost()` treats zero token usage as the one exception:
zero tokens cost zero regardless of whether the provider/model's price is
known (zero times any finite price is still zero), so a cache hit — whose
synthetic `("cache", "response")` provider/model never has a price row —
is always `price_known=True`, `cost_amount=0.0` rather than being wrongly
reported as an unpriced request.
- `retrieve_batch()`'s `item.prompt_tokens or None` treated a legitimately
reported zero token count the same as a missing one (Python's falsy-zero),
so a genuinely confirmed zero-usage batch item was silently downgraded to
an unmeasured estimate instead of staying a real, priced `measured` `0`.
`PgLlmBatchBackend.retrieve()` and `BatchResultItem` now carry an explicit
`usage_valid` tri-state (confirmed-typed non-negative counts vs. missing/
malformed usage), and `retrieve_batch()` reads that instead of relying on
truthiness. That same estimation fallback also passed a hardcoded
empty-content placeholder instead of the request actually submitted, so a
large prompt whose provider marked usage invalid was undercounted to
near-zero prompt tokens — understating batch cost. (CodeRabbit review)
`submit_batch()` now computes a real prompt-token estimate per
`custom_id` before submission and stores it as a `prompt_token_estimates`
field on the durable `BatchJob` record itself, rather than the raw
submitted messages (Devin review: a batch registry can be Valkey-backed
with a multi-day retention shared across processes, and a submitted
prompt may be ZDR-flagged or otherwise sensitive — a token count carries
no reconstructable prompt content). Publishing it on the existing
`BatchJob` write also means an accepted job still has exactly one
publication write, so a metadata-only estimate can never orphan an
already-accepted (and possibly billed) backend job behind a raised
exception with no job id ever returned to the caller. `retrieve_batch()`
reads that stored estimate instead of falling back to an empty
placeholder. (Devin review) A job accepted before this fix has no
`prompt_token_estimates` at all, and its original request never lived
anywhere durable that a post-fix retrieval could still read — so
`retrieve_batch()` now also reads a legacy, pre-fix `batch_requests`
registry entry (still populated only for jobs submitted before this
change; nothing writes new entries there anymore) whenever a custom_id
has no stored estimate, computes the real prompt-token count from it the
same way a fresh submission would have, and persists that estimate back
onto the job so a re-retrieval never repeats the lookup. That legacy
lookup initially gated on whether `job.prompt_token_estimates` was
non-empty at all — wrong once any one custom_id had already picked up an
estimate (including from an earlier partial retrieval of the same job),
since every other still-unestimated custom_id would then silently stop
being looked up for the rest of that job's lifetime. It now gates on
whether the current retrieval actually has an item that still needs the
legacy lookup, so a legacy job's estimates can be filled in correctly
across as many partial retrievals as it takes; a job that has never
needed the legacy path (everything submitted after this fix) still never
touches that registry.
- `CostRoutingCoordinator._record_race_endpoint_usage()` no longer silently
drops a completed, billable race-loser call's spend when its usage payload
can't be parsed. It now writes a `measurement_status="unavailable"` ledger
Expand Down
25 changes: 20 additions & 5 deletions contextual_orchestrator/batch_routing.py
Original file line number Diff line number Diff line change
Expand Up @@ -131,8 +131,8 @@ def cheapest_upstream(
Cost-optimising upstream selection for load balancing: given candidate
provider/model pairs, price each against the configurable price table for a
representative request shape and return the cheapest. Unpriced candidates
cost ``0`` and are treated as free (explicit, so a missing price is visible
rather than silently expensive). Ties keep input order.
are excluded because an unknown price is not free. Ties keep input order;
``None`` is returned when no candidate has a known price.
"""
if not candidates:
return None
Expand All @@ -141,9 +141,11 @@ def cheapest_upstream(
for candidate in candidates:
provider = candidate.get("provider", "")
model = candidate.get("model", "")
cost, _currency = price_book.compute_cost(
cost, _currency, price_known = price_book.compute_cost(
provider, model, assumed_prompt_tokens, assumed_completion_tokens
)
if not price_known:
continue
if best_cost is None or cost < best_cost:
best_cost = cost
best = candidate
Expand Down Expand Up @@ -190,6 +192,9 @@ class BatchJob:
# HTTP callers bind this opaque digest to the authenticated principal;
# library-only jobs may remain unowned for standalone use.
owner_id: Optional[str] = None
# Prompt-token fallback estimates are safe metadata, stored atomically with
# the job handle rather than retaining submitted prompt text.
prompt_token_estimates: Dict[str, int] = field(default_factory=dict)


@dataclass
Expand All @@ -203,6 +208,7 @@ class BatchResultItem:
attribution: Dict[str, Any] = field(default_factory=dict)
model: str = "contextual-orchestrator"
mode: str = "auto"
usage_valid: Optional[bool] = None


class BatchBackend(Protocol):
Expand Down Expand Up @@ -403,17 +409,26 @@ async def _download() -> Dict[str, Any]:
body = (entry.get("response") or {}).get("body", {})
answer = _extract_answer(body)
usage = body.get("usage", {}) or {}
raw_prompt_tokens = usage.get("prompt_tokens")
raw_completion_tokens = usage.get("completion_tokens")
usage_valid = (
type(raw_prompt_tokens) is int
and raw_prompt_tokens >= 0
and type(raw_completion_tokens) is int
and raw_completion_tokens >= 0
)
raw_request = tracked.get(custom_id)
request = BatchRequest(**raw_request) if raw_request else None
items.append(
BatchResultItem(
custom_id=custom_id,
answer=answer,
prompt_tokens=int(usage.get("prompt_tokens", 0)),
completion_tokens=int(usage.get("completion_tokens", 0)),
prompt_tokens=raw_prompt_tokens if usage_valid else 0,
completion_tokens=raw_completion_tokens if usage_valid else 0,
attribution=dict(request.attribution) if request else {},
model=request.model if request else "contextual-orchestrator",
mode=request.mode if request else "auto",
usage_valid=usage_valid,
)
)
return items
Expand Down
Loading
Loading