Problem
The engine persists provider-reported token usage when a stream is cancelled, but the conversation usage endpoint and budget guard ignore it. Repeated cancelled turns can keep consuming model tokens while the conversation reports zero usage.
Reviewed main at b9acf1e6cc74484120009f626cdb92b1616178c1. Reproduced with Python 3.12, langgraph 1.2.11 and langchain-core 1.6.2, using the real Extra engine and deterministic fake models. No live LLM or external tool was called.
Severity: P2 — a client that starts and cancels turns in a loop bypasses the per-conversation token budget entirely.
Reproduction and observed result
- Create a
ConversationService with max_tokens=5 and the same run repository as its engine.
- Use the existing
BlockingModel fixture, which emits a chunk reporting 4 input + 1 output tokens, then waits.
- Consume the first answer delta and close the conversation stream.
- Inspect the run, query conversation usage, and prepare another turn.
Observed output:
run status=cancelled
run usage=5, conversation usage=0, next turn accepted
Expected: the five reported tokens remain charged to the conversation. Its usage should be 5/5 and the next turn should raise ConversationTokenBudgetExceeded (HTTP 429 through the API).
Cause
This also leaves reported usage from a pending approval outside the budget until a final assistant response is persisted.
Suggested fix and acceptance criteria
Use authoritative run usage associated with the conversation, or an idempotent conversation-level usage ledger independent of assistant-message persistence.
Executable regression
Save as tests/test_review_regression.py and run python -m pytest -q -s tests/test_review_regression.py from the repository root. It reuses existing test helpers. The test currently fails on its behavioral assertion.
import pytest
from agent_engine.engine.langgraph.engine import LangGraphEngine
from agent_engine.runs.in_memory import InMemoryRunRepository
from agent_manager.application import ConversationService
from agent_manager.domain import Principal
from agent_manager.infrastructure.persistence.memory_repository import MemoryRepository
from tests.engine.test_stream_lifecycle import BlockingModel, _spec
@pytest.mark.asyncio
async def test_cancelled_usage_counts_toward_conversation_budget(tmp_path):
model = BlockingModel()
runs = InMemoryRunRepository()
async with LangGraphEngine(tmp_path, model_factory=lambda *a, **kw: model,
run_repository=runs) as engine:
await engine.build(_spec())
service = ConversationService(engine, MemoryRepository(), run_repository=runs,
max_tokens=5)
user = Principal.external('review-user')
cid = await service.create(user)
events = service.stream(cid, 'hello', user)
async for event in events:
if event.type == 'answer_delta':
break
await events.aclose()
history = await service.history(cid, user)
run = await runs.get(history[0].run_id)
assert run.status.value == 'cancelled'
assert (run.input_tokens, run.output_tokens) == (4, 1)
usage = await service.usage(cid, user)
# Observe whether the next turn gets rejected without running another model.
next_turn = await service.prepare_turn(cid, 'another request', user)
await service.cancel_turn(next_turn)
print(f'run usage=5, conversation usage={usage.used_tokens}, next turn accepted')
assert usage.used_tokens == 5
From the Extra bug review of 2026-09-05 at b9acf1e. Review environment note: make check passed formatting, lint and mypy; baseline pytest reported 957 passed, 6 failed and 7 errors (API failures from a missing SOCKS dependency, plus three uninvestigated local-MCP example failures), so the baseline suite was not green in that environment.
Problem
The engine persists provider-reported token usage when a stream is cancelled, but the conversation usage endpoint and budget guard ignore it. Repeated cancelled turns can keep consuming model tokens while the conversation reports zero usage.
Reviewed
mainatb9acf1e6cc74484120009f626cdb92b1616178c1. Reproduced with Python 3.12, langgraph 1.2.11 and langchain-core 1.6.2, using the real Extra engine and deterministic fake models. No live LLM or external tool was called.Severity: P2 — a client that starts and cancels turns in a loop bypasses the per-conversation token budget entirely.
Reproduction and observed result
ConversationServicewithmax_tokens=5and the same run repository as its engine.BlockingModelfixture, which emits a chunk reporting 4 input + 1 output tokens, then waits.Observed output:
Expected: the five reported tokens remain charged to the conversation. Its usage should be 5/5 and the next turn should raise
ConversationTokenBudgetExceeded(HTTP 429 through the API).Cause
ConversationService.usageandprepare_turnconsult only the conversation repository'sget_token_usage.MemoryRepository.get_token_usagesums tokens on persisted messages. A cancelled turn has no assistant message holding those tokens.SqlRepository.get_token_usageuses the same message-only accounting. The executable reproduction uses memory; the SQL exposure is supported by code inspection.This also leaves reported usage from a pending approval outside the budget until a final assistant response is persisted.
Suggested fix and acceptance criteria
Use authoritative run usage associated with the conversation, or an idempotent conversation-level usage ledger independent of assistant-message persistence.
Executable regression
Save as
tests/test_review_regression.pyand runpython -m pytest -q -s tests/test_review_regression.pyfrom the repository root. It reuses existing test helpers. The test currently fails on its behavioral assertion.From the Extra bug review of 2026-09-05 at
b9acf1e. Review environment note:make checkpassed formatting, lint and mypy; baseline pytest reported 957 passed, 6 failed and 7 errors (API failures from a missing SOCKS dependency, plus three uninvestigated local-MCP example failures), so the baseline suite was not green in that environment.