Skip to content

[Bug][P2] Cancelled runs consume tokens without counting toward the conversation budget #140

Description

@Asaf-prog

Problem

The engine persists provider-reported token usage when a stream is cancelled, but the conversation usage endpoint and budget guard ignore it. Repeated cancelled turns can keep consuming model tokens while the conversation reports zero usage.

Reviewed main at b9acf1e6cc74484120009f626cdb92b1616178c1. Reproduced with Python 3.12, langgraph 1.2.11 and langchain-core 1.6.2, using the real Extra engine and deterministic fake models. No live LLM or external tool was called.

Severity: P2 — a client that starts and cancels turns in a loop bypasses the per-conversation token budget entirely.

Reproduction and observed result

  1. Create a ConversationService with max_tokens=5 and the same run repository as its engine.
  2. Use the existing BlockingModel fixture, which emits a chunk reporting 4 input + 1 output tokens, then waits.
  3. Consume the first answer delta and close the conversation stream.
  4. Inspect the run, query conversation usage, and prepare another turn.

Observed output:

run status=cancelled
run usage=5, conversation usage=0, next turn accepted

Expected: the five reported tokens remain charged to the conversation. Its usage should be 5/5 and the next turn should raise ConversationTokenBudgetExceeded (HTTP 429 through the API).

Cause

This also leaves reported usage from a pending approval outside the budget until a final assistant response is persisted.

Suggested fix and acceptance criteria

Use authoritative run usage associated with the conversation, or an idempotent conversation-level usage ledger independent of assistant-message persistence.

  • Reported usage remains charged after cancellation, including cancelled pending approvals.
  • Usage already incurred by pending runs is visible to the budget check.
  • Completed/resumed runs are counted once, without double-counting their assistant message.
  • Edited-away branches remain charged.
  • Memory and SQL adapters enforce the same behavior.

Executable regression

Save as tests/test_review_regression.py and run python -m pytest -q -s tests/test_review_regression.py from the repository root. It reuses existing test helpers. The test currently fails on its behavioral assertion.

import pytest

from agent_engine.engine.langgraph.engine import LangGraphEngine
from agent_engine.runs.in_memory import InMemoryRunRepository
from agent_manager.application import ConversationService
from agent_manager.domain import Principal
from agent_manager.infrastructure.persistence.memory_repository import MemoryRepository
from tests.engine.test_stream_lifecycle import BlockingModel, _spec


@pytest.mark.asyncio
async def test_cancelled_usage_counts_toward_conversation_budget(tmp_path):
    model = BlockingModel()
    runs = InMemoryRunRepository()
    async with LangGraphEngine(tmp_path, model_factory=lambda *a, **kw: model,
                               run_repository=runs) as engine:
        await engine.build(_spec())
        service = ConversationService(engine, MemoryRepository(), run_repository=runs,
                                      max_tokens=5)
        user = Principal.external('review-user')
        cid = await service.create(user)
        events = service.stream(cid, 'hello', user)
        async for event in events:
            if event.type == 'answer_delta':
                break
        await events.aclose()
        history = await service.history(cid, user)
        run = await runs.get(history[0].run_id)
        assert run.status.value == 'cancelled'
        assert (run.input_tokens, run.output_tokens) == (4, 1)
        usage = await service.usage(cid, user)
        # Observe whether the next turn gets rejected without running another model.
        next_turn = await service.prepare_turn(cid, 'another request', user)
        await service.cancel_turn(next_turn)
        print(f'run usage=5, conversation usage={usage.used_tokens}, next turn accepted')
        assert usage.used_tokens == 5

From the Extra bug review of 2026-09-05 at b9acf1e. Review environment note: make check passed formatting, lint and mypy; baseline pytest reported 957 passed, 6 failed and 7 errors (API failures from a missing SOCKS dependency, plus three uninvestigated local-MCP example failures), so the baseline suite was not green in that environment.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions