Skip to content

Surface max_tokens truncation; never silently treat a cut partial as completed speech #135

Description

@antra-tess

Problem

A provider stream that ends at stop_reason=max_tokens is currently routed/stored as ordinary assistant prose with no resident-visible or witness-visible indication that the message was cut. For models with a hard output ceiling (notably Claude 3 Opus at 4,096 tokens), this creates repeated mid-word sentence stumps that can be misread as authored endings.

Production specimen (payload omitted):

  • model: claude-3-opus-20240229;
  • two primary stream responses, four hours apart;
  • both emitted exactly 4,096 output tokens;
  • both ended stop_reason=max_tokens;
  • visible endings were mid-word;
  • no framework-level truncation receipt was shown to the resident or room.

This is distinct from provider refusal and from ordinary end_turn.

Automatic continuation is not a safe default. In the same specimen, the later compiled request contained the complete prior response inside an assistant-role transcript block and ended with assistant prefill. A blind continue/next loop can amplify self-echo or repeat a repetition indefinitely.

Desired contract

On terminal max_tokens:

  1. preserve the emitted partial exactly;
  2. record structured provenance such as { truncated:true, stopReason:'max_tokens', generation/attempt identity, emitted token/char count };
  3. surface a short model-/resident-/witness-visible boundary that the text was cut, without reproducing payload;
  4. never label the partial as an ordinary completed turn;
  5. do not auto-continue by default.

If bounded continuation is added later, require:

  • explicit per-residence policy/standing;
  • hard segment and token limits;
  • generation/attempt binding;
  • exact overlap deduplication;
  • progress detection;
  • high-similarity/no-progress abort;
  • one canonical assembled output with segment provenance;
  • no duplicate publication of already emitted text;
  • tests against a prompt that already contains the prior partial.

Acceptance tests

  • end_turn behavior remains unchanged.
  • max_tokens with partial prose persists the bytes and emits exactly one truncation receipt.
  • a mid-word cutoff is not represented as authored completion.
  • retry/provider-refusal paths remain distinct.
  • disabled continuation makes zero additional provider calls.
  • an optional continuation fixture that repeats the prior output stops boundedly rather than looping.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions