Skip to content

[Bug] llama.cpp: "Cannot have 2 or more assistant messages at the end of the list" — stuck during compaction with tool calls disabled #367

Description

@mossaab

Description

During a long session using the llama.cpp backend (Qwen3.8-27B), the request returns:

{"error":{"code":400,"message":"Cannot have 2 or more assistant messages at the end of the list.","type":"invalid_request_error"}}

OpenFox then injects a compaction <system-reminder> ("Tool calls are not possible at this stage. STOP and produce a summary...") and the session gets stuck: tool calls are disabled, and since llama.cpp already rejects the message list, compaction cannot complete → corrupted state with no way to continue in that session.

Hypothesis

The agentic loop produces "assistant" turns containing only tool calls (no text) → serialized as consecutive / trailing empty assistant messages. The Qwen template in llama.cpp is stricter than vLLM/sglang/OpenAI about this structure and rejects the request. In my case, a reviewer sub-agent dumping full git diffs pushed context close to the limit, triggering compaction right when the message list was already invalid.

Architecture / call chain

OpenFox → LiteLLM router (:PORT) → llama-server (:PORT) → Qwen3.8-27B

The 400 is returned by llama-server (OpenAI-compatible error format) and passed through by LiteLLM unchanged — verified via [curl isolation test / LiteLLM logs: ...].

Note: with the router in between, OpenFox may detect the backend as unknown
(see #258). Reproduced both directly against llama-server [yes/no] and via the
LiteLLM router. If it only reproduces behind LiteLLM, that suggests its tool-call
transformation turns empty assistant (tool-only) turns into bare consecutive
assistant messages.

Steps to reproduce

  1. Backend: llama-server (Qwen3.x), provider set to llamacpp in OpenFox (via LiteLLM router).
  2. Session with many tool calls in a loop (e.g., a reviewer sub-agent dumping full diffs of ~40 files, 1686 insertions).
  3. Context approaching the window limit → compaction triggers + 400 error above.

Environment

  • OpenFox: v2.0.149
  • Backend: llama.cpp b11007-mix-3e83366, model Qwen3.8-27B, effective context ~70k tokens (note whether --jinja is enabled)
  • LiteLLM router in front of llama-server — version: v1.100.0, registered model name: qwen3.8-27b (litellm_params.model: openai/... )
  • OS: linux x64, Shell: zsh

Workaround that worked

Switched to a different model with 256k context (via the same setup). The larger window avoided hitting the overflow/compaction path, so the session no longer breaks. This suggests the bug is most likely triggered when compaction engages on an already-invalid trailing message sequence — i.e., it's the combination of "context near limit" + "trailing empty assistant turns", not one alone.

Expected behavior

The message list should be normalized before each chat-completions call (no ≥2 consecutive trailing assistant messages; generation always starts from a user/tool turn). At minimum, the llama.cpp structure/overflow error should trigger a clean compaction instead of getting stuck with tools disabled.

Related issues

Related to #151, #211, #334.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions