Description
During a long session using the llama.cpp backend (Qwen3.8-27B), the request returns:
{"error":{"code":400,"message":"Cannot have 2 or more assistant messages at the end of the list.","type":"invalid_request_error"}}
OpenFox then injects a compaction <system-reminder> ("Tool calls are not possible at this stage. STOP and produce a summary...") and the session gets stuck: tool calls are disabled, and since llama.cpp already rejects the message list, compaction cannot complete → corrupted state with no way to continue in that session.
Hypothesis
The agentic loop produces "assistant" turns containing only tool calls (no text) → serialized as consecutive / trailing empty assistant messages. The Qwen template in llama.cpp is stricter than vLLM/sglang/OpenAI about this structure and rejects the request. In my case, a reviewer sub-agent dumping full git diffs pushed context close to the limit, triggering compaction right when the message list was already invalid.
Architecture / call chain
OpenFox → LiteLLM router (:PORT) → llama-server (:PORT) → Qwen3.8-27B
The 400 is returned by llama-server (OpenAI-compatible error format) and passed through by LiteLLM unchanged — verified via [curl isolation test / LiteLLM logs: ...].
Note: with the router in between, OpenFox may detect the backend as unknown
(see #258). Reproduced both directly against llama-server [yes/no] and via the
LiteLLM router. If it only reproduces behind LiteLLM, that suggests its tool-call
transformation turns empty assistant (tool-only) turns into bare consecutive
assistant messages.
Steps to reproduce
- Backend: llama-server (Qwen3.x), provider set to llamacpp in OpenFox (via LiteLLM router).
- Session with many tool calls in a loop (e.g., a reviewer sub-agent dumping full diffs of ~40 files, 1686 insertions).
- Context approaching the window limit → compaction triggers + 400 error above.
Environment
- OpenFox: v2.0.149
- Backend: llama.cpp b11007-mix-3e83366, model Qwen3.8-27B, effective context ~70k tokens (note whether
--jinja is enabled)
- LiteLLM router in front of llama-server — version: v1.100.0, registered model name: qwen3.8-27b (
litellm_params.model: openai/... )
- OS: linux x64, Shell: zsh
Workaround that worked
Switched to a different model with 256k context (via the same setup). The larger window avoided hitting the overflow/compaction path, so the session no longer breaks. This suggests the bug is most likely triggered when compaction engages on an already-invalid trailing message sequence — i.e., it's the combination of "context near limit" + "trailing empty assistant turns", not one alone.
Expected behavior
The message list should be normalized before each chat-completions call (no ≥2 consecutive trailing assistant messages; generation always starts from a user/tool turn). At minimum, the llama.cpp structure/overflow error should trigger a clean compaction instead of getting stuck with tools disabled.
Related issues
Related to #151, #211, #334.
Description
During a long session using the llama.cpp backend (Qwen3.8-27B), the request returns:
OpenFox then injects a compaction
<system-reminder>("Tool calls are not possible at this stage. STOP and produce a summary...") and the session gets stuck: tool calls are disabled, and since llama.cpp already rejects the message list, compaction cannot complete → corrupted state with no way to continue in that session.Hypothesis
The agentic loop produces "assistant" turns containing only tool calls (no text) → serialized as consecutive / trailing empty
assistantmessages. The Qwen template in llama.cpp is stricter than vLLM/sglang/OpenAI about this structure and rejects the request. In my case, a reviewer sub-agent dumping full git diffs pushed context close to the limit, triggering compaction right when the message list was already invalid.Architecture / call chain
OpenFox → LiteLLM router (:PORT) → llama-server (:PORT) → Qwen3.8-27B
The 400 is returned by llama-server (OpenAI-compatible error format) and passed through by LiteLLM unchanged — verified via [curl isolation test / LiteLLM logs: ...].
Note: with the router in between, OpenFox may detect the backend as
unknown(see #258). Reproduced both directly against llama-server [yes/no] and via the
LiteLLM router. If it only reproduces behind LiteLLM, that suggests its tool-call
transformation turns empty assistant (tool-only) turns into bare consecutive
assistant messages.
Steps to reproduce
Environment
--jinjais enabled)litellm_params.model: openai/... )Workaround that worked
Switched to a different model with 256k context (via the same setup). The larger window avoided hitting the overflow/compaction path, so the session no longer breaks. This suggests the bug is most likely triggered when compaction engages on an already-invalid trailing message sequence — i.e., it's the combination of "context near limit" + "trailing empty assistant turns", not one alone.
Expected behavior
The message list should be normalized before each chat-completions call (no ≥2 consecutive trailing assistant messages; generation always starts from a user/tool turn). At minimum, the llama.cpp structure/overflow error should trigger a clean compaction instead of getting stuck with tools disabled.
Related issues
Related to #151, #211, #334.