Pre-generation path (chat.params) aborts session at context threshold instead of switching to largeContextModel
Environment
- opencode: v1.18.30
- opencode-auto-fallback: v0.4.59
- Config: per-agent
fallback chains + largeContextModel for prod/build agents (primary model opencode/big-pickle, 200K context / 160K input limit; large context model opencode/glm-5.2, 1M context)
Symptom
Sessions whose context exceeds ~95% of the model's input limit are aborted by the plugin's pre-generation check instead of being switched to the configured largeContextModel:
2026-09-14T06:00:06.073Z [INFO] Pre-generation: context at threshold, aborting {"sessionID":"ses_...","usage":152354,"limit":200000,"atThreshold":true}
Real-world impact:
- A cron-driven research session was killed 9 seconds after start (06:00:06) because its stored context was already above the threshold — the large context fallback never ran.
- A long-running research session (167K/200K) was repeatedly aborted the same way; the only way to continue was a manual
/compact or a manual model switch.
Root cause
The large-context switch is implemented in the idle path (session.status handler), but the pre-generation path in the chat.params handler only checks the threshold and aborts — it never attempts the large context model switch:
// dist/index.js (v0.4.59, chat.params handler)
if (!getLargeContextPhase(input.sessionID)) {
if (input.agent) {
const threshold = await checkContextThreshold(input.sessionID, context, logger);
if (threshold.atThreshold) {
await logger.info("Pre-generation: context at threshold, aborting", { /* ... */ });
await abortSessionSafely(input.sessionID, context);
return;
}
}
}
So whenever a message is sent while the session is already above the threshold (scheduled/cron prompts, prompt_async, or a user sending before the idle handler gets a chance), the session dies instead of escalating to the large context model.
Fix (patched locally)
Resolve the agent the same way the idle path does (getSessionOriginalAgent(sessionID) ?? input.agent), look up the agent's largeContextModel, and switch before aborting. The abort remains as the fallback for unknown agents, cooldown, same-model, or insufficient context increase:
if (!getLargeContextPhase(input.sessionID)) {
if (input.agent) {
const threshold = await checkContextThreshold(input.sessionID, context, logger);
if (threshold.atThreshold) {
const agent = getSessionOriginalAgent(input.sessionID) ?? input.agent;
const lcfParsed = agent ? getAgentLargeContextModel(config, agent) : null;
if (lcfParsed && isRegisteredAgent(agent)) {
const curModel = getCurrentModel(input.sessionID);
if (curModel && !isSameModel(curModel, lcfParsed) && !isModelInCooldown(lcfParsed.providerID, lcfParsed.modelID)) {
const largeLimit = getModelContextLimit(formatModelKey(lcfParsed));
const minRatio = getAgentMinContextRatio(config, agent);
if (!largeLimit || !shouldSkipLargeContextFallback(threshold.limit, largeLimit, minRatio)) {
await logger.info("Pre-generation: context at threshold, switching to large model", { /* ... */ });
await abortSessionSafely(input.sessionID, context);
const switched = await handleLargeContextSwitch(
input.sessionID,
lcfParsed,
context,
logger,
`Context at ${(threshold.usage / threshold.limit * 100).toFixed(1)}%`
);
if (switched) return;
}
}
}
await logger.info("Pre-generation: context at threshold, aborting", { /* ... */ });
await abortSessionSafely(input.sessionID, context);
return;
}
}
}
Verification (live)
Sent POST /session/:id/message with model: big-pickle + agent: prod to a session stuck at 167K/200K (previously aborted repeatedly):
Before patch:
2026-09-14T08:35:38.170Z [INFO] Pre-generation: context at threshold, aborting {"sessionID":"ses_...","usage":166746,"limit":200000,"atThreshold":true}
After patch:
2026-09-14T09:28:38.074Z [INFO] Session status: busy {"sessionID":"ses_...","phase":"active","agent":"prod"}
2026-09-14T09:28:38.130Z [INFO] Model changed, cooldown reset {"sessionID":"ses_...","model":"opencode/glm-5.2","previousModel":"opencode/big-pickle"}
2026-09-14T09:28:38.371Z [INFO] Detected model context limit {"sessionID":"ses_...","model":"opencode/glm-5.2","contextLimit":1000000}
2026-09-14T09:28:38.373Z [INFO] Applying fallback model params {"sessionID":"ses_...","model":"opencode/glm-5.2"}
The in-flight message was aborted safely, the session switched to the large context model, the synthetic continuation prompt was answered, compaction ran on the 1M-context model, and the session then returned to the default model (~54K context) and resumed work — no manual intervention needed.
Happy to turn this into a PR against src/ if useful.
Pre-generation path (
chat.params) aborts session at context threshold instead of switching tolargeContextModelEnvironment
fallbackchains +largeContextModelforprod/buildagents (primary modelopencode/big-pickle, 200K context / 160K input limit; large context modelopencode/glm-5.2, 1M context)Symptom
Sessions whose context exceeds ~95% of the model's input limit are aborted by the plugin's pre-generation check instead of being switched to the configured
largeContextModel:Real-world impact:
/compactor a manual model switch.Root cause
The large-context switch is implemented in the idle path (
session.statushandler), but the pre-generation path in thechat.paramshandler only checks the threshold and aborts — it never attempts the large context model switch:So whenever a message is sent while the session is already above the threshold (scheduled/cron prompts,
prompt_async, or a user sending before the idle handler gets a chance), the session dies instead of escalating to the large context model.Fix (patched locally)
Resolve the agent the same way the idle path does (
getSessionOriginalAgent(sessionID) ?? input.agent), look up the agent'slargeContextModel, and switch before aborting. The abort remains as the fallback for unknown agents, cooldown, same-model, or insufficient context increase:Verification (live)
Sent
POST /session/:id/messagewithmodel: big-pickle+agent: prodto a session stuck at 167K/200K (previously aborted repeatedly):Before patch:
After patch:
The in-flight message was aborted safely, the session switched to the large context model, the synthetic continuation prompt was answered, compaction ran on the 1M-context model, and the session then returned to the default model (~54K context) and resumed work — no manual intervention needed.
Happy to turn this into a PR against
src/if useful.