Skip to content

feat(web): continue interrupted assistant responses - #1717

Merged
JustVugg merged 5 commits into
JustVugg:devfrom
enitimeago:feat/web-continue
Sep 24, 2026
Merged

JustVugg merged 5 commits into
JustVugg:devfrom
enitimeago:feat/web-continue

Conversation

@enitimeago

Copy link
Copy Markdown
Contributor

Summary

Closes #1699.

Adds Continue to the last assistant message when the response was cut off by max_tokens, is stopped manually, or is interrupted by an error. Responses whose streams end without a finish reason also remain continuable.

Continue resends the conversation ending on that message and appends the generated text to the same response, using #1402's continuation support. The action is hidden during generation and after a normal completion. As requested in #1699, the authenticated /health response now includes continue_assistant, next to the scheduler fields. It is true when continuation is enabled and the loaded family is one of CONTINUATION_FAMILIES (every shipped family today).

The UI requires an explicit true, so servers with COLI_CONTINUE_ASSISTANT=0 or without the field do not show Continue. Unauthenticated health probes retain their existing response.

Sending and continuing share the streaming path. Before continuing, the UI trims trailing whitespace from the assistant message to satisfy the server's validation, and updates the displayed text to match. An abort or error preserves the text already received. The Continue label is added for languages with existing UI coverage: English, Italian, and Indonesian, with the existing English fallback for the other locales.

Validation

  • make -C c check
  • CUDA validation: N/A — no CUDA changes.
  • Performance measurements: N/A — no performance claims.

Added coverage:

  • Resending the final assistant message with trailing whitespace removed and appending streamed text to it without adding a message. This exercises the transcript helpers through streamChat with a mocked HTTP response.
  • Recording finish status on the target message, keeping it out of the request, and allowing continuation after an interrupted stream, including one that closes without a finish reason.
  • Showing continuation support only for an explicit continue_assistant: true.
  • Reporting the enabled and disabled continuation setting in authenticated /health responses, and omitting it from unauthenticated responses.

All 27 web tests and tsc -b pass. All 215 tests in test_openai_server passed at 42969b6; subsequent changes only affect web files.

These tests cover the lib/ helpers, not App.tsx's use of them. A UI rewrite that bypassed those helpers could still pass them. The web tests only cover lib/ today, and closing that gap in my understanding requires jsdom and @testing-library/react as devDependencies for a component test. As such these were left out, but I can proceed to add such coverage to this PR if you would prefer it.

UI tested end to end with GLM-5.3-Flash int4 on CPU: both a response cut off by the token limit and a manually stopped response continued in place, and a normally completed response showed no Continue action.

No visual change from the prototype screenshots attached to #1699:

Screenshot 2026-09-21 at 19-17-06 colibrì

After completed turn:

Screenshot 2026-09-21 at 19-26-19 colibrì

Compatibility

  • The default CPU build remains dependency-free
  • No model files, generated binaries, or benchmark artifacts are included

No new web dependencies or chat request fields. The API change is the additional field in authenticated /health responses.

@JustVugg

Copy link
Copy Markdown
Owner

Thanks, this is the shape I asked for in #1699, and the /health field plus the explicit-true check are right. It conflicts with dev in web/src/App.tsx: two PRs merged this week touched the same send() you refactor into runStream — #1697 restored the reasoning stream (onReasoning, the reasoning field, TTFT from the first thinking token, a stop during thinking keeps the turn) and #1709 gave send a consumeDraft parameter and made regenerate resend the last user message's pictures (resendFrom). Please rebase and carry both through runStream, so a continued turn also streams reasoning and neither regenerate nor send loses what those two fixed; their vitest cases (api.test.ts reasoning stream, chat.test.ts) must stay green next to yours. I will merge it as soon as it is clean and green; if that lands before the tag it is in 1.12.1.

enitimeago and others added 5 commits September 24, 2026 19:21
A "Continue" action on the last assistant message resends the transcript ending
on that turn and streams the model's resumption into the same bubble, instead of
opening a new turn -- the client-side surface for the server's
COLI_CONTINUE_ASSISTANT path (default on across the shipped families). Against an
older engine that does not continue, the same request folds the turn into a
completed one and appends a fresh cue, so it reads as an ordinary re-answer.

The streaming core (metrics, abort, error handling) is extracted from send() into
a shared runStream(payload, targetId); send() and the new continueTurn() differ
only in what they send and which bubble they stream into, so api.ts is unchanged.

continueTurn() mirrors the server's refusals so the UI never trips them: it
strips trailing whitespace (the server 400s on it, since the template strips it
and the model would resume from different bytes) and sets the bubble to those
exact bytes first, so what is shown is what it resumes from; it guards an
empty/whitespace-only turn; and on abort or error it keeps the existing text on
screen, dropping only a bubble that never got a token. A continued turn streams
back as content, not reasoning_content, which the existing stream path already
handles.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Continuing a turn the model already closed hands it a prompt that ends where the
model chose to stop, so it re-emits its stop token and adds nothing -- a dead
button. Show Continue only when the last turn is actually open: it hit max_tokens
(finishReason "length") or was stopped by hand (an abort, which sets no result,
so it is tracked separately as "aborted"). A turn the model closed itself
(finishReason "stop") no longer offers it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The web UI's Continue button needs to know whether a trailing assistant
turn will actually be continued: with COLI_CONTINUE_ASSISTANT=0 (or a
family without an open-turn shape) the server answers it fresh, and a
Continue that quietly does that is worse than no button. Expose the
effective switch next to the scheduler fields, behind the same auth gate.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With COLI_CONTINUE_ASSISTANT=0, or against a server that predates the
field, a trailing assistant turn is answered fresh, and a Continue button
there would present a re-answer as a resumption. Gate the button on the
new /health field, treating an absent one as off.

Move the resend and the delta append into lib/transcript.ts so vitest can
pin them: the continued request ends on the assistant turn, and the
resumption lands in that same bubble instead of opening a new one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Whether Continue is offered was decided by one app-wide lastFinish.
Clearing it on slot switches and archive restores kept it from leaking
into another conversation, but also dropped Continue from an interrupted
reply once you came back to it. And a run that failed mid-stream never
set it, so a partial reply showed or hid Continue depending on how the
previous reply had ended.

Record the finish on the assistant message runStream streamed into, and
decide from that message alone. Besides the server's finish_reason it
can be "aborted" or "error" when the client stopped or lost the stream,
or "incomplete" when the stream closed without a finish_reason. colibri
always sends one in its final chunk, and its SSE body is close-delimited,
so an engine failure after the 200 reaches the browser as a clean end of
stream with no finish; treating that as "stop" hid Continue on a reply
that was cut off. All three are open turns and offer Continue.

The field is never sent: streamChat serialises only role, content and
images, which the transcript test now pins with a finish-carrying
fixture.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@enitimeago

enitimeago commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor Author

Thanks, this is the shape I asked for in #1699, and the /health field plus the explicit-true check are right. It conflicts with dev in web/src/App.tsx: two PRs merged this week touched the same send() you refactor into runStream — #1697 restored the reasoning stream (onReasoning, the reasoning field, TTFT from the first thinking token, a stop during thinking keeps the turn) and #1709 gave send a consumeDraft parameter and made regenerate resend the last user message's pictures (resendFrom). Please rebase and carry both through runStream, so a continued turn also streams reasoning and neither regenerate nor send loses what those two fixed; their vitest cases (api.test.ts reasoning stream, chat.test.ts) must stay green next to yours. I will merge it as soon as it is clean and green; if that lands before the tag it is in 1.12.1.

@JustVugg Rebased onto dev, preserving both fixes through runStream:

Checks: vitest 34/34 (including reasoning-stream, chat and Continue tests), npm run build,python3 -m unittest tests.test_openai_server: 237 OK, end-to-end manual test on GLM-5.3-Flash.

Continue still requires answer text: a turn stopped during thinking retains its reasoning but does not offer Continue.

@JustVugg
JustVugg merged commit 6b5671c into JustVugg:dev Sep 24, 2026
29 checks passed
monotophic added a commit to monotophic/colibri that referenced this pull request Sep 28, 2026
…erve_thinking (JustVugg#1767) and the engine's graceful exit (JustVugg#1734) under the logprobs work

Dev's other motion into c/openai_server.py comes along unchanged: brio's
normalize defaults to sum (JustVugg#1753), the authed /health reports
continue_assistant (JustVugg#1717), and JustVugg#1777 also brings qwen36 vision checkpoints
through the qwen38 image path and the registry's default model id.

The conflicting import list in c/tests/test_openai_server.py takes the union
of both sides, which keeps dev's _dsv4_tool_calls import (unused in dev as
well).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants