feat(web): continue interrupted assistant responses - #1717
Conversation
|
Thanks, this is the shape I asked for in #1699, and the |
A "Continue" action on the last assistant message resends the transcript ending on that turn and streams the model's resumption into the same bubble, instead of opening a new turn -- the client-side surface for the server's COLI_CONTINUE_ASSISTANT path (default on across the shipped families). Against an older engine that does not continue, the same request folds the turn into a completed one and appends a fresh cue, so it reads as an ordinary re-answer. The streaming core (metrics, abort, error handling) is extracted from send() into a shared runStream(payload, targetId); send() and the new continueTurn() differ only in what they send and which bubble they stream into, so api.ts is unchanged. continueTurn() mirrors the server's refusals so the UI never trips them: it strips trailing whitespace (the server 400s on it, since the template strips it and the model would resume from different bytes) and sets the bubble to those exact bytes first, so what is shown is what it resumes from; it guards an empty/whitespace-only turn; and on abort or error it keeps the existing text on screen, dropping only a bubble that never got a token. A continued turn streams back as content, not reasoning_content, which the existing stream path already handles. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Continuing a turn the model already closed hands it a prompt that ends where the model chose to stop, so it re-emits its stop token and adds nothing -- a dead button. Show Continue only when the last turn is actually open: it hit max_tokens (finishReason "length") or was stopped by hand (an abort, which sets no result, so it is tracked separately as "aborted"). A turn the model closed itself (finishReason "stop") no longer offers it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The web UI's Continue button needs to know whether a trailing assistant turn will actually be continued: with COLI_CONTINUE_ASSISTANT=0 (or a family without an open-turn shape) the server answers it fresh, and a Continue that quietly does that is worse than no button. Expose the effective switch next to the scheduler fields, behind the same auth gate. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With COLI_CONTINUE_ASSISTANT=0, or against a server that predates the field, a trailing assistant turn is answered fresh, and a Continue button there would present a re-answer as a resumption. Gate the button on the new /health field, treating an absent one as off. Move the resend and the delta append into lib/transcript.ts so vitest can pin them: the continued request ends on the assistant turn, and the resumption lands in that same bubble instead of opening a new one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Whether Continue is offered was decided by one app-wide lastFinish. Clearing it on slot switches and archive restores kept it from leaking into another conversation, but also dropped Continue from an interrupted reply once you came back to it. And a run that failed mid-stream never set it, so a partial reply showed or hid Continue depending on how the previous reply had ended. Record the finish on the assistant message runStream streamed into, and decide from that message alone. Besides the server's finish_reason it can be "aborted" or "error" when the client stopped or lost the stream, or "incomplete" when the stream closed without a finish_reason. colibri always sends one in its final chunk, and its SSE body is close-delimited, so an engine failure after the 200 reaches the browser as a clean end of stream with no finish; treating that as "stop" hid Continue on a reply that was cut off. All three are open turns and offer Continue. The field is never sent: streamChat serialises only role, content and images, which the transcript test now pins with a finish-carrying fixture. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
a1c4b27 to
96467ec
Compare
@JustVugg Rebased onto
Checks: vitest 34/34 (including reasoning-stream, chat and Continue tests), Continue still requires answer text: a turn stopped during thinking retains its reasoning but does not offer Continue. |
…erve_thinking (JustVugg#1767) and the engine's graceful exit (JustVugg#1734) under the logprobs work Dev's other motion into c/openai_server.py comes along unchanged: brio's normalize defaults to sum (JustVugg#1753), the authed /health reports continue_assistant (JustVugg#1717), and JustVugg#1777 also brings qwen36 vision checkpoints through the qwen38 image path and the registry's default model id. The conflicting import list in c/tests/test_openai_server.py takes the union of both sides, which keeps dev's _dsv4_tool_calls import (unused in dev as well).
Summary
Closes #1699.
Adds Continue to the last assistant message when the response was cut off by
max_tokens, is stopped manually, or is interrupted by an error. Responses whose streams end without a finish reason also remain continuable.Continue resends the conversation ending on that message and appends the generated text to the same response, using #1402's continuation support. The action is hidden during generation and after a normal completion. As requested in #1699, the authenticated
/healthresponse now includescontinue_assistant, next to the scheduler fields. It is true when continuation is enabled and the loaded family is one ofCONTINUATION_FAMILIES(every shipped family today).The UI requires an explicit
true, so servers withCOLI_CONTINUE_ASSISTANT=0or without the field do not show Continue. Unauthenticated health probes retain their existing response.Sending and continuing share the streaming path. Before continuing, the UI trims trailing whitespace from the assistant message to satisfy the server's validation, and updates the displayed text to match. An abort or error preserves the text already received. The Continue label is added for languages with existing UI coverage: English, Italian, and Indonesian, with the existing English fallback for the other locales.
Validation
make -C c checkAdded coverage:
streamChatwith a mocked HTTP response.continue_assistant: true./healthresponses, and omitting it from unauthenticated responses.All 27 web tests and
tsc -bpass. All 215 tests intest_openai_serverpassed at 42969b6; subsequent changes only affect web files.These tests cover the
lib/helpers, notApp.tsx's use of them. A UI rewrite that bypassed those helpers could still pass them. The web tests only coverlib/today, and closing that gap in my understanding requiresjsdomand@testing-library/reactas devDependencies for a component test. As such these were left out, but I can proceed to add such coverage to this PR if you would prefer it.UI tested end to end with GLM-5.3-Flash int4 on CPU: both a response cut off by the token limit and a manually stopped response continued in place, and a normally completed response showed no Continue action.
No visual change from the prototype screenshots attached to #1699:
After completed turn:
Compatibility
No new web dependencies or chat request fields. The API change is the additional field in authenticated
/healthresponses.