Skip to content

serve: opt-in embedded chat UI at /ui - #808

Open
ghazni101 wants to merge 20 commits into
warpfront:betafrom
ghazni101:feat/serve-chat-ui-beta
Open

ghazni101 wants to merge 20 commits into
warpfront:betafrom
ghazni101:feat/serve-chat-ui-beta

Conversation

@ghazni101

Copy link
Copy Markdown
Contributor

Summary

A zero-dependency chat frontend embedded in the serve binary, off by default: hipfire serve --ui (or --open, serve.ui = true, HIPFIRE_SERVE_UI=1) serves it at /ui on the same listener and the same /v1/chat/completions path every other client uses. Chat history lives in the browser's IndexedDB — serve stores nothing. Off is byte-for-byte the old behavior (/ and /ui 404).

Which surface(s) does this touch?

  • serve — crates/hipfire-cli/src/serve (gateway routes, session-reset flag), crates/hipfire-config (serve.ui key)
  • control plane — hipfire-cli flags, config key, docs
  • daemon slots.rs — unchanged: beta already clamps max_tokens_fit via hipfire_generate::common::fit_max_tokens_if

Details

  • Assets: assets/ui/ (index/app/db/render/style), embedded via include_str!; no new crates, no JS toolchain. History sidebar with pin/search/export/import, message branching (edit/regenerate forks), markdown + syntax highlighting + TeX-unicode math, reasoning disclosure, effort picker, image attach gated on the model's vision capability, live TTFT/tok-s strip, mobile layout, theme switcher.
  • Security: unauthenticated control surface, so opt-in only; non-loopback binds warn on stderr (and in the detached-child announce); every UI route carries CSP default-src 'none'; script-src 'self'; frame-ancestors 'none'; form-action 'none', X-Frame-Options: DENY, nosniff, no-cache. Unit tests assert the CSP value (clickjacking class closed, not just header presence).
  • API surface: /health and /v1/models architecture publish the resident model's reasoning contract and effort rungs (served from meta.loaded, no runtime lock); capabilities gains chat_ui.
  • Gateway fixes carried (same bugs, also filed against master): mid-generation daemon errors with rolled_back: false now poison session state so the next request cold-resets (PR serve: cold-reset the next request after a mid-generation failure #806); completion_timings no longer reports wall tok_s as decode_tok_s (PR serve: completion timings no longer report wall tok/s as decode tok/s #807). Both are needed for the UI to be correct; keeping them here too so this branch is self-contained — drop either hunk if the master fix lands first.
  • Port of feat/serve-chat-ui-v040 (master-based) onto beta; serve/complete/mod/http were refactored since, so hunks were re-applied against beta's shapes (slot-side clamp already exists on beta and was dropped).

Test plan

  • cargo build --release -p hipfire-cli -p hipfire-daemon clean
  • cargo test -p hipfire-cli — 312/312 (incl. new ui route/header, reasoning-facts, poison-classification, serve_ui_url/host_is_loopback unit tests)
  • scripts/check-crate-maps.py --check, scripts/check-env-docs.py, scripts/check-lifecycle.py, scripts/leanup-ratchets.sh green; crate maps regenerated
  • Live smoke on gfx1101 (--ui --no-prewarm): / 302 → /ui 200, all assets 200 with correct content types + full header set, non-UI paths 404, /health reports chat_ui: true and the reasoning fields. Earlier iteration verified headless at 320/360/390/768px (no overflow/overlap) and a full streamed chat turn against qwen3.5-4b.mq4 multi-slot.
  • serve_harness.py battery — the gateway changes are additive (new routes + flag) but the poison-fix touches session lifecycle; recommend a harness pass on a flagship artifact. Daemon binary unchanged.

Hardware validation request (optional)

{
  "routes": [
    {"mode": "battery", "tag": "qwen3.8:27b-mq4-xts"}
  ],
  "claim": "serve gateway unchanged for non-UI requests; session cold-reset path behaves on poisoned multi-slot sessions"
}

@ghazni101

Copy link
Copy Markdown
Contributor Author

Empirical evidence (gfx1101, ROCm/HIP 7.15 container local/rocm-build:1, qwen3.5-4b.mq4):

Routes/gating:

  • HIPFIRE_SERVE_UI unset: /, /ui, /ui/app.js all 404 (byte-for-byte off behavior)
  • HIPFIRE_SERVE_UI=1 and --ui: / → 302 → /ui 200; /ui/{app,db,render}.js, /ui/style.css all 200 with correct content types; /ui/../etc, /ui/nope.js 404; every UI route carries the full CSP + X-Frame-Options: DENY + nosniff
  • 0.0.0.0 bind logs: WARNING: chat UI is unauthenticated and bound to 0.0.0.0
  • /health after load: reasoning_contract: "qwen_jinja", reasoning_efforts: [low,medium,high,xhigh,max]; same fields under /v1/models → architecture; capabilities.chat_ui: true

End-to-end turn: UI-shaped streamed request (stream:true, stream_options.include_usage, enable_thinking:false) → clean SSE deltas, coherent answer, usage present. Carried poison fix verified live on beta multi-slot: open_think 500 (rolled_back=false) → "session state poisoned — next request cold-resets" logged + enforced.

serve_harness.py battery: 5/5 turns, finish=stop, 0 runaway/empty/attractor (battery-808.json, daemon md5 a9756e6651a220474ecaab2f1ab39315).

CI note: the unit + integration tests failure is peacemaker-ir codec tests (pinned llvm-mc: No such file or directory) — identical failures on plain-beta runs (e.g. run 37068182226), a runner-environment issue unrelated to this diff. New-code clippy warning introduced by the port (unused ClientError import) is fixed in HEAD. In-browser visual check not included — no display environment in the test container; earlier iteration was headless-verified at 320/360/390/768px with zero overflow/overlaps.

@ghazni101
ghazni101 force-pushed the feat/serve-chat-ui-beta branch from b2d8573 to 2d9042d Compare October 2, 2026 23:16
Bjoern Agent added 3 commits October 3, 2026 08:39
`(b >> 2j) & 0x3 <= 3` is always true; clippy's deny-level bad_bit_mask
stopped `cargo clippy --all-targets` at hipfire-quantize. Pin the block
geometry instead.
The no-GPU CI runner has no /opt/rocm/core-10.0, so peacemaker-ir's codec
gates panicked in the lib-test step and 14 hipfire-isa integration tests
in the GPU-free step. Follow the existing table-gate convention: skip with
a log line when llvm-mc is absent; the gates still run where it exists.
clippy's deny-by-default mut_from_ref stopped `cargo clippy --all-targets`
at hipfire-xdna once hipfire-quantize compiled. Bo is a cloneable handle to
one shared mapping, so &mut self would not prevent aliasing; exclusivity is
the documented caller contract. Same fix as a1ac7f4 on
feat/qwen38-flash-next.
Bjoern Agent and others added 11 commits October 3, 2026 15:59
`(b >> 2j) & 0x3 <= 3` is always true; clippy's deny-level bad_bit_mask
stopped `cargo clippy --all-targets` at hipfire-quantize. Pin the block
geometry instead.
The no-GPU CI runner has no /opt/rocm/core-10.0, so peacemaker-ir's codec
gates panicked in the lib-test step and 15 hipfire-isa integration tests
in the GPU-free step. Follow the existing table-gate convention: skip with
a log line when llvm-mc is absent; the gates still run where it exists.
clippy's deny-by-default mut_from_ref stopped `cargo clippy --all-targets`
at hipfire-xdna once hipfire-quantize compiled. Bo is a cloneable handle to
one shared mapping, so &mut self would not prevent aliasing; exclusivity is
the documented caller contract. Same fix as a1ac7f4 on
feat/qwen38-flash-next.
…ad lifecycle

a31e438 (land/041t) grew crates/hipfire-daemon/src/main.rs to 5392 lines
without moving the ceiling, so beta's own gates job fails. The 15 lines are
no_model_loaded_suffix (an admitted single-device/pp load failure now says
the prior model is gone) and its unit test.

RATCHET-RAISE: daemon_lines 5377 -> 5392, traded for the Flash-Next unload-first load lifecycle (a31e438) that fixes the FN<->VMM swap refusal and names the missing model on admitted-load failure
…E_UI)

Port of feat/serve-chat-ui-v040 onto the 0.4.x beta line (the v040
branch is authored against master; serve/complete.rs, serve/mod.rs and
serve/http.rs have since been refactored on beta, so the Rust hunks
were re-applied against the beta shapes rather than rebased).

What ships:

- Zero-dependency chat frontend: five static assets embedded via
  include_str! (index/app/db/render/style), served by a new
  serve/ui.rs matcher that shares the listener and the existing
  /v1/chat/completions path. IndexedDB history, branching,
  markdown/highlighting, effort picker, live TTFT/tok-s strip,
  mobile layout, ember palette.
- Enablement in precedence order: --ui, --open (implies --ui,
  launches the browser), serve.ui = true, HIPFIRE_SERVE_UI=1.
  Off is byte-for-byte the old behavior: / and /ui both 404.
  Non-loopback binds warn on stderr; detached children get --ui.
- Security headers on every UI route: strict CSP (default-src
  'none', frame-ancestors 'none', form-action 'none'),
  X-Frame-Options: DENY, nosniff, no-cache — the page is an
  unauthenticated control surface and is not frameable.
- /health and /v1/models publish the resident model's reasoning
  contract and effort rungs (from meta.loaded — no runtime lock),
  and capabilities gains chat_ui.
- Carries the two gateway fixes the UI surfaced: mid-generation
  daemon failures without attested rollback poison the session so
  the next request cold-resets (poisons_session_state), and
  completion_timings no longer reports wall tok/s as decode tok/s.
  (slots.rs unchanged: beta already clamps max_tokens_fit via
  hipfire_generate::common::fit_max_tokens_if.)

Verified on gfx1101 against a fresh --ui --no-prewarm serve:
/ -> 302 /ui; /ui 200 with the full header set; all four assets 200
with correct content types; non-UI paths still 404; /health reports
chat_ui: true and the reasoning fields. cargo test -p hipfire-cli
312/312; crate maps regenerated.
The module-level use sat unused on beta; the ported test now resolves
through it instead of shadowing with its own import (clears the
unused-import advisory).
The UI held a single global in-flight guard (state.stream), serializing
every chat through one request at a time even when serve.multi_slot runs
the daemon's concurrent slot engine. Streams are now keyed by
conversation id (state.streams: Map), with curStream()/anyStreaming()
helpers.

- send/regenerate/edit/tool-reply/variant-switch block only on the same
  conversation's stream; other chats generate in parallel
- stop button and Esc abort the open conversation's stream only
- deleteConv/clearAll/bc del+reload abort the matching streams
- sidebar conv-streaming dot shows for every in-flight chat
- per-conversation debounced saves (a shared timer dropped the second
  stream's writes)
- view updates/scroll only for the open conversation; streaming view is
  nulled when switching away and rebound on return
- doc title shows ● Generating… while any stream runs hidden
- busy toast rewritten: 'keep working in another chat'

Verified against a stubbed /v1/chat/completions SSE backend driving the
real page in headless Chromium over CDP: two conversations streamed
concurrently (streams.size=2, both lens growing), same-chat send
refused mid-stream, stop aborted only the open chat, deleteConv aborted
its stream, sidebar dots tracked in-flight convs.
…lot thinking-on)

Every thinking-on (or thinking-auto) request on the multi-slot route failed
closed with 'reasoning prefix mismatch: expected OpenThink got ClosedThink'.

The route derived the continuation/reasoning framing as OpenThink /
ClosedThink / Plain from `has_think`, which is the *tokenizer's* inventory
of a `<think>` special token — not the prompt. Standard Qwen3-family
templates emit no opener when thinking is enabled and let the model generate
`<think>` itself (only enable_thinking=false emits a closed
`<think>\n\n</think>\n\n` block), so a thinking-on render ends in a bare
`assistant\n`: started_in_think=false and has_think=true fabricated
ClosedThink, which then failed the equality check against the request-derived
expected OpenThink.

The Jinja render owns the generation suffix, so the framing is read back from
what the template actually emitted: new render_assistant_prefix() (sibling of
render_tail_opens_think, tail-only for the same reason) maps a closed
`</think>` tail to ClosedThink, an open `<think>` tail to OpenThink, and
anything else to Plain. The VL and ChatFrame builders frame the assistant
turn themselves from expected_prefix and keep returning it.

The fail-closed equality check is removed: it compared two independent
derivations of the same quantity, and the render is the authoritative one.
started_in_think (what emit_generation_start latches) is now simply
prefix == OpenThink.

Reproduced against the MiMo-V2.6-Distill-Qwen-9B container on this branch:
'hi' with no thinking field, enable_thinking=true both returned the mismatch
error; enable_thinking=false worked. Also present on upstream/beta and
upstream/master (slots.rs:1168/1172) — not a regression from this branch.

Test: render_tail_think_tests::rendered_framing_is_read_from_the_suffix.
…hinking

Second gate on the same thinking-on path, reached once the prefix check was
fixed: 'think_mode mismatch with enable_thinking' (typed internal, so a 500).

The route derived think_mode from the prompt framing (started_in_think) and
then failed closed unless it agreed with the request's enable_thinking. Those
are different quantities: enable_thinking is a Jinja *input*, and a thinking-on
Qwen3-family template frames no opener — the model emits <think> itself, so
started_in_think is false while enable_thinking is true. The parser must start
outside the reasoning span and open it on that token, which is what the
framing-derived value already says.

Drop the equality check and keep think_mode = framing. Also drop the now-stale
'think' case from the post-start error-release seam list (the remaining three
still cover the internal class).
- Deleted-chat toast no longer covers the composer: #toasts anchors below the
  topbar (top: calc(var(--topbar-h) + 12px)) instead of the bottom edge, and
  the enter/leave keyframes drop from above to match.
- Each sidebar row gains a visible delete button beside the three-dot menu,
  wired to the same deleteConv path (undo toast included). It shares the
  three-dot button's reveal: hidden until row hover/active/focus, always
  visible on touch/narrow. Hover tints it with the error colour.
- Drop the topbar chat-title button (redundant with the sidebar row): the
  Rename action now targets the sidebar row, opening the sidebar first when
  it is collapsed. Removes .conv-title/.conv-title-input and the responsive
  hide.

The delete entry stays in the three-dot menu as before.
@ghazni101
ghazni101 force-pushed the feat/serve-chat-ui-beta branch from 31cea59 to 96388a7 Compare October 3, 2026 14:28
…beta

# Conflicts:
#	crates/hipfire-daemon/map.md
#	docs/env-vars.md
…ilure suffix helper out of main.rs

land/041t + land/041u grew daemon main.rs to 5392 against the 5377
ceiling, leaving beta's own gates job red for every branch based on its
head. Raising the ceiling needs the maintainer-only 'ratchet-raise'
label, so pay the debt instead: no_model_loaded_suffix and its test
move to request_guards (the module whose charter is exactly 'small
load/generate message helpers kept out of main.rs'), semantics
untouched — main.rs lands at 5372. Also refreshes the stale crate maps
(engine/generate/quantize/cli/config/daemon).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant