Conversation
|
Empirical evidence (gfx1101, ROCm/HIP 7.15 container Routes/gating:
End-to-end turn: UI-shaped streamed request (
CI note: the |
b2d8573 to
2d9042d
Compare
`(b >> 2j) & 0x3 <= 3` is always true; clippy's deny-level bad_bit_mask stopped `cargo clippy --all-targets` at hipfire-quantize. Pin the block geometry instead.
The no-GPU CI runner has no /opt/rocm/core-10.0, so peacemaker-ir's codec gates panicked in the lib-test step and 14 hipfire-isa integration tests in the GPU-free step. Follow the existing table-gate convention: skip with a log line when llvm-mc is absent; the gates still run where it exists.
clippy's deny-by-default mut_from_ref stopped `cargo clippy --all-targets` at hipfire-xdna once hipfire-quantize compiled. Bo is a cloneable handle to one shared mapping, so &mut self would not prevent aliasing; exclusivity is the documented caller contract. Same fix as a1ac7f4 on feat/qwen38-flash-next.
`(b >> 2j) & 0x3 <= 3` is always true; clippy's deny-level bad_bit_mask stopped `cargo clippy --all-targets` at hipfire-quantize. Pin the block geometry instead.
The no-GPU CI runner has no /opt/rocm/core-10.0, so peacemaker-ir's codec gates panicked in the lib-test step and 15 hipfire-isa integration tests in the GPU-free step. Follow the existing table-gate convention: skip with a log line when llvm-mc is absent; the gates still run where it exists.
clippy's deny-by-default mut_from_ref stopped `cargo clippy --all-targets` at hipfire-xdna once hipfire-quantize compiled. Bo is a cloneable handle to one shared mapping, so &mut self would not prevent aliasing; exclusivity is the documented caller contract. Same fix as a1ac7f4 on feat/qwen38-flash-next.
…ad lifecycle a31e438 (land/041t) grew crates/hipfire-daemon/src/main.rs to 5392 lines without moving the ceiling, so beta's own gates job fails. The 15 lines are no_model_loaded_suffix (an admitted single-device/pp load failure now says the prior model is gone) and its unit test. RATCHET-RAISE: daemon_lines 5377 -> 5392, traded for the Flash-Next unload-first load lifecycle (a31e438) that fixes the FN<->VMM swap refusal and names the missing model on admitted-load failure
…E_UI) Port of feat/serve-chat-ui-v040 onto the 0.4.x beta line (the v040 branch is authored against master; serve/complete.rs, serve/mod.rs and serve/http.rs have since been refactored on beta, so the Rust hunks were re-applied against the beta shapes rather than rebased). What ships: - Zero-dependency chat frontend: five static assets embedded via include_str! (index/app/db/render/style), served by a new serve/ui.rs matcher that shares the listener and the existing /v1/chat/completions path. IndexedDB history, branching, markdown/highlighting, effort picker, live TTFT/tok-s strip, mobile layout, ember palette. - Enablement in precedence order: --ui, --open (implies --ui, launches the browser), serve.ui = true, HIPFIRE_SERVE_UI=1. Off is byte-for-byte the old behavior: / and /ui both 404. Non-loopback binds warn on stderr; detached children get --ui. - Security headers on every UI route: strict CSP (default-src 'none', frame-ancestors 'none', form-action 'none'), X-Frame-Options: DENY, nosniff, no-cache — the page is an unauthenticated control surface and is not frameable. - /health and /v1/models publish the resident model's reasoning contract and effort rungs (from meta.loaded — no runtime lock), and capabilities gains chat_ui. - Carries the two gateway fixes the UI surfaced: mid-generation daemon failures without attested rollback poison the session so the next request cold-resets (poisons_session_state), and completion_timings no longer reports wall tok/s as decode tok/s. (slots.rs unchanged: beta already clamps max_tokens_fit via hipfire_generate::common::fit_max_tokens_if.) Verified on gfx1101 against a fresh --ui --no-prewarm serve: / -> 302 /ui; /ui 200 with the full header set; all four assets 200 with correct content types; non-UI paths still 404; /health reports chat_ui: true and the reasoning fields. cargo test -p hipfire-cli 312/312; crate maps regenerated.
The module-level use sat unused on beta; the ported test now resolves through it instead of shadowing with its own import (clears the unused-import advisory).
The UI held a single global in-flight guard (state.stream), serializing every chat through one request at a time even when serve.multi_slot runs the daemon's concurrent slot engine. Streams are now keyed by conversation id (state.streams: Map), with curStream()/anyStreaming() helpers. - send/regenerate/edit/tool-reply/variant-switch block only on the same conversation's stream; other chats generate in parallel - stop button and Esc abort the open conversation's stream only - deleteConv/clearAll/bc del+reload abort the matching streams - sidebar conv-streaming dot shows for every in-flight chat - per-conversation debounced saves (a shared timer dropped the second stream's writes) - view updates/scroll only for the open conversation; streaming view is nulled when switching away and rebound on return - doc title shows ● Generating… while any stream runs hidden - busy toast rewritten: 'keep working in another chat' Verified against a stubbed /v1/chat/completions SSE backend driving the real page in headless Chromium over CDP: two conversations streamed concurrently (streams.size=2, both lens growing), same-chat send refused mid-stream, stop aborted only the open chat, deleteConv aborted its stream, sidebar dots tracked in-flight convs.
…lot thinking-on) Every thinking-on (or thinking-auto) request on the multi-slot route failed closed with 'reasoning prefix mismatch: expected OpenThink got ClosedThink'. The route derived the continuation/reasoning framing as OpenThink / ClosedThink / Plain from `has_think`, which is the *tokenizer's* inventory of a `<think>` special token — not the prompt. Standard Qwen3-family templates emit no opener when thinking is enabled and let the model generate `<think>` itself (only enable_thinking=false emits a closed `<think>\n\n</think>\n\n` block), so a thinking-on render ends in a bare `assistant\n`: started_in_think=false and has_think=true fabricated ClosedThink, which then failed the equality check against the request-derived expected OpenThink. The Jinja render owns the generation suffix, so the framing is read back from what the template actually emitted: new render_assistant_prefix() (sibling of render_tail_opens_think, tail-only for the same reason) maps a closed `</think>` tail to ClosedThink, an open `<think>` tail to OpenThink, and anything else to Plain. The VL and ChatFrame builders frame the assistant turn themselves from expected_prefix and keep returning it. The fail-closed equality check is removed: it compared two independent derivations of the same quantity, and the render is the authoritative one. started_in_think (what emit_generation_start latches) is now simply prefix == OpenThink. Reproduced against the MiMo-V2.6-Distill-Qwen-9B container on this branch: 'hi' with no thinking field, enable_thinking=true both returned the mismatch error; enable_thinking=false worked. Also present on upstream/beta and upstream/master (slots.rs:1168/1172) — not a regression from this branch. Test: render_tail_think_tests::rendered_framing_is_read_from_the_suffix.
…hinking Second gate on the same thinking-on path, reached once the prefix check was fixed: 'think_mode mismatch with enable_thinking' (typed internal, so a 500). The route derived think_mode from the prompt framing (started_in_think) and then failed closed unless it agreed with the request's enable_thinking. Those are different quantities: enable_thinking is a Jinja *input*, and a thinking-on Qwen3-family template frames no opener — the model emits <think> itself, so started_in_think is false while enable_thinking is true. The parser must start outside the reasoning span and open it on that token, which is what the framing-derived value already says. Drop the equality check and keep think_mode = framing. Also drop the now-stale 'think' case from the post-start error-release seam list (the remaining three still cover the internal class).
- Deleted-chat toast no longer covers the composer: #toasts anchors below the topbar (top: calc(var(--topbar-h) + 12px)) instead of the bottom edge, and the enter/leave keyframes drop from above to match. - Each sidebar row gains a visible delete button beside the three-dot menu, wired to the same deleteConv path (undo toast included). It shares the three-dot button's reveal: hidden until row hover/active/focus, always visible on touch/narrow. Hover tints it with the error colour. - Drop the topbar chat-title button (redundant with the sidebar row): the Rename action now targets the sidebar row, opening the sidebar first when it is collapsed. Removes .conv-title/.conv-title-input and the responsive hide. The delete entry stays in the three-dot menu as before.
31cea59 to
96388a7
Compare
…beta # Conflicts: # crates/hipfire-daemon/map.md # docs/env-vars.md
…erve-chat-ui-beta
…ilure suffix helper out of main.rs land/041t + land/041u grew daemon main.rs to 5392 against the 5377 ceiling, leaving beta's own gates job red for every branch based on its head. Raising the ceiling needs the maintainer-only 'ratchet-raise' label, so pay the debt instead: no_model_loaded_suffix and its test move to request_guards (the module whose charter is exactly 'small load/generate message helpers kept out of main.rs'), semantics untouched — main.rs lands at 5372. Also refreshes the stale crate maps (engine/generate/quantize/cli/config/daemon).
… (main.rs is back under it)
Summary
A zero-dependency chat frontend embedded in the serve binary, off by default:
hipfire serve --ui(or--open,serve.ui = true,HIPFIRE_SERVE_UI=1) serves it at/uion the same listener and the same/v1/chat/completionspath every other client uses. Chat history lives in the browser's IndexedDB — serve stores nothing. Off is byte-for-byte the old behavior (/and/ui404).Which surface(s) does this touch?
crates/hipfire-cli/src/serve(gateway routes, session-reset flag),crates/hipfire-config(serve.uikey)hipfire-cliflags, config key, docsslots.rs— unchanged: beta already clampsmax_tokens_fitviahipfire_generate::common::fit_max_tokens_ifDetails
assets/ui/(index/app/db/render/style), embedded viainclude_str!; no new crates, no JS toolchain. History sidebar with pin/search/export/import, message branching (edit/regenerate forks), markdown + syntax highlighting + TeX-unicode math, reasoning disclosure, effort picker, image attach gated on the model's vision capability, live TTFT/tok-s strip, mobile layout, theme switcher.default-src 'none'; script-src 'self'; frame-ancestors 'none'; form-action 'none',X-Frame-Options: DENY, nosniff, no-cache. Unit tests assert the CSP value (clickjacking class closed, not just header presence)./healthand/v1/modelsarchitecturepublish the resident model's reasoning contract and effort rungs (served frommeta.loaded, no runtime lock);capabilitiesgainschat_ui.rolled_back: falsenow poison session state so the next request cold-resets (PR serve: cold-reset the next request after a mid-generation failure #806);completion_timingsno longer reports walltok_sasdecode_tok_s(PR serve: completion timings no longer report wall tok/s as decode tok/s #807). Both are needed for the UI to be correct; keeping them here too so this branch is self-contained — drop either hunk if the master fix lands first.feat/serve-chat-ui-v040(master-based) onto beta; serve/complete/mod/http were refactored since, so hunks were re-applied against beta's shapes (slot-side clamp already exists on beta and was dropped).Test plan
cargo build --release -p hipfire-cli -p hipfire-daemoncleancargo test -p hipfire-cli— 312/312 (incl. new ui route/header, reasoning-facts, poison-classification,serve_ui_url/host_is_loopbackunit tests)scripts/check-crate-maps.py --check,scripts/check-env-docs.py,scripts/check-lifecycle.py,scripts/leanup-ratchets.shgreen; crate maps regenerated--ui --no-prewarm):/302 →/ui200, all assets 200 with correct content types + full header set, non-UI paths 404,/healthreportschat_ui: trueand the reasoning fields. Earlier iteration verified headless at 320/360/390/768px (no overflow/overlap) and a full streamed chat turn against qwen3.5-4b.mq4 multi-slot.serve_harness.py battery— the gateway changes are additive (new routes + flag) but the poison-fix touches session lifecycle; recommend a harness pass on a flagship artifact. Daemon binary unchanged.Hardware validation request (optional)
{ "routes": [ {"mode": "battery", "tag": "qwen3.8:27b-mq4-xts"} ], "claim": "serve gateway unchanged for non-UI requests; session cold-reset path behaves on poisoned multi-slot sessions" }