Repository navigation
[Release] 2026-09-24: v2.6.0 — phone remote bridge over LAN and Tailscale HTTPS, phone microphone, safer goal mode, invented-URL guard - #291
Merged
Conversation
… turn is not an empty loop (#258) goal_last_answer now returns a copy of the last answer carrying every tool call of its turn (BUDGET_SPENT / TOKEN_BUDGET_SPENT / THINK_ONLY_NUDGE do not open a turn). On robin's 2026-09-23 session the brake read the forced round's empty content as "empty" and cut 157 + 52 messages of real work, each cut a 106k-token COLD re-prefill. Replay on session.json: before empty=True, cut at 212; after 24 calls seen, empty=False, no cut. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
… empty answer After BUDGET_SPENT the forced round came back reasoning-only (engine.log 2026-09-23 20:14:38Z: content chunks 0, reasoning chunks 710, finish tool_calls): the window showed nothing, and goal mode's brake judged the 24-round turn by its last message -- three such turns cut 157 messages of work, and the next stopped the goal at 6/9. - run_turn: a final round (forced, or past the #150 nudge) with no visible text stores and shows its reasoning under SURFACED_MARK (tail kept past 6000 chars); with no reasoning either, a line counting the calls that ran. No extra round, no prompt change; reasoning_content stays stored. - goal_turn_empty / goal_turn_mark: a turn with any tool call or result is never empty, and the echo fingerprint covers the whole turn's calls. A surfaced stand-in still counts as empty, so #202's loops still brake. - crow_gui _goal_brake uses both. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
A mid-goal goal_set reported "carried: done 6/9" and left every note in goal.json empty: _carry_goal_marks copied `status` only, so the evidence notes, seconds and tokens -- and the goal's spent -- were reset, and #250's "done after failed needs a note" lost the note it quotes. Carried now for matched done/failed steps: status, note, started, started_tokens, seconds, tokens, delegated (text stays the new plan's); and for the goal: spent with last_spent, delegated, created, started with last_at -- in pairs, so the next step is not billed the whole goal. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…row notes are never a title (#261) A rollover is the live chat's earlier half, not a chat of its own. The rail no longer lists rollover-*.json (except the one open in the window); the archive drawer lists them in place, files untouched, no "restore" row. Titles come from crow_core.chat_title: every Crow user-role note (rollover, budget, token budget, think nudge, abort, goal nudge) is skipped, the #224 root notice is stripped, and a segment opening with a rollover note carries the previous archive's title, so the live chat keeps its name across cuts. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
robin, 2026-09-23: the grey machinery lines clutter long sessions. Moved to a new rotating log, <state_dir>/log/crow.log (~/.local/state/crow/log/crow.log on Linux, %LOCALAPPDATA%\Crow\log\crow.log on Windows; 1 MiB x 3 backups, local time with UTC offset, one [kind] per line): - goal mode: N messages of an empty loop dropped from the history (_goal_cut) - goal mode, step N: the same failure keeps coming back -- ... (the nudge to the model is unchanged) - discarded a degenerate reply / kept the re-asked reply (#217, run_turn) - the restored cache did not hold (window and terminal) - tool budget spent after N rounds (terminal) None was ever in `messages`. One classifier (note_is_log_only) guards Api.push, the terminal's turn_note and clean_notes, so a reopened chat saved by an earlier build no longer draws them either. Kept in the chat: the rollover card and "rolled over at N tokens", goal mode stopped/paused, server reboot/load-wait lines, alarms, Memory updated, carried across the cut, and every answer to a command. Replayed over the live state (read-only): notes band 16 -> 10, exactly the six machinery lines removed. test_crow_core 1320 OK (3 skipped), test_crow 452 OK (3 skipped), test_crow_gui 760 OK (2 skipped), ruff clean. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
A third of every full window on robin's 2026-09-23 diorama run was tool output the model had already acted on (33.0-39.9 % of the request tokens, 97 % of it behind the last five rounds; rendered with the model's template and tokenizer, within 0.05 % of the server's count). The only response to a full window was the 0.9 cut. At 0.65 of the window (default on the local server, off remote, settings.json `context_clear_at` / `--context-clear-at`, 0 = off) every tool result older than the last 5 rounds becomes a one-line stub naming the tool, its target, the size and how to get it again; the originals go to session/cleared/<stamp>.json first. Assistant messages, reasoning, tool calls and the head stay byte-identical. A batch runs only if it frees at least 10 % of the window, so the prefix cache pays one re-prefill per batch, not per round. One line per batch to logs/crow.log, nothing in the chat. A cleared read_file drops its #215-H read stamp; the goal counter rebases. Replay through clear_before_roll on the chained run (790 messages): the two rollovers move from message 188/493 to 299/669 (55 and 88 rounds later) for 2 batches, 158,541 re-prefilled tokens; each recorded context alone would not have reached the cut. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…d line has priority again go() sent every Enter during a running turn to stop() since 4860300, so #165's "a typed line always has priority; the engine only runs when the queue is empty" was unreachable from the live chat. A non-empty plain line mid-turn now goes through the existing send() -> _queued -> _pump path: the page holds it (hint "queued -- it goes in when this turn ends") and draws it at the turn's idle; _pump runs it ahead of _goal_nudge and now calls _goal_reset for it. send() joins a second queued line for the same chat instead of replacing the first. The Stop button (press()) and Escape stay Stop; an empty Enter and slash lines other than /delegate, /subtasks too. test_crow_gui 765 OK (2 skipped; 9 new, 6 red on 0c3e392, 3 counter-probes), test_crow_core 1315 OK (3 skipped), test_crow 451 OK (3 skipped), ruff clean. refs #264 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…overage (#265) The diorama filled 7-22 % of each render on a grainy black ground; at ~1,030 px per visual token the model graded it from ~40-120 tokens. read_image now adds the content box, enlarged to a 1024-px long edge, as a second image and states the coverage; render_page prints metrics: lines (box, coverage, colours, luma), warns under 25 % coverage or 16 colours and saves <shot>-crop.png. Pure Python on the #213 PNG decoder. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…e judge's bar (#267) 2026-09-23 diorama run: 9/9 done over a small box in a black frame; the final notes replayed through goal_done_refusal at 0c3e392 were 9/9 accepted. On a visual goal (keyword guess, goal.json "visual" overrides) a done must now cite an image written after goal.created, and when a judge verdict is stored on the step its lowest score must reach JUDGE_THRESHOLD (default 8; settings.json judge_threshold, --judge-threshold). Planning-only steps are exempt. Replay: 6 of 9 refused. The #250 tests move to a non-visual title. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
2026-09-23: the model scored its own black-frame capture 9+ on every criterion. New tool `judge`: one request with the capture(s) (newest render + its -crop.png if present), the rubric and the goal/step text -- never the transcript. Rubric order: /goal `accept:` lines (stored beside check: in SESSION_DIR), PLAN.md criteria, the caller's, a default visual rubric. Spot order: providers.json `judge` pin, the delegate spot and its fallbacks (catalogue rows now keep `vision` from input_modalities; known blind rows skipped), then this turn's own endpoint in a fresh request, refused when /props is blind. JSON scores, min, weakest 3, stored on the goal step for #267's gate. Live 2026-09-24 on render-20260923-232855.png: four free OpenRouter vision models, min 2 each. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…lover at 6 (#268) 2026-09-23 22:30-22:48: seven chain.html captures in a row at 99.2-99.4 % one colour while every turn called tools; neither #202's brake nor its trouble classes fire on captures. goal_render_scan counts, per page, the render_page captures that are blank-ish (#213/#175 warnings, the crop branch's distinct-colour warn, a decoded share >= 0.98) or near-identical to the previous capture of that page (metrics: lines within 3 %/2 when present, else share within 0.005 and size within 3 %). At 3: one nudge to bisect (clear colour, one cube, passes one by one). At 6: _run rolls on a flag, and the carry lists each stuck capture and the files written before it; one forced roll per step. _goal_brake is untouched (#258/#259). Replay of the four 2026-09-23 transcripts: nudges before messages 35 and 295 (rollover-210418) and 314 (rollover-225111); no roll fired. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…ve one (#268) Found by running the breaker over #265's real metrics lines on three 2026-09-23 captures. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
# Conflicts: # cli/crow.py # cli/crow_gui.py # justfile
… test counts t-judge's loop breaker pushes "goal mode, step N: ... the context rolls over" -- a prefix #262 routes to crow.log. The test now expects it there and not in the chat. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
… and the spot is not memoed dead _run_subtask read sub.cancelled only after the verdict: an attempt that came back "failed" after the user's Stop was written into the chain, marked the spot dead in _SPOT_DEAD and closed "failed". The mark is now read right after each attempt, before anything is decided. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
… serves all three file tools edit_file ran no syntax check (#251 covered write_file/append_file only). It now runs syntax_check after the edit; on a failure the pre-edit text is parsed in a scratch copy, and the result says whether this edit broke the file or it was broken already. settings.json `syntax_checks` maps an extension to an argv ({path} = the file); an empty row switches a built-in off, malformed rows are dropped one by one. Replay of the 2026-09-23 diorama run: 0 of 66 JS/HTML edits broke a file -- this is parity and one place for new formats, not a measured failure. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…ct it, and render_page says why it rendered in software 2026-09-24 the diorama run saved "Machine has NO GPU" into its project memory (the card is an RTX 5090) and repeated it three times; 49 of 49 render_page captures on 2026-09-23 fell back to SwiftShader while serve held the card, and the result never said why. #270: crow_platform.machine_facts() (OS, CPU, RAM, GPU name + total VRAM, static, once per process) as a head line after the working area, with the rule that a tool's limit is not the machine's; memory_contradiction() beside memory_threat refuses no-GPU/CPU-only notes that do not name the card and "a tool changes bytes" notes, on add and replace. #271: render_gl_reason() appended to a software render result. Tests: 6 new, 5 red without the fix (the sixth is the negative probe); 2,633 OK (skipped=8), ruff clean. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…format, rotated crow_log() wrote <state>/logs/crow.log in bare UTC without rotation, while #262's log_note() writes LOG_FILE = <state>/log/crow.log with local time and offset, rotated; both were live on 2026-09-24 (10:17:56 context-clear in one, 10:36:26 goal line in the other). crow_log is log_note(line, "context") now; window.md and settings.md name the one file. Tests: the two clearing tests read LOG_FILE and assert no second log; red without the fix (FileNotFoundError on log/crow.log), green with it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…s note is stored
2026-09-24 diorama run: the engine began step 2 at 10:07, the model re-reported
it running at 11:21 (session.json msg 148), and goal_step_begin reset `started`
without billing the open window -- goal.json read seconds 0, tokens 0, and the
call's note ("... NOT verified ...") was dropped while the tool answered ok.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…ml?view=albedo was "no such page" on 2026-09-24 (session.json msg 36/37) and the model stored "file:// rejects a ?query" in memory tool_render_page handed the whole argument to os.path.isfile, so the diorama prompt's required index.html?shot=<view>&w=<px> could never be rendered. _page_target now checks the whole path first (a file really named with ? or # wins), else the part before the first ?/#, and re-attaches the suffix to a percent-encoded file:// URL built by _file_url (drive letter and UNC per RFC 8089, identical to Path.as_uri() on POSIX; a pure function so the NT form is testable on Linux, PureWindowsPath.as_uri/nturl2path are deprecated in 3.14). A missing page names the FILE. _source_line (#253) takes the path part of a console source URL, which now carries the query. #175's _LAST_CAPTURES and _cache_key already key on the full URL/arguments: two views are two pages (pinned by a test). Tool description and docs/reference/tools.md say queries work. Tests: 7 new (ALocalPageKeepsItsQueryTests), 7/7 red on 0c3e392 (6 fail, 1 AttributeError), green with the fix. test_crow_core 1322 OK (skipped=3), test_crow 451 OK (skipped=3), test_crow_gui 756 OK (skipped=2), ruff clean, check_shared_core 82/82, check_operating_point 10/10. Not verified: a live browser render with a query. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
# Conflicts: # CHANGELOG.md
…le shows the full text approve_pending now returns (saved, failed). A tool refusal (over the limit, no match, contradiction, threat), a duplicate (success without an action) and an expired entry each come back in `failed` with a short reason instead of being swallowed; expired entries are kept aside in _EXPIRED at drop time so the answer can name them. Each miss is logged to crow.log. answer_memory pushes a "Memory: not saved -- ..." note (the glow still fires for what was saved); the terminal prints the same line. pending_view adds `full` (untruncated content) and, for replace/remove, `old` (the whole entry old_text matches right now, read-only; None when there is no single match) plus `find`. The Memory Consolidation tile gets three states: collapsed, previews, full text (a replace as "- old" above "+ new", an add marked "appended at the end"). Same page on the phone layer; nothing in REMOTE_CSS hides the new rows. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…#285 tile, head cost; architecture counts incl. crow_remote.py Audit of 2026-09-24 against 3e0ccc2: the model's own `memory` calls write unasked at every level (only the review is gated); the tile has three states and says what did not land (#285); the 633-char head-cost example and "byte 0" were outdated. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
index_sources() took its list from the rail, and #261 keeps rollover segments out of the rail, so 0 of 6 segments on disk were searchable. The list is split: rail_sources() is the chats of their own (live session.json, session/chat-*.json, archiv/chat-*.json), index_sources() adds session/ and archiv/ rollover-*.json. A segment hit is labelled "<chat title> (before the cut, <date>)", and the tool prints each hit's path for a following read_file. The missing session.json row: "neu" writes the open chat to session/chat-*.json and deletes session.json until the new chat's first turn ends. A search in that turn dropped session.json's rows as gone and never saw session/chat-*.json. index.db was last written 14:02:40.580Z, 19 ms after the first reply of a fresh chat asked for two tools. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
Without the user's accept: lines, judge_rubric(criteria, step_text) puts "delivers: <step text>" first; PLAN.md, the caller's criteria or the default rubric only add to it and can no longer replace it. tool_judge passes the text of the step _judge_step picked, and the stored verdict records rubric_source and the criteria as judged. Step 3 'G-buffer and voxel volume' had passed on step 2's camera criteria the model re-sent. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…; the bar follows goal.json
- goal_step 'failed' counts failures on the step: the first sends it back
once (answer carries `retry`, the next nudge quotes the failure note and
asks for a different approach), the second marks it `skipped` and
next_step moves on.
- goal_next_open, the seam head and goal_set's `next` pass over skipped
steps; a goal whose steps are all done or skipped ends `partial`
("complete with N skipped"), never `done`.
- /goal skip <n> [reason]: one core command for window, phone and
terminal; a done step is not skipped, a running one is billed first.
- Goal bar (desktop and the thin phone bar): skipped mark, note as tooltip,
"step N skipped" / "Complete · step N skipped" in the head.
- The bar follows goal.json: every round (only when a step or the goal
changed state), before every turn, and on `/goal`. A hand edit stayed
hidden until the end of the next turn because push_goal ran only after a
turn and on goal_set/goal_step results.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
A remote host must already appear in the conversation: the user's words (URL or bare name), a tool result (URL), the goal's title/steps or PLAN.md (read from disk at the miss). Otherwise the call comes back as "error: refused: <host> appears nowhere in this conversation -- a URL you made up?" plus a next step, before any socket or browser. Loopback always passes; a subdomain of a named host counts. Always on, yolo included. The set is rebuilt per turn from the conversation (like _MANDATED) and fed after every tool result in run_turn. Across a rollover the note carries the newest 60 hosts in one Crow-written line ahead of the digest, and the next turn re-derives the set from the new context; the model's digest cannot carry a host. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…, dialog choice crow_remote: tailscale_state() reads `tailscale status --json` and `tailscale serve status --json` (no sudo) into missing / down / https-off / serve-missing / funnel / ready. With a ts.net name the server additionally binds 127.0.0.1:<port> (the `tailscale serve --bg --https=443` target) where only Host <name>[:443] and Origin https://<name> pass; a bare 127.0.0.1 Host stays 421 as before. The cookie gets Secure on that origin only. Never 0.0.0.0 or the 100.x address in plain HTTP; Funnel on the name means no loopback listener at all. A failed loopback bind leaves the LAN mirror up. Window: the Remote dialog offers "HTTPS · <name>" as a network choice (remote_use_https, desktop-only, setting remote_https) with one status line per state and the one-time serve command to copy; its QR carries the same single-use code on the ts.net origin once serve points here. /remote and the titlebar icon re-probe and restart once when the tailnet came up after start. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
Phone page (secure context + mediaDevices + MediaRecorder): the mic records
on the phone -- first tap starts, second tap stops, a level ring from an
AnalyserNode -- and uploads the clip to /upload?kind=audio with the same
cookie, Host and Origin guards. The server writes it and answers 202 at once;
the window transcribes it on a thread with crow_voice.transcribe_file()
(same model and settings as stop(); PyAV decode via faster-whisper's own
decode_audio) and pushes {"k":"heard"} to THAT phone only, which puts the
words into its input field, never sent. The clip is deleted either way.
Plain HTTP keeps the keyboard-dictation hint, reworded as in #290. The
desktop's own dictation (its "mic" pushes) no longer lands on a phone.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
docs/user-guide/remote.md: /remote, pairing, the LAN and HTTPS addresses, the phone microphone (#290), guards, troubleshooting. remote-tailscale.md: cost and privacy, phone app, per-OS install, admin console (MagicDNS, Enable HTTPS, CertDomains check), the one-time `tailscale serve` line, the steps in Crow, checks, a troubleshooting table per dialog status line, undo. Commands cited against the Tailscale KB; what could not be verified is marked. Settings reference gains remote_https; docs index and README link both pages. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…erver not found, login 403, two pairing origins) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…one log line per phone dictation
Scope amendment after robin's first live test on the HTTPS address: the ring
moved, nothing arrived (the second tap was not obvious), and the first
dictation spent ~52 s loading the model with no word on the phone.
- Phone (REMOTE_JS section 6): MediaRecorder timeslice 1500 ms; every chunk
uploads the recording so far as /upload?kind=audio&partial=1&seq=N. The
partial shows greyed after the text already in the field (field read-only
while provisional); the final replaces only the partial. Every dictation
APPENDS (one space unless empty or ending in whitespace). Auto-stop 2 s
below the level after >200 ms of speech; silence before speech stops
nothing. Hints: listening / writing / loading the speech model. The button
shows a filled stop square while recording, ring kept, 44 px target.
- Server: /upload parses partial and seq (a partial without seq is 400).
Api._remote_audio keeps one lane per phone: one transcription at a time,
only the newest waiting partial, finals first, partials not newer than the
last final dropped (waiting, running or late). {"k":"heard","loading":true}
while crow_voice.model_loaded() is False. One crow.log line per final
(kind voice): bytes, seconds, transcribe ms, error.
- Measured (faster-whisper small int8, CPU): 5 s clip 1.25-1.4 s, 10 s 1.4 s,
15 s 1.5 s; model load from cache 0.85 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…TPS setup as printed steps Opt-in installer component for #249 stage 5. Both installers read where the machine stands through read-only `tailscale status --json` and `tailscale serve status --json` (the same state table as crow_remote.tailscale_state) and print only the missing steps: the install line (pacman on Arch and Arch-likes, kb/1031's script elsewhere, winget Tailscale.Tailscale on Windows), systemctl enable --now tailscaled, tailscale up, Enable HTTPS in the admin console, `tailscale serve --bg --https=443 http://127.0.0.1:<remote_port>` and the phone app links. Neither runs sudo or elevates. On Windows -Tailscale prints and exits without downloading. install.sh --selftest: 39 checks (was 22), the probe driven against a fake tailscale on a PATH holding nothing else, with an argv log proving only the two read calls ran. install.ps1 -Selftest: the same states through a fake seam (not run here: no PowerShell on this machine). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
- README: optional-components table (--voice, --tailscale / -Tailscale), Tailscale download/App Store/Google Play links in the Phone row; 28 tools, #288 host refusal, #289 retry-then-skip and /goal skip. - install.md, linux.md, remote.md, remote-tailscale.md (new section 0: the installer route and every download link), settings.md, docs index. - window.md crow.log row and dictation model row, remote.md phone clips, goals-and-subagents.md #284 refusal, architecture.md line counts (wc -l at b9cac62), testing.md and justfile: the fourth suite (test_crow_remote) and re-measured case counts. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
VERSION 2.6.0 in cli/crow.py, install.ps1's -Version default, the README badge and manifests/operating-point.json. CHANGELOG: Unreleased becomes 2.6.0 — 2026-09-24 (release summary; Added / Changed / Fixed / Known limitations, every item with its issue numbers), new empty Unreleased. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
… (13 glyphs incl. U+1F3A4) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…a drive-letter file:// source _TROUBLE_PATHLIKE knew only '/' paths, so two identical edit_file refusals on D:\...\a.py and b.py keyed apart (EditFileMissTests on windows-latest). _source_line read 'file://D:\x' as a UNC host; the test built that URL by hand and now uses _file_url, and the code accepts a drive-letter netloc. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…n Windows too (TimeoutError after SYN retries) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
This was referenced Sep 24, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Release v2.6.0: 91 commits since
origin/main(v2.5.0). Crow now runs on the phone. The remote bridge (#249) pairs over the LAN with a QR code, and away from home it runs over Tailscale HTTPS; the installers gain an optional--tailscale/-Tailscalestep. The phone microphone (#290) shows a live transcript, stops on silence and appends to what is already typed. Goal mode got safer: the fresh-eyes judge rubric names the step (#266/#286), visual steps need evidence (#267), a render loop is broken off (#268/#277), and a failed step is retried once, then skipped, with/goal skip(#289). Tools: invented URLs are refused (#288),edit_filegets a syntax check and hints (#269/#276/#283), WebGL errors carry hints (#253),session_searchcovers rollover segments (#287). Memory: the consolidation tile says what did not land and shows the full text (#285). Also: crow.log, context clearing (#262/#263). Docs are synced.Scope
Resolved issues (close on merge)
Resolved in part (stays open)
Already closed, shipped by this PR
Checks (head
7e7246f, Linux, crow venv)install.sh --selftesttailscale, sudo guard)install.ps1 -SelftestCHANGELOG
The phone becomes a second view of the session, and goal runs get a fresh-eyes judge. A paired phone mirrors
the open session on the home LAN and, through Tailscale serve, over HTTPS from anywhere, with its own microphone
(#249 stages 1 and 5, #290). Goal mode gains a judge that scores captures in a fresh request, an evidence gate for
visual steps, a render-loop breaker, retry-then-skip for failed steps and
/goal skip(#265–#268, #277, #286,#289). Long runs keep more of the window: old tool results are cleared below the rollover, and Crow's own status
lines move to
crow.log(#262, #263). The file tools share one syntax-check table and keep line endings (#269,#276, #283);
render_page,build_bundle, memory,session_searchand delegate each close a failure seen in the2026-09-23/24 diorama goal runs (#270–#275, #278, #279, #284, #285, #287, #288).
Everything on local
mainsince 2.5.0 (origin/main 0c3e392): 87 commits, 2026-09-23 to 2026-09-24. Mostmeasurements are against robin's 2026-09-23 diorama goal run and its 2026-09-24 follow-ups. Tickets stay open until
robin's live check unless the ticket says otherwise.
Added
/remote on(or the phone icon left of code/git/browser in thetitle bar) starts
cli/crow_remote.py, a standard-library HTTP server on one LAN address and a fixed port(
remote_port, default 8765;0.0.0.0is refused). The Remote dialog shows a QR code with a single-use 120 spairing code; a new device waits up to 60 s for Allow on the desktop, otherwise it is denied. A paired phone
keeps an HttpOnly, SameSite=Strict cookie for 400 days, of which only the sha256 is stored (devices in
secrets.json); 5 failed pairings lock pairing for 10 minutes, and revoking a device ends its stream at once.Host (421), Origin (403) and cookie (401) guards on every request. The phone gets the session over SSE (a
per-device replay ring of 5,000 events, snapshot otherwise, 15 s heartbeat, text coalesced to ≤ 4 updates/s) and
calls an allowlist of 58 proxied page methods; each client has its own view, so a phone never moves the desktop
and the other way round. The phone layer (below 700 px): a one-row composer, rail/code/git/browser as drawers, a
44 px goal bar, pinned approvals, image upload (20 MiB), its own home-screen icon and title, and an auto-hiding
header in the iOS home-screen app.
/remote on|off|status|devices|forget <name>; settingsremote_enabled,remote_port,remote_host. The QR encoder (byte mode, level M, versions 1–6) matches segno module for moduleon 20 URLs.
127.0.0.1:<remote_port>fortailscale serve --bg --https=443, where only the machine'sts.netHost andOrigin pass (a bare
127.0.0.1Host stays 421) and the cookie isSecure. Crow never runs a Tailscale commandthat changes anything: it reads
tailscale status --jsonandtailscale serve status --jsonwithout sudo intosix states (missing, down, https-off, serve-missing, funnel, ready), and the Remote dialog offers
HTTPS ·
<name>with one status line per state and the one-time serve command to copy (settingremote_https). Never plain HTTP on0.0.0.0or the 100.x address; with Funnel on for the name there is noloopback listener at all. Setup guide:
docs/user-guide/remote-tailscale.md.install.sh --tailscale/install.ps1 -Tailscalereadwhat is already done (read-only
tailscale status --json/tailscale serve status --json, through the samestate table as the Remote dialog) and print only the missing steps: the install command for the distribution
(
sudo pacman -S tailscaleon Arch/Omarchy and Arch-likes, the officialcurl -fsSL https://tailscale.com/install.sh | shelsewhere,winget install --id Tailscale.Tailscale -eon Windows),sudo systemctl enable --now tailscaled,tailscale up, Enable HTTPS in the admin console,tailscale serve --bg --https=443 http://127.0.0.1:<remote_port>(in an elevated shell on Windows) and the phoneapp when no iOS/Android device is in the tailnet yet. On Linux the section follows the install; on Windows
-Tailscaleprints it and exits without downloading anything. Neither installer runs sudo or elevates.install.sh --selftestdrives every state against a faketailscaleon PATH (39 checks);install.ps1 -Selftestdoes the same through a fake seam.and uploads the clip with the same guards; the window transcribes it with
crow_voice(same model and settingsas the desktop's dictation) and puts the words into that phone's input, never sent. Every 1.5 s the recording so
far is transcribed as a greyed partial; recording stops after 2 s of silence once speech was heard; the button
shows listening / writing / loading the speech model. One transcription per phone at a time, only the newest
partial waits, finals first. One
crow.logline per final dictation. Measured 2026-09-24 (faster-whisper smallint8, CPU): 5 s clip 1.25–1.4 s, 10 s 1.4 s, 15 s 1.5 s; model load from cache 0.85 s. Plain HTTP keeps the
keyboard-dictation hint.
judgetool sends images and a rubric to a vision modelin a fresh request with no history and stores the score in goal.json; without a pin it uses the local model. The
rubric leads with the step's own text ("delivers: "); the goal's
accept:lines, PLAN.md, the caller'scriteria or a default rubric only add to it (judge: a visual step can pass on a rubric the model picks -- step 3 'G-buffer and voxel volume' was accepted on step 2's structure/camera criteria (8/9/8) #286: step 3 "G-buffer and voxel volume" had passed on step 2's
camera criteria). The verdict records its rubric source. Live on the 2026-09-23 "9+" final frame: 4 vision models
scored it at most 2.
goal_step doneon a visual step needs a capture from this goal anda judge minimum of at least
judge_threshold(default 8;--judge-threshold, settingsjudge_threshold).force a rollover that carries what was tried. A capture with a luma mean ≤ 4/255 counts as stuck, and captures
are counted in every turn, including turns cut by a same-failure nudge or a mid-turn rollover (#268's loop breaker never fired on 9 near-black captures in step 2 (2026-09-24): a dark frame with a lit patch is not 'blank', and turns across a mid-turn rollover or behind a #202 nudge are never counted #277: on
2026-09-24, 9 near-black captures of one page, luma mean 1–2/255, got 0 nudges; the replay nudges at session
message 13 and forces the roll at 155). A black streak gets a concrete bisect (one pass in isolation;
getError,checkFramebufferStatus,readPixels)./goal skip(goal: a 'failed' step is sent back as the next step forever (step 4 looped ~40 min after the model failed it honestly), and robin has no /goal skip #289). The firstfailedsends the step backwith the failure note and asks for a different approach; the second marks it
skippedand the goal moves on. Agoal whose steps are all done or skipped ends
partial("complete with N skipped"), neverdone./goal skip <n> [reason]works in the window, on the phone and in the terminal. The goal bar shows skipped stepsand follows goal.json every round, before every turn and on
/goal, so a hand edit shows at once.enlarged content crop for
read_image. 2026-09-23: the scene filled 7–22 % of 62 renders.become a one-line stub, in batches that free at least 10 %; the originals go to
session/cleared/.context_clear_atin settings,--context-clear-at. Replay: rollovers 55 and 88 rounds later.~/.local/state/crow/log/crow.logwith local time and offset,rotated, instead of into the chat: goal brake, same-failure streak, degenerate round, cache, budget, context
clearing (Long goal runs roll over with a third of the window spent on tool results the model already acted on: no tool-result clearing below the 0.9 cut #263), the rollover line and The boundary refuses write_file, and the model reaches the same path with run_command in the next call #98's boundary alarm, chat-open and
/goalsetup notes, the "no folder"note, a browser-panel crash (The browser panel's WebKit process segfaults in the NVIDIA driver right after each GPU render_page: 512 MiB headroom ignores Crow's own panel as a second GPU client next to serve #279) and rejected memory writes (Memory Consolidation tile: 'save to memory' can drop a write without a word (refusal or 5-min expiry), and the tile shows only 160 chars and never the entry a replace removes #285). The roll card, the
/goalstatus answerand the Browser panel runs foreign pages in Crow's own web context: no memory kill (kill threshold 0), no sandbox, pywebview bridge injected, logins lost on restart #226 memory-ceiling stop stay in the chat.
edit_fileparses what it left (node for JS/HTML) and sayswhether this edit broke the file or it was broken before; settings
syntax_checksmaps an extension to a commandfor
write_file,append_fileandedit_file. Replay of 2026-09-23: 0 of 66 JS/HTML edits broke a file —parity, not a measured failure.
and total VRAM (probed once per process) and says a tool's limit is not the machine's;
memoryrefuses ano-GPU/CPU-only note that does not name the real card, and a "tool changes bytes" note. 2026-09-24: the diorama
run had saved "Machine has NO GPU" on an RTX 5090.
server holds it, and that the machine has the GPU. 2026-09-23: 49 of 49 captures fell back.
Changed
render_pageandfetch_urlrefuse a remote host that appears nowherein the conversation (the user's words, a tool result, the goal, PLAN.md) before any socket or browser opens, with
a next step; loopback always passes, a subdomain of a named host counts, and it is on in every mode, yolo
included. Across a rollover a Crow-written line carries the newest 60 hosts.
"Memory: not saved -- " in the window and the terminal, and goes to
crow.log. The Memory Consolidationtile has three states (collapsed, previews, full text) and shows a replace as "- old" above "+ new".
searchable, because Every rollover adds its archived segment to the chat rail as a separate chat, titled by Crow's own note ('[The tool budget for this turn is …') #261 took them out of the rail the index read. A segment hit is labelled
" (before the cut, )", every hit prints its path, and a search in the first turn of a new chat
no longer drops the chat just put aside.
A line opening with
/that is not a Crow command (a path) is queued too; the button reads↑ Queuewhile thebox holds a line; a Crow command waits in the box; an IME-confirming Enter is not a submit. Stop pauses goal
mode (Stop does not stop goal mode: run_turn consumes INTERRUPT before the pump asks, so the engine starts the next turn at once #282): Stop used to end one turn while the engine started the next at once; now the engine pauses with
a note until the next line you send.
oldwith the closest text (edit_file answers a missed 'old' with one bare line: 16 of 147 diorama edits failed, 4 of them on whitespace alone, and the model retried blind (4 in a row on mesh.js, goal brake twice) #276): line numbers, what differs, whitespace-onlynamed; a uniform indentation drift with one unambiguous window is applied and said. 16 of 147 edits missed on
2026-09-23/24; replayed, 3 now land and 11 of the other 12 get the region.
Crow's text on its left border (the
●column is gone). Headless Chromium 2026-09-24, 32 layouts: 716–780 px(user) and 40 px (Crow) off before, 0.0 px after.
×in the same place hides the card and its 306 pxcolumn reserve; nothing is cancelled, and a new subtask or a jump brings it back.
presence_penalty(serve: a request without temperature decodes greedy, and top_k 0 is greedy / top_k > 64 silently clamped on the device - absent fields must come from the model card row of the thinking mode crow-nest#111) no longer say crow-nest reads an absent field as 1.5(it is 0 since crow-nest The four behaviour changes cli/crow.py took in the split, and the gate each one stands on #91). No behaviour change: Crow sends 0.0 explicitly.
Fixed
(measured 2026-09-24 at 57ed521). The file is now read raw and only the matched span is replaced, in the ending
it had: 40/0 stays 40/0, a 20/20 mixed file stays 20/20.
(The browser panel's WebKit process segfaults in the NVIDIA driver right after each GPU render_page: 512 MiB headroom ignores Crow's own panel as a second GPU client next to serve #279). 2026-09-24: the panel's WebKitWebProcess segfaulted in
libnvidia-eglcore5–10 s after each of two GPUrender_pagecaptures (~560 MiB free), and after six crashes the window stopped repainting. render_page's GPUbound is now 1,536 MiB while the panel is open or holds a page (512 MiB otherwise); a folded or hidden panel
unloads its page and loads nothing; after a crash the dead view is hidden and the window redrawn, and a watched
page is reloaded once; during a local-model turn the panel renders without hardware acceleration.
as visual because of "verify"; the step's leading phrase decides now, and the refusal names the planning
exemption.
render_pageafter an edit replayed the old capture (render_page after an edit is answered from the per-turn _SEEN cache ('The result was, and still is:') with the pre-edit capture; the model learned to change width/height to get a real render #273).render_pagejoinsrun_command/build_bundleinNEVER_CACHED.render_pagedropped a local page's?query/#fragment(render_page refuses a local page with a ?query or #fragment ('no such page: index.html?view=albedo'), so the diorama prompt's required screenshot mode index.html?shot=<view>&w=<px> cannot be used and the model stored 'file:// rejects a ?query' in memory #272):index.html?view=albedocame back"no such page" (2026-09-24, session.json msg 36/37), and the model stored "file:// rejects a ?query" in memory.
The file is checked without the suffix, which goes back onto a percent-encoded
file://URL (RFC 8089 for driveletters and UNC).
~/.cache/deno, takesbundlerin settings.json or--bundler PATH, and its error names the settings file andthe
node_modules/.bin/esbuildlink and says anexportdoes not reach Crow. Measured 2026-09-24: findsesbuild 0.25.5 in
~/.cache/deno/dl/(was: none).decoded one more round (298–4,290 tokens, 11,240 / 271 s in total). A budget turn is now 24 tool rounds + the
answer (25 requests, was 26).
goal_step runningon the running step restarted its clock (goal_step 'running' resets the running step's clock and tokens #275): step 2's 73.5 min and its tokens vanishedfrom goal.json on 2026-09-24. A repeated
runningkeepsstartedand stores its note.(cold prefills of 106k). A turn that ran tools is never "empty".
reasoning is surfaced now.
"interrupted" now.
every subtask of the 2026-09-24 diorama wave spent 1 of 3 attempts on a spot that can never answer; it is a spot
refusal now (memoed, next spot).
up after ~6 s;
POST /pairanswers 202 and the page polls/pair/wait), a paired phone ignores a stale pairingcode, a phone page load no longer re-recorded every note (1 "no folder" note became 32), and Safari's toolbar
tint and scrolling stay as loaded after a drawer closes.
Known limitations
live on robin's iPhone (iOS 27, Safari and Chrome); the auto-hiding home-screen header and the drawer fixes were
checked in Chromium emulation only. Android is unverified.
/remoteis window-only; the terminal answers"stage 2".
sudoandinstall.ps1 -Tailscalewere not run on a Windows machine, andinstall.ps1 -Selftestwas not run for thisrelease (no PowerShell on the Linux release machine).
is refused.
check_gui_prereqs(not in CI) reports 2 of 3 prerequisites: point (ii) fails on 13 glyphs missing fromGoogle Sans Code (both shipped faces, 26 problems) — the 12 of 2.5.0 (0c3e392) plus U+1F3A4 🎤, new with the phone
microphone (Phone: the 🎤 records nothing -- getUserMedia needs a secure context and the mirror is plain HTTP; record on the phone over Tailscale HTTPS and transcribe with crow_voice on the PC #290, 4d753af;
cli/crow_core.py:18037,cli/crow_gui.py:8740). Measured 2026-09-24 at 0a0304b.sampling_no_thinking's presence_penalty 1.5 (presence_penalty 1.5 in the unused no-thinking sampling row #246).🤖 Generated with Claude Code
https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv