Skip to content

[Release] 2026-09-24: v2.6.0 — phone remote bridge over LAN and Tailscale HTTPS, phone microphone, safer goal mode, invented-URL guard - #291

Merged
nibor1896 merged 93 commits into
mainfrom
release-2026-09-24
Sep 24, 2026
Merged

nibor1896 merged 93 commits into
mainfrom
release-2026-09-24

Conversation

@nibor1896

Copy link
Copy Markdown
Owner

Summary

Release v2.6.0: 91 commits since origin/main (v2.5.0). Crow now runs on the phone. The remote bridge (#249) pairs over the LAN with a QR code, and away from home it runs over Tailscale HTTPS; the installers gain an optional --tailscale / -Tailscale step. The phone microphone (#290) shows a live transcript, stops on silence and appends to what is already typed. Goal mode got safer: the fresh-eyes judge rubric names the step (#266/#286), visual steps need evidence (#267), a render loop is broken off (#268/#277), and a failed step is retried once, then skipped, with /goal skip (#289). Tools: invented URLs are refused (#288), edit_file gets a syntax check and hints (#269/#276/#283), WebGL errors carry hints (#253), session_search covers rollover segments (#287). Memory: the consolidation tile says what did not land and shows the full text (#285). Also: crow.log, context clearing (#262/#263). Docs are synced.

Scope

Group Count
Close on merge (shipped and tested; live check noted in each) 9
Resolved in part, stays open 1 (#249: stages 2–4)
Already closed on live or measured evidence, shipped by this PR 32

Resolved issues (close on merge)

Resolved in part (stays open)

Already closed, shipped by this PR

Checks (head 7e7246f, Linux, crow venv)

Check Result
test_crow_remote + test_crow_core + test_crow + test_crow_gui 2,880 OK (8 skipped)
ruff clean
check_shared_core 83/83
check_operating_point 10/10
install.sh --selftest 39 checks PASS (fake tailscale, sudo guard)
install.ps1 -Selftest not run (no PowerShell on this machine)
check_gui_prereqs fails on 13 glyphs missing from Google Sans Code (12 known since 2.5.0, plus 🎤); not in CI

CHANGELOG

The phone becomes a second view of the session, and goal runs get a fresh-eyes judge. A paired phone mirrors
the open session on the home LAN and, through Tailscale serve, over HTTPS from anywhere, with its own microphone
(#249 stages 1 and 5, #290). Goal mode gains a judge that scores captures in a fresh request, an evidence gate for
visual steps, a render-loop breaker, retry-then-skip for failed steps and /goal skip (#265–#268, #277, #286,
#289). Long runs keep more of the window: old tool results are cleared below the rollover, and Crow's own status
lines move to crow.log (#262, #263). The file tools share one syntax-check table and keep line endings (#269,
#276, #283); render_page, build_bundle, memory, session_search and delegate each close a failure seen in the
2026-09-23/24 diorama goal runs (#270–#275, #278, #279, #284, #285, #287, #288).

Everything on local main since 2.5.0 (origin/main 0c3e392): 87 commits, 2026-09-23 to 2026-09-24. Most
measurements are against robin's 2026-09-23 diorama goal run and its 2026-09-24 follow-ups. Tickets stay open until
robin's live check unless the ticket says otherwise.

Added

Changed

Fixed

Known limitations

🤖 Generated with Claude Code

https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv

nibor1896 and others added 30 commits September 23, 2026 22:28
… turn is not an empty loop (#258)

goal_last_answer now returns a copy of the last answer carrying every tool
call of its turn (BUDGET_SPENT / TOKEN_BUDGET_SPENT / THINK_ONLY_NUDGE do
not open a turn). On robin's 2026-09-23 session the brake read the forced
round's empty content as "empty" and cut 157 + 52 messages of real work,
each cut a 106k-token COLD re-prefill. Replay on session.json: before
empty=True, cut at 212; after 24 calls seen, empty=False, no cut.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
… empty answer

After BUDGET_SPENT the forced round came back reasoning-only (engine.log
2026-09-23 20:14:38Z: content chunks 0, reasoning chunks 710, finish
tool_calls): the window showed nothing, and goal mode's brake judged the
24-round turn by its last message -- three such turns cut 157 messages of
work, and the next stopped the goal at 6/9.

- run_turn: a final round (forced, or past the #150 nudge) with no visible
  text stores and shows its reasoning under SURFACED_MARK (tail kept past
  6000 chars); with no reasoning either, a line counting the calls that ran.
  No extra round, no prompt change; reasoning_content stays stored.
- goal_turn_empty / goal_turn_mark: a turn with any tool call or result is
  never empty, and the echo fingerprint covers the whole turn's calls. A
  surfaced stand-in still counts as empty, so #202's loops still brake.
- crow_gui _goal_brake uses both.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
A mid-goal goal_set reported "carried: done 6/9" and left every note in
goal.json empty: _carry_goal_marks copied `status` only, so the evidence
notes, seconds and tokens -- and the goal's spent -- were reset, and #250's
"done after failed needs a note" lost the note it quotes.

Carried now for matched done/failed steps: status, note, started,
started_tokens, seconds, tokens, delegated (text stays the new plan's);
and for the goal: spent with last_spent, delegated, created, started with
last_at -- in pairs, so the next step is not billed the whole goal.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…row notes are never a title (#261)

A rollover is the live chat's earlier half, not a chat of its own. The rail
no longer lists rollover-*.json (except the one open in the window); the
archive drawer lists them in place, files untouched, no "restore" row.
Titles come from crow_core.chat_title: every Crow user-role note (rollover,
budget, token budget, think nudge, abort, goal nudge) is skipped, the #224
root notice is stripped, and a segment opening with a rollover note carries
the previous archive's title, so the live chat keeps its name across cuts.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
robin, 2026-09-23: the grey machinery lines clutter long sessions. Moved to
a new rotating log, <state_dir>/log/crow.log (~/.local/state/crow/log/crow.log
on Linux, %LOCALAPPDATA%\Crow\log\crow.log on Windows; 1 MiB x 3 backups,
local time with UTC offset, one [kind] per line):

- goal mode: N messages of an empty loop dropped from the history (_goal_cut)
- goal mode, step N: the same failure keeps coming back -- ... (the nudge
  to the model is unchanged)
- discarded a degenerate reply / kept the re-asked reply (#217, run_turn)
- the restored cache did not hold (window and terminal)
- tool budget spent after N rounds (terminal)

None was ever in `messages`. One classifier (note_is_log_only) guards
Api.push, the terminal's turn_note and clean_notes, so a reopened chat
saved by an earlier build no longer draws them either. Kept in the chat:
the rollover card and "rolled over at N tokens", goal mode stopped/paused,
server reboot/load-wait lines, alarms, Memory updated, carried across the
cut, and every answer to a command.

Replayed over the live state (read-only): notes band 16 -> 10, exactly the
six machinery lines removed. test_crow_core 1320 OK (3 skipped), test_crow
452 OK (3 skipped), test_crow_gui 760 OK (2 skipped), ruff clean.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
A third of every full window on robin's 2026-09-23 diorama run was tool
output the model had already acted on (33.0-39.9 % of the request tokens,
97 % of it behind the last five rounds; rendered with the model's template
and tokenizer, within 0.05 % of the server's count). The only response to a
full window was the 0.9 cut.

At 0.65 of the window (default on the local server, off remote,
settings.json `context_clear_at` / `--context-clear-at`, 0 = off) every tool
result older than the last 5 rounds becomes a one-line stub naming the tool,
its target, the size and how to get it again; the originals go to
session/cleared/<stamp>.json first. Assistant messages, reasoning, tool
calls and the head stay byte-identical. A batch runs only if it frees at
least 10 % of the window, so the prefix cache pays one re-prefill per batch,
not per round. One line per batch to logs/crow.log, nothing in the chat.
A cleared read_file drops its #215-H read stamp; the goal counter rebases.

Replay through clear_before_roll on the chained run (790 messages): the
two rollovers move from message 188/493 to 299/669 (55 and 88 rounds later)
for 2 batches, 158,541 re-prefilled tokens; each recorded context alone
would not have reached the cut.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…d line has priority again

go() sent every Enter during a running turn to stop() since 4860300, so
#165's "a typed line always has priority; the engine only runs when the
queue is empty" was unreachable from the live chat. A non-empty plain line
mid-turn now goes through the existing send() -> _queued -> _pump path: the
page holds it (hint "queued -- it goes in when this turn ends") and draws it
at the turn's idle; _pump runs it ahead of _goal_nudge and now calls
_goal_reset for it. send() joins a second queued line for the same chat
instead of replacing the first. The Stop button (press()) and Escape stay
Stop; an empty Enter and slash lines other than /delegate, /subtasks too.

test_crow_gui 765 OK (2 skipped; 9 new, 6 red on 0c3e392, 3 counter-probes),
test_crow_core 1315 OK (3 skipped), test_crow 451 OK (3 skipped), ruff clean.

refs #264

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…overage (#265)

The diorama filled 7-22 % of each render on a grainy black ground; at
~1,030 px per visual token the model graded it from ~40-120 tokens.
read_image now adds the content box, enlarged to a 1024-px long edge, as a
second image and states the coverage; render_page prints metrics: lines
(box, coverage, colours, luma), warns under 25 % coverage or 16 colours and
saves <shot>-crop.png. Pure Python on the #213 PNG decoder.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…e judge's bar (#267)

2026-09-23 diorama run: 9/9 done over a small box in a black frame; the final
notes replayed through goal_done_refusal at 0c3e392 were 9/9 accepted. On a
visual goal (keyword guess, goal.json "visual" overrides) a done must now cite
an image written after goal.created, and when a judge verdict is stored on the
step its lowest score must reach JUDGE_THRESHOLD (default 8; settings.json
judge_threshold, --judge-threshold). Planning-only steps are exempt. Replay:
6 of 9 refused. The #250 tests move to a non-visual title.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
2026-09-23: the model scored its own black-frame capture 9+ on every
criterion. New tool `judge`: one request with the capture(s) (newest
render + its -crop.png if present), the rubric and the goal/step text --
never the transcript. Rubric order: /goal `accept:` lines (stored beside
check: in SESSION_DIR), PLAN.md criteria, the caller's, a default visual
rubric. Spot order: providers.json `judge` pin, the delegate spot and its
fallbacks (catalogue rows now keep `vision` from input_modalities; known
blind rows skipped), then this turn's own endpoint in a fresh request,
refused when /props is blind. JSON scores, min, weakest 3, stored on the
goal step for #267's gate. Live 2026-09-24 on render-20260923-232855.png:
four free OpenRouter vision models, min 2 each.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…lover at 6 (#268)

2026-09-23 22:30-22:48: seven chain.html captures in a row at 99.2-99.4 %
one colour while every turn called tools; neither #202's brake nor its
trouble classes fire on captures. goal_render_scan counts, per page, the
render_page captures that are blank-ish (#213/#175 warnings, the crop
branch's distinct-colour warn, a decoded share >= 0.98) or near-identical
to the previous capture of that page (metrics: lines within 3 %/2 when
present, else share within 0.005 and size within 3 %). At 3: one nudge to
bisect (clear colour, one cube, passes one by one). At 6: _run rolls on a
flag, and the carry lists each stuck capture and the files written before
it; one forced roll per step. _goal_brake is untouched (#258/#259).
Replay of the four 2026-09-23 transcripts: nudges before messages 35 and
295 (rollover-210418) and 314 (rollover-225111); no roll fired.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…ve one (#268)

Found by running the breaker over #265's real metrics lines on three
2026-09-23 captures.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
# Conflicts:
#	cli/crow.py
#	cli/crow_gui.py
#	justfile
… test counts

t-judge's loop breaker pushes "goal mode, step N: ... the context rolls
over" -- a prefix #262 routes to crow.log. The test now expects it there
and not in the chat.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
… and the spot is not memoed dead

_run_subtask read sub.cancelled only after the verdict: an attempt that
came back "failed" after the user's Stop was written into the chain,
marked the spot dead in _SPOT_DEAD and closed "failed". The mark is now
read right after each attempt, before anything is decided.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
… serves all three file tools

edit_file ran no syntax check (#251 covered write_file/append_file only).
It now runs syntax_check after the edit; on a failure the pre-edit text is
parsed in a scratch copy, and the result says whether this edit broke the
file or it was broken already. settings.json `syntax_checks` maps an
extension to an argv ({path} = the file); an empty row switches a built-in
off, malformed rows are dropped one by one. Replay of the 2026-09-23
diorama run: 0 of 66 JS/HTML edits broke a file -- this is parity and
one place for new formats, not a measured failure.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…ct it, and render_page says why it rendered in software

2026-09-24 the diorama run saved "Machine has NO GPU" into its project
memory (the card is an RTX 5090) and repeated it three times; 49 of 49
render_page captures on 2026-09-23 fell back to SwiftShader while serve
held the card, and the result never said why.

#270: crow_platform.machine_facts() (OS, CPU, RAM, GPU name + total VRAM,
static, once per process) as a head line after the working area, with the
rule that a tool's limit is not the machine's; memory_contradiction()
beside memory_threat refuses no-GPU/CPU-only notes that do not name the
card and "a tool changes bytes" notes, on add and replace.
#271: render_gl_reason() appended to a software render result.

Tests: 6 new, 5 red without the fix (the sixth is the negative probe);
2,633 OK (skipped=8), ruff clean.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…format, rotated

crow_log() wrote <state>/logs/crow.log in bare UTC without rotation, while
#262's log_note() writes LOG_FILE = <state>/log/crow.log with local time
and offset, rotated; both were live on 2026-09-24 (10:17:56 context-clear
in one, 10:36:26 goal line in the other). crow_log is log_note(line,
"context") now; window.md and settings.md name the one file.

Tests: the two clearing tests read LOG_FILE and assert no second log;
red without the fix (FileNotFoundError on log/crow.log), green with it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…s note is stored

2026-09-24 diorama run: the engine began step 2 at 10:07, the model re-reported
it running at 11:21 (session.json msg 148), and goal_step_begin reset `started`
without billing the open window -- goal.json read seconds 0, tokens 0, and the
call's note ("... NOT verified ...") was dropped while the tool answered ok.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…ml?view=albedo was "no such page" on 2026-09-24 (session.json msg 36/37) and the model stored "file:// rejects a ?query" in memory

tool_render_page handed the whole argument to os.path.isfile, so the
diorama prompt's required index.html?shot=<view>&w=<px> could never be
rendered. _page_target now checks the whole path first (a file really
named with ? or # wins), else the part before the first ?/#, and
re-attaches the suffix to a percent-encoded file:// URL built by
_file_url (drive letter and UNC per RFC 8089, identical to
Path.as_uri() on POSIX; a pure function so the NT form is testable on
Linux, PureWindowsPath.as_uri/nturl2path are deprecated in 3.14). A
missing page names the FILE. _source_line (#253) takes the path part of
a console source URL, which now carries the query. #175's
_LAST_CAPTURES and _cache_key already key on the full URL/arguments:
two views are two pages (pinned by a test). Tool description and
docs/reference/tools.md say queries work.

Tests: 7 new (ALocalPageKeepsItsQueryTests), 7/7 red on 0c3e392
(6 fail, 1 AttributeError), green with the fix. test_crow_core 1322 OK
(skipped=3), test_crow 451 OK (skipped=3), test_crow_gui 756 OK
(skipped=2), ruff clean, check_shared_core 82/82, check_operating_point
10/10. Not verified: a live browser render with a query.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
nibor1896 and others added 18 commits September 24, 2026 20:04
…le shows the full text

approve_pending now returns (saved, failed). A tool refusal (over the
limit, no match, contradiction, threat), a duplicate (success without an
action) and an expired entry each come back in `failed` with a short
reason instead of being swallowed; expired entries are kept aside in
_EXPIRED at drop time so the answer can name them. Each miss is logged to
crow.log. answer_memory pushes a "Memory: not saved -- ..." note (the glow
still fires for what was saved); the terminal prints the same line.

pending_view adds `full` (untruncated content) and, for replace/remove,
`old` (the whole entry old_text matches right now, read-only; None when
there is no single match) plus `find`.

The Memory Consolidation tile gets three states: collapsed, previews,
full text (a replace as "- old" above "+ new", an add marked "appended
at the end"). Same page on the phone layer; nothing in REMOTE_CSS hides
the new rows.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…#285 tile, head cost; architecture counts incl. crow_remote.py

Audit of 2026-09-24 against 3e0ccc2: the model's own `memory` calls
write unasked at every level (only the review is gated); the tile has
three states and says what did not land (#285); the 633-char head-cost
example and "byte 0" were outdated.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
index_sources() took its list from the rail, and #261 keeps rollover
segments out of the rail, so 0 of 6 segments on disk were searchable.
The list is split: rail_sources() is the chats of their own (live
session.json, session/chat-*.json, archiv/chat-*.json), index_sources()
adds session/ and archiv/ rollover-*.json. A segment hit is labelled
"<chat title> (before the cut, <date>)", and the tool prints each hit's
path for a following read_file.

The missing session.json row: "neu" writes the open chat to
session/chat-*.json and deletes session.json until the new chat's first
turn ends. A search in that turn dropped session.json's rows as gone and
never saw session/chat-*.json. index.db was last written 14:02:40.580Z,
19 ms after the first reply of a fresh chat asked for two tools.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
Without the user's accept: lines, judge_rubric(criteria, step_text) puts
"delivers: <step text>" first; PLAN.md, the caller's criteria or the
default rubric only add to it and can no longer replace it. tool_judge
passes the text of the step _judge_step picked, and the stored verdict
records rubric_source and the criteria as judged. Step 3 'G-buffer and
voxel volume' had passed on step 2's camera criteria the model re-sent.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…; the bar follows goal.json

- goal_step 'failed' counts failures on the step: the first sends it back
  once (answer carries `retry`, the next nudge quotes the failure note and
  asks for a different approach), the second marks it `skipped` and
  next_step moves on.
- goal_next_open, the seam head and goal_set's `next` pass over skipped
  steps; a goal whose steps are all done or skipped ends `partial`
  ("complete with N skipped"), never `done`.
- /goal skip <n> [reason]: one core command for window, phone and
  terminal; a done step is not skipped, a running one is billed first.
- Goal bar (desktop and the thin phone bar): skipped mark, note as tooltip,
  "step N skipped" / "Complete · step N skipped" in the head.
- The bar follows goal.json: every round (only when a step or the goal
  changed state), before every turn, and on `/goal`. A hand edit stayed
  hidden until the end of the next turn because push_goal ran only after a
  turn and on goal_set/goal_step results.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
A remote host must already appear in the conversation: the user's words
(URL or bare name), a tool result (URL), the goal's title/steps or PLAN.md
(read from disk at the miss). Otherwise the call comes back as
"error: refused: <host> appears nowhere in this conversation -- a URL you
made up?" plus a next step, before any socket or browser. Loopback always
passes; a subdomain of a named host counts. Always on, yolo included.

The set is rebuilt per turn from the conversation (like _MANDATED) and fed
after every tool result in run_turn. Across a rollover the note carries the
newest 60 hosts in one Crow-written line ahead of the digest, and the next
turn re-derives the set from the new context; the model's digest cannot
carry a host.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…, dialog choice

crow_remote: tailscale_state() reads `tailscale status --json` and
`tailscale serve status --json` (no sudo) into missing / down / https-off /
serve-missing / funnel / ready. With a ts.net name the server additionally
binds 127.0.0.1:<port> (the `tailscale serve --bg --https=443` target) where
only Host <name>[:443] and Origin https://<name> pass; a bare 127.0.0.1 Host
stays 421 as before. The cookie gets Secure on that origin only. Never
0.0.0.0 or the 100.x address in plain HTTP; Funnel on the name means no
loopback listener at all. A failed loopback bind leaves the LAN mirror up.

Window: the Remote dialog offers "HTTPS · <name>" as a network choice
(remote_use_https, desktop-only, setting remote_https) with one status line
per state and the one-time serve command to copy; its QR carries the same
single-use code on the ts.net origin once serve points here. /remote and the
titlebar icon re-probe and restart once when the tailnet came up after start.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
Phone page (secure context + mediaDevices + MediaRecorder): the mic records
on the phone -- first tap starts, second tap stops, a level ring from an
AnalyserNode -- and uploads the clip to /upload?kind=audio with the same
cookie, Host and Origin guards. The server writes it and answers 202 at once;
the window transcribes it on a thread with crow_voice.transcribe_file()
(same model and settings as stop(); PyAV decode via faster-whisper's own
decode_audio) and pushes {"k":"heard"} to THAT phone only, which puts the
words into its input field, never sent. The clip is deleted either way.
Plain HTTP keeps the keyboard-dictation hint, reworded as in #290. The
desktop's own dictation (its "mic" pushes) no longer lands on a phone.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
docs/user-guide/remote.md: /remote, pairing, the LAN and HTTPS addresses,
the phone microphone (#290), guards, troubleshooting. remote-tailscale.md:
cost and privacy, phone app, per-OS install, admin console (MagicDNS,
Enable HTTPS, CertDomains check), the one-time `tailscale serve` line, the
steps in Crow, checks, a troubleshooting table per dialog status line, undo.
Commands cited against the Tailscale KB; what could not be verified is
marked. Settings reference gains remote_https; docs index and README link
both pages.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…erver not found, login 403, two pairing origins)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…one log line per phone dictation

Scope amendment after robin's first live test on the HTTPS address: the ring
moved, nothing arrived (the second tap was not obvious), and the first
dictation spent ~52 s loading the model with no word on the phone.

- Phone (REMOTE_JS section 6): MediaRecorder timeslice 1500 ms; every chunk
  uploads the recording so far as /upload?kind=audio&partial=1&seq=N. The
  partial shows greyed after the text already in the field (field read-only
  while provisional); the final replaces only the partial. Every dictation
  APPENDS (one space unless empty or ending in whitespace). Auto-stop 2 s
  below the level after >200 ms of speech; silence before speech stops
  nothing. Hints: listening / writing / loading the speech model. The button
  shows a filled stop square while recording, ring kept, 44 px target.
- Server: /upload parses partial and seq (a partial without seq is 400).
  Api._remote_audio keeps one lane per phone: one transcription at a time,
  only the newest waiting partial, finals first, partials not newer than the
  last final dropped (waiting, running or late). {"k":"heard","loading":true}
  while crow_voice.model_loaded() is False. One crow.log line per final
  (kind voice): bytes, seconds, transcribe ms, error.
- Measured (faster-whisper small int8, CPU): 5 s clip 1.25-1.4 s, 10 s 1.4 s,
  15 s 1.5 s; model load from cache 0.85 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…TPS setup as printed steps

Opt-in installer component for #249 stage 5. Both installers read where the
machine stands through read-only `tailscale status --json` and `tailscale
serve status --json` (the same state table as crow_remote.tailscale_state)
and print only the missing steps: the install line (pacman on Arch and
Arch-likes, kb/1031's script elsewhere, winget Tailscale.Tailscale on
Windows), systemctl enable --now tailscaled, tailscale up, Enable HTTPS in
the admin console, `tailscale serve --bg --https=443
http://127.0.0.1:<remote_port>` and the phone app links. Neither runs sudo
or elevates. On Windows -Tailscale prints and exits without downloading.

install.sh --selftest: 39 checks (was 22), the probe driven against a fake
tailscale on a PATH holding nothing else, with an argv log proving only
the two read calls ran. install.ps1 -Selftest: the same states through a
fake seam (not run here: no PowerShell on this machine).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
- README: optional-components table (--voice, --tailscale / -Tailscale),
  Tailscale download/App Store/Google Play links in the Phone row; 28 tools,
  #288 host refusal, #289 retry-then-skip and /goal skip.
- install.md, linux.md, remote.md, remote-tailscale.md (new section 0: the
  installer route and every download link), settings.md, docs index.
- window.md crow.log row and dictation model row, remote.md phone clips,
  goals-and-subagents.md #284 refusal, architecture.md line counts (wc -l at
  b9cac62), testing.md and justfile: the fourth suite (test_crow_remote) and
  re-measured case counts.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
VERSION 2.6.0 in cli/crow.py, install.ps1's -Version default, the README
badge and manifests/operating-point.json. CHANGELOG: Unreleased becomes
2.6.0 — 2026-09-24 (release summary; Added / Changed / Fixed / Known
limitations, every item with its issue numbers), new empty Unreleased.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
… (13 glyphs incl. U+1F3A4)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…a drive-letter file:// source

_TROUBLE_PATHLIKE knew only '/' paths, so two identical edit_file refusals
on D:\...\a.py and b.py keyed apart (EditFileMissTests on windows-latest).
_source_line read 'file://D:\x' as a UNC host; the test built that URL by
hand and now uses _file_url, and the code accepts a drive-letter netloc.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
…n Windows too (TimeoutError after SYN retries)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmVJHThNzKH4NWHvQffwYv
@nibor1896
nibor1896 merged commit 21d7505 into main Sep 24, 2026
4 checks passed
This was referenced Sep 24, 2026
@nibor1896
nibor1896 deleted the release-2026-09-24 branch September 24, 2026 22:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Memory Consolidation tile: 'save to memory' can drop a write without a word (refusal or 5-min expiry), and the tile shows only 160 chars and never the entry a replace removes The browser panel's WebKit process segfaults in the NVIDIA driver right after each GPU render_page: 512 MiB headroom ignores Crow's own panel as a second GPU client next to serve The model believes the machine has no GPU (RTX 5090) and saves it to memory: nothing in its context says what the machine is, and the memory gate stores contradicting facts Goal mode accepts 'done' without evidence: 9/9 complete over a page that draws nothing, with step notes saying 'done-with-deviation only in spirit' and a failed step flipped to done with no note delegate: a subtask stopped while its spot is failing ends 'failed' and memos the spot dead, instead of 'interrupted' (#143 E2 contract)

1 participant