Focus stops hiding the agent's work: folded activity names what it did - #658
Merged
Merged
Conversation
… how much
"In Focus mode I can't see anything" was argued from code for weeks. A live
capture settles it: one turn searched the web and dispatched a subagent, and
the moment it settled the transcript read
Thought process · 2 steps ⌄
<the answer>
Nothing on screen recorded that a search or a subagent had happened. In Focus
the Run-details column is folded too (deliberately, e306c08), so that one muted
row was the ENTIRE record of the agent's work. Studio folds identically — the
transcript is a single shared mount — but there the Agents tab and the centre
column mask it.
Why it never reproduced in the journeys: `absorbThoughtActivity` folds tool and
fleet cards into the reasoning span that produced them, and `ThinkingBlock`
closes when `autoOpen = failedCount > 0 || activityRunning` goes false. A model
that emits no reasoning never enters that path, so its cards stay top-level
rows. Every journey ran such a model. A real user on Claude Opus gets the fold
on every turn.
The fold itself is right and stays. PRD-03 ("the answer wins") exists because
six stacked cards once took 55% of the transcript and buried a one-line answer.
But PRD-03's own reference table shows the half that got dropped: Codex folds
to `Worked for 26s ›`, while Claude Desktop keeps a NOUN — `Read depth.py ›`.
We copied the fold and lost the noun. The code even knew the risk: "the count
is the only thing that tells a reader work is folded … without it, absorbing
the tool cards would BE hiding them." A count says work exists; it takes a noun
to say what it was.
So the label learns to talk:
Thought process · Searched the web · Dispatched 1 subagent ⌄
Worked for 8.3s · Read 3 files · Edited 1 file ⌄
`describeActivity` is pure and feeds BOTH fold sites (the thought block and the
PRD-03 run group). It asks the tool registry first, because that is the
product's one statement of what a tool is; counts FILES not calls, so three
edits to one file read "Edited 1 file"; keeps a connector call's own name
rather than calling it "1 step"; reads in first-occurrence order so the line
tells the run as it went; and folds the long tail into "+N more".
Two layout details that a green suite would not have caught:
* The digest is the one unbounded string in a `width: fit-content` header, and
the Studio chat column is ~335px. It is therefore the item that yields
(`flex: 0 1 auto`, ellipsis, full sentence on `title`) while the label and
the chevron stay `flex: none` — otherwise the row pushes its own chevron, the
only affordance the control has, off the edge. Asserted as a shrink CONTRACT
because jsdom runs no layout.
* The count was drawn in `--color-text-subtle` (~3.3:1 here, under AA), which
was fine for decoration. A sentence carrying the record of the work is
content, so it takes the label's `--color-text-muted` (~7:1).
The count survives as `data-steps` for anything that needs the number, and as
the fallback label, so a caller that cannot describe its members still never
folds them silently. `Worked for` stays the group's prefix — the TR-16 journey
pins it.
Verified on the real screen, before and after, not only in jsdom.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…gn and our own splash do "The fonts look off" has four candidate causes in TYPOGRAPHY-PROPOSAL.md, and this is the one that is certain, face-independent and two lines. Every design mock sets `-webkit-font-smoothing: antialiased` on body (design-kit/app-v3/copilot.css:110). So does our own boot splash, on `.boot` (apps/desktop/renderer/BootProgress.css:34). The app body did not — so text rendering visibly CHANGED the moment the splash unmounted, inside one React root: stems went from the splash's thin grayscale to Blink's heavier default. On a dark theme that default reads as slightly bold and slightly smeared, which is a fair description of the complaint and has nothing to do with the typeface. The proposal marked this option's visual payoff "NOT verified — I did not run the app and did not measure a pixel." It is verified now: a like-for-like 1:1 crop of the same Focus transcript, before and after, shows light-on-dark stems thinning to their designed weight. Captured from the real Electron window, not headless Chromium and not jsdom. macOS-only properties; inert on Windows, where ClearType owns this. Holds whichever way the proposal's A-vs-C typeface decision goes, which is why it can land ahead of that decision. Revert = delete two lines. Deliberately NOT included: B2, the whole-pixel size ladder. It moves 450 call sites by 0.2-0.48px, the design uses BOTH 12px and 12.5px so retuning `xs` swaps which group is wrong, and the proposal itself says not to merge it without a parity run. That wants its own change and its own look at the app. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… a reasoning model
`transcript_rendering` was green only by accident of which model ran it. Every
phase was written against a model that emits no reasoning, where a tool or
fleet card is a top-level transcript row. On a reasoning model — what a real
user on Claude Opus has — `absorbThoughtActivity` folds every card into the
"Thought process" block, which closes when the run settles. Pointing the suite
at Anthropic turned TR-1, TR-2, TR-7 and TR-8 red with messages that all read
like product bugs:
no required new fleet card appeared … 'text': '', 'fleetStatus': 'done'
page.click: Timeout … element is not visible
fleet card missing 'Dispatched' copy: ''
`innerText` is layout-aware, so a card inside a closed disclosure reads `''`,
and Playwright will not click what has no visible box. Same root cause as the
TR-3 fix two PRs ago, now handled once instead of per phase:
* `JS_WITH_REVEALED` — open every disclosure above a node, read, RESTORE, all in
one evaluate. Reading `innerText` forces the layout flush, so the text is real
and nothing repaints. Restoring matters: TR-16 asserts every settled group IS
collapsed, and a snapshot that left older groups open would fail it three
phases later with no trail back. `card_snapshots` and both Focus readers use
it.
* `reveal()` — the mutating sibling, for phases that go on to CLICK. A user who
wants to operate a folded card opens the fold first. TR-1's first click and
`ensure_native_disclosure` use it; `ensure_fleet_card_interaction` drops its
private copy.
Two more stale assumptions, both in the Focus readers, both from when each was
its own journey with its own boot:
* **They took the FIRST card in the DOM.** Grouped phases share a conversation,
so by TR-8 the first fleet card is TR-2's, six turns old. They now skip the
cards that existed before the phase sent and take the newest — selecting card
HOSTS (`[data-tool-status]`), because the bare testid prefix also matches a
card's own `…-details` child and "the last match" was a fragment of a card.
* **TR-7 learned which tool ran by reading copy.** The title is model-authored
now ("math.isqrt docs"), so the raw name is on no pixel. `ToolCallCard` stamps
`data-tool-name` on its root, beside `data-tool-status` and for the same
reason, on both render arms.
Result on Anthropic (a reasoning model): TR-1, 2, 3, 6, 7, 8, 9 all pass.
Not fixed here, and worth knowing before trusting a run: the DEFAULT journey
provider is Virtuals, and its catalog currently preselects Gemini 2.5 Pro,
which fails every run with `external_service_error` ("Service unavailable").
That is a gateway/default-model problem, not a transcript one — run with
`JOURNEY_PROVIDER=anthropic` until it is looked at.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…e 2026-08-27 `build-and-audit` is red on dev, and has been since something that is not a code change: `ci-frontend` was GREEN on 2026-08-27 (ed72f5c) and is red on today's dev HEAD (7193fbc), with no dependency edit in between. New advisories were published against versions already pinned. The gate is `npm audit --audit-level=high`, so the 6 moderates do not matter and these 8 do: critical astro <=7.2.7 RCE via AVIF image optimization high sharp <0.35.4 libheif GHSA-g89c-p67h-r497 / -2jg2-4ch7-h545 high svgo 4.0.0-4.0.2 removeScripts sanitiser bypasses (x2) high smol-toml <=1.7.0 DoS via malformed TOML high js-yaml 4.0.0-4.3.1 maxTotalMergeKeys does not bound CPU high browserslist <=4.28.6 unbounded cache growth -> OOM high fast-uri 3.0.0-3.1.5 host confusion via skipped IDN canonicalisation high @xmldom/xmldom <=0.8.14 XML fragment injection via EntityReference Every move is a patch/minor INSIDE the existing ranges, so no package.json changes and no new direct dependency: astro 7.2.0 -> 7.3.3 (satisfies ^7.2.0), sharp 0.35.3 -> 0.35.4, svgo 4.0.2 -> 4.1.0, smol-toml 1.7.0 -> 1.8.0, js-yaml 4.3.1 -> 4.3.2, browserslist 4.28.6 -> 4.29.0, fast-uri 3.1.5 -> 3.1.8, xmldom 0.8.13 -> 0.8.15. All of it is the `apps/website` (Astro) tree — nothing in `apps/desktop` or `packages/*` moves. WHY `--package-lock-only` AND NOT `npm audit fix`. This repo's lockfile does not survive a full re-resolve (peer conflict in the electron-builder tree plus website hoisting), and a regenerated lockfile reds the SBOM gate. A targeted `audit fix --package-lock-only` touches only the advisory subtrees and leaves node_modules alone. Verified stable: `npm ci` succeeds against the result and does NOT mutate the lockfile further, which is the property a regeneration breaks. On the SBOM step that runs right after the audit: `cyclonedx-npm` fails LOCALLY on both this lockfile and the one before it, identically — `npm ls` reports 26 pre-existing peer problems (@electron/asar, @electron/universal, ejs, all from app-builder-lib) on a clean install of either. So it is not something this change introduces, and the step passed in CI as recently as 2026-08-27, which is the run this commit restores the job to. Unrelated to the rest of this PR; it is here because the PR cannot go green without it while dev is red. Verified: audit gate passes, `npm ci` clean, no manifest changes, chat-surface and desktop typecheck at 0 errors, desktop builds, and the website — the only tree that moved — builds on the bumped Astro. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…e goes green
`build-and-audit` is red at "Generate SBOM". `cyclonedx-npm` shells out to
`npm ls --json --long --all`, which exits ELSPROBLEMS with 7 extraneous and 3
invalid packages:
extraneous @img/sharp-wasm32 + its @emnapi/*, @napi-rs/wasm-runtime,
@tybys/wasm-util subtree — orphaned, nothing declares an edge
invalid app-builder-lib/node_modules/{@electron/asar@3.4.1,
@electron/universal@2.0.3, ejs@3.1.10} — the pre-override
versions the root `overrides` block exists to eliminate
CORRECTION TO 15c297c. That commit reports the SBOM breakage as pre-existing
and "not something this change introduces". Measured on CI's own toolchain that
is not so, and the difference is the npm version:
dev lockfile 15c297c lockfile
npm 10.9.8 (CI) npm ls exit 0 exit 1 — 7 extraneous, 3 invalid
npm 11.9.0 (local) exit 1 — 11 probs exit 1 — 10 probs
Under npm 11.9 both trees look broken, which is what made it read as
pre-existing. Under the npm CI actually runs — 10.9.8, bundled with the
node 22.23.2 the workflow pins — dev is spotless and only the new lockfile
fails, reproducing the CI error package-for-package. The last green SBOM run
(2026-08-27, ed72f5c) uploaded its artifact off that same dev lockfile, which
is the authoritative check the local one contradicted.
The mechanism: 15c297c used `npm audit fix --package-lock-only` on npm 11.9,
whose resolver nests un-overridden copies under app-builder-lib and leaves the
stale sharp wasm32 subtree behind when sharp 0.35.4 drops it. npm 10.9.8 then
calls both out.
Note that npm 11.9 was not a gratuitous choice: npm 10.9.8 CANNOT run
`npm audit fix` in this repo at all — it dies on `"playwright": "$playwright"`
with `Unable to resolve reference $playwright`. Plain
`npm install --package-lock-only` and `npm update` are unaffected, so the
advisory work is done with those instead. Worth replacing that `$playwright`
reference with its literal version if we want `audit fix` back on npm 10.
WHAT THIS DOES. Starts from dev's clean lockfile and re-applies the same 8
bumps on npm 10.9.8 — `npm update` for seven, and a re-resolve inside the
unchanged `^7.2.0` range for astro. Same end state, no nested override
violations, no orphan subtree: astro 7.3.3, sharp 0.35.4, svgo 4.1.0,
smol-toml 1.8.0, js-yaml 4.3.2, browserslist 4.29.0, fast-uri 3.1.8,
@xmldom/xmldom 0.8.15. It is 303 lines SMALLER than the lockfile it replaces,
because it drops that subtree rather than adding to it.
Only package-lock.json moves; every package.json is byte-identical. astro now
resolves at apps/website/node_modules/astro rather than the root hoist — npm
10's own placement, not a hand edit.
Verified on node 22.23.2 / npm 10.9.8, the workflow's pinned pair: `npm ci`
clean, all six package typechecks at 0 errors, frontend builds, `npm audit
--audit-level=high` passes, `npm ls --json --long --all` exits 0, the SBOM step
produces a valid CycloneDX 1.6 document with 869 components, and the website —
the only tree that moved — builds on the bumped Astro.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Three long-standing complaints, investigated by looking at the real screen instead of arguing from code. Two turned out to be one bug; one turned out not to be a bug at all.
1 · "In Focus mode I can't see anything" — root-caused and fixed
A live capture settles it. One turn searched the web and dispatched a subagent. While it ran, Focus showed everything — live reasoning, the tool card, the fleet card with its child and a progress bar. The moment it settled, the transcript read:
Nothing on screen recorded that a search or a subagent had happened. In Focus the Run-details column is folded too (deliberately —
e306c084), so that one muted row was the entire record of the agent's work. Studio folds identically — the transcript is one shared mount — but there the Agents tab and centre column mask it.Why it never reproduced in the journeys.
absorbThoughtActivityfolds tool/fleet cards into the reasoning span that produced them, andThinkingBlockcloses onceautoOpen = failedCount > 0 || activityRunninggoes false. A model that emits no reasoning never enters that path, so its cards stay top-level rows. Every journey ran such a model. A real user on Claude Opus gets the fold on every single turn.The fold is right and stays. PRD-03 ("the answer wins") exists because six stacked cards once took 55% of the transcript and buried a one-line answer. But its own reference table shows the half that got dropped: Codex folds to
Worked for 26s ›, while Claude Desktop keeps a noun —Read depth.py ›. We copied the fold and lost the noun. The code even flagged the risk: "the count is the only thing that tells a reader work is folded… without it, absorbing the tool cards would BE hiding them." A count says work exists; it takes a noun to say what it was.So the label learns to talk:
describeActivityis pure and feeds both fold sites. It asks the tool registry first (the product's one statement of what a tool is); counts files, not calls, so three edits to one file read "Edited 1 file"; keeps a connector call's own name rather than "1 step"; reads in first-occurrence order so the line tells the run as it went; folds the long tail into+N more.Two layout details a green suite would not have caught:
width: fit-contentheader, and the Studio chat column is ~335px. It is the item that yields (flex: 0 1 auto, ellipsis, full sentence ontitle) while label and chevron stayflex: none— otherwise the row pushes its own chevron, the only affordance the control has, off the edge. Asserted as a shrink contract, since jsdom runs no layout.--color-text-subtle(~3.3:1 here, under AA) — fine for decoration. A sentence carrying the record of the work is content, so it takes--color-text-muted(~7:1).The count survives as
data-stepsand as the fallback label.Worked forstays the group's prefix — TR-16 pins it.2 · "The fonts look off" — the certain, face-independent half
TYPOGRAPHY-PROPOSAL.mdlists four candidate causes. This lands the one that is certain and two lines: every design mock sets-webkit-font-smoothing: antialiasedonbody, and so does our own boot splash — so text rendering visibly changed the moment the splash unmounted, inside one React root. On a dark theme Blink's default reads slightly bold and slightly smeared.The proposal marked this option's payoff "NOT verified — I did not run the app and did not measure a pixel." It is verified now: a like-for-like 1:1 crop of the same Focus transcript, before and after, from the real Electron window.
Deliberately not included: B2, the whole-pixel size ladder. It moves 450 call sites by 0.2–0.48px, the design uses both 12px and 12.5px so retuning
xsswaps which group is wrong, and the proposal says not to merge it without a parity run. The A-vs-C typeface choice also stays open — it is a brand decision, and this change holds either way.3 · "Can't work with subagents" (TR-8) — not a Focus bug
TR-8 failed with
fleet card missing 'Dispatched' copy: '', which reads like Focus dropping subagents. It doesn't. The phase read the first fleet card in the DOM; grouped phases share a conversation, so by TR-8 that is TR-2's card, six turns old and folded — andinnerTextof a folded card is''.Pointing the suite at a reasoning model then turned TR-1, TR-2 and TR-7 red for the same reason, so the fix is shared rather than per-phase:
JS_WITH_REVEALED— open every disclosure above a node, read, restore, in one evaluate. Restoring matters: TR-16 asserts every settled group is collapsed, and a snapshot that left older groups open would fail it three phases later with no trail back.reveal()— the mutating sibling, for phases that go on to click.ToolCallCardstampsdata-tool-namebesidedata-tool-status, on both render arms.Verification
ChatsArchive FR-G.3,canvasLifecycle PRD-B3) that fail identically on a cleandev.Worth knowing before trusting a journey run
The default journey provider is Virtuals, and its catalog currently preselects Gemini 2.5 Pro, which fails every run with
external_service_error("Service unavailable"). That is a gateway/default-model problem, not a transcript one, and is not fixed here. UseJOURNEY_PROVIDER=anthropicuntil it is looked at.Not in this PR
The subagent-control PRD (inspect / interrupt / steer a child) is still unbuilt. This PR makes subagent work visible; it does not make it controllable.
🤖 Generated with Claude Code