Skip to content

Focus stops hiding the agent's work: folded activity names what it did - #658

Merged
parthpahwa1 merged 6 commits into
devfrom
fix/folded-activity-names-its-work
Sep 20, 2026
Merged

parthpahwa1 merged 6 commits into
devfrom
fix/folded-activity-names-its-work

Conversation

@0x-copilot-dev

Copy link
Copy Markdown
Owner

Three long-standing complaints, investigated by looking at the real screen instead of arguing from code. Two turned out to be one bug; one turned out not to be a bug at all.

1 · "In Focus mode I can't see anything" — root-caused and fixed

A live capture settles it. One turn searched the web and dispatched a subagent. While it ran, Focus showed everything — live reasoning, the tool card, the fleet card with its child and a progress bar. The moment it settled, the transcript read:

Thought process · 2 steps ⌄
<the answer>

Nothing on screen recorded that a search or a subagent had happened. In Focus the Run-details column is folded too (deliberately — e306c084), so that one muted row was the entire record of the agent's work. Studio folds identically — the transcript is one shared mount — but there the Agents tab and centre column mask it.

Why it never reproduced in the journeys. absorbThoughtActivity folds tool/fleet cards into the reasoning span that produced them, and ThinkingBlock closes once autoOpen = failedCount > 0 || activityRunning goes false. A model that emits no reasoning never enters that path, so its cards stay top-level rows. Every journey ran such a model. A real user on Claude Opus gets the fold on every single turn.

The fold is right and stays. PRD-03 ("the answer wins") exists because six stacked cards once took 55% of the transcript and buried a one-line answer. But its own reference table shows the half that got dropped: Codex folds to Worked for 26s ›, while Claude Desktop keeps a noun — Read depth.py ›. We copied the fold and lost the noun. The code even flagged the risk: "the count is the only thing that tells a reader work is folded… without it, absorbing the tool cards would BE hiding them." A count says work exists; it takes a noun to say what it was.

So the label learns to talk:

Thought process · Searched the web · Dispatched 1 subagent ⌄
Worked for 8.3s · Read 3 files · Edited 1 file ⌄

describeActivity is pure and feeds both fold sites. It asks the tool registry first (the product's one statement of what a tool is); counts files, not calls, so three edits to one file read "Edited 1 file"; keeps a connector call's own name rather than "1 step"; reads in first-occurrence order so the line tells the run as it went; folds the long tail into +N more.

Two layout details a green suite would not have caught:

  • The digest is the one unbounded string in a width: fit-content header, and the Studio chat column is ~335px. It is the item that yields (flex: 0 1 auto, ellipsis, full sentence on title) while label and chevron stay flex: none — otherwise the row pushes its own chevron, the only affordance the control has, off the edge. Asserted as a shrink contract, since jsdom runs no layout.
  • The count was drawn in --color-text-subtle (~3.3:1 here, under AA) — fine for decoration. A sentence carrying the record of the work is content, so it takes --color-text-muted (~7:1).

The count survives as data-steps and as the fallback label. Worked for stays the group's prefix — TR-16 pins it.

2 · "The fonts look off" — the certain, face-independent half

TYPOGRAPHY-PROPOSAL.md lists four candidate causes. This lands the one that is certain and two lines: every design mock sets -webkit-font-smoothing: antialiased on body, and so does our own boot splash — so text rendering visibly changed the moment the splash unmounted, inside one React root. On a dark theme Blink's default reads slightly bold and slightly smeared.

The proposal marked this option's payoff "NOT verified — I did not run the app and did not measure a pixel." It is verified now: a like-for-like 1:1 crop of the same Focus transcript, before and after, from the real Electron window.

Deliberately not included: B2, the whole-pixel size ladder. It moves 450 call sites by 0.2–0.48px, the design uses both 12px and 12.5px so retuning xs swaps which group is wrong, and the proposal says not to merge it without a parity run. The A-vs-C typeface choice also stays open — it is a brand decision, and this change holds either way.

3 · "Can't work with subagents" (TR-8) — not a Focus bug

TR-8 failed with fleet card missing 'Dispatched' copy: '', which reads like Focus dropping subagents. It doesn't. The phase read the first fleet card in the DOM; grouped phases share a conversation, so by TR-8 that is TR-2's card, six turns old and folded — and innerText of a folded card is ''.

Pointing the suite at a reasoning model then turned TR-1, TR-2 and TR-7 red for the same reason, so the fix is shared rather than per-phase:

  • JS_WITH_REVEALED — open every disclosure above a node, read, restore, in one evaluate. Restoring matters: TR-16 asserts every settled group is collapsed, and a snapshot that left older groups open would fail it three phases later with no trail back.
  • reveal() — the mutating sibling, for phases that go on to click.
  • The Focus readers take the newest card host, not the first match.
  • TR-7 learned which tool ran by reading copy; the title is model-authored now, so ToolCallCard stamps data-tool-name beside data-tool-status, on both render arms.

Verification

  • chat-surface: 4294 passed; the 2 failures are the long-standing named pair (ChatsArchive FR-G.3, canvasLifecycle PRD-B3) that fail identically on a clean dev.
  • chat-surface + desktop typecheck clean.
  • Journeys on Anthropic (a reasoning model): TR-1, 2, 3, 6, 7, 8, 9 pass.
  • Every visual claim above was checked against a screenshot of the running app.

Worth knowing before trusting a journey run

The default journey provider is Virtuals, and its catalog currently preselects Gemini 2.5 Pro, which fails every run with external_service_error ("Service unavailable"). That is a gateway/default-model problem, not a transcript one, and is not fixed here. Use JOURNEY_PROVIDER=anthropic until it is looked at.

Not in this PR

The subagent-control PRD (inspect / interrupt / steer a child) is still unbuilt. This PR makes subagent work visible; it does not make it controllable.

🤖 Generated with Claude Code

0x-copilot-dev and others added 3 commits September 20, 2026 16:34
… how much

"In Focus mode I can't see anything" was argued from code for weeks. A live
capture settles it: one turn searched the web and dispatched a subagent, and
the moment it settled the transcript read

    Thought process · 2 steps ⌄
    <the answer>

Nothing on screen recorded that a search or a subagent had happened. In Focus
the Run-details column is folded too (deliberately, e306c08), so that one muted
row was the ENTIRE record of the agent's work. Studio folds identically — the
transcript is a single shared mount — but there the Agents tab and the centre
column mask it.

Why it never reproduced in the journeys: `absorbThoughtActivity` folds tool and
fleet cards into the reasoning span that produced them, and `ThinkingBlock`
closes when `autoOpen = failedCount > 0 || activityRunning` goes false. A model
that emits no reasoning never enters that path, so its cards stay top-level
rows. Every journey ran such a model. A real user on Claude Opus gets the fold
on every turn.

The fold itself is right and stays. PRD-03 ("the answer wins") exists because
six stacked cards once took 55% of the transcript and buried a one-line answer.
But PRD-03's own reference table shows the half that got dropped: Codex folds
to `Worked for 26s ›`, while Claude Desktop keeps a NOUN — `Read depth.py ›`.
We copied the fold and lost the noun. The code even knew the risk: "the count
is the only thing that tells a reader work is folded … without it, absorbing
the tool cards would BE hiding them." A count says work exists; it takes a noun
to say what it was.

So the label learns to talk:

    Thought process · Searched the web · Dispatched 1 subagent ⌄
    Worked for 8.3s · Read 3 files · Edited 1 file ⌄

`describeActivity` is pure and feeds BOTH fold sites (the thought block and the
PRD-03 run group). It asks the tool registry first, because that is the
product's one statement of what a tool is; counts FILES not calls, so three
edits to one file read "Edited 1 file"; keeps a connector call's own name
rather than calling it "1 step"; reads in first-occurrence order so the line
tells the run as it went; and folds the long tail into "+N more".

Two layout details that a green suite would not have caught:

* The digest is the one unbounded string in a `width: fit-content` header, and
  the Studio chat column is ~335px. It is therefore the item that yields
  (`flex: 0 1 auto`, ellipsis, full sentence on `title`) while the label and
  the chevron stay `flex: none` — otherwise the row pushes its own chevron, the
  only affordance the control has, off the edge. Asserted as a shrink CONTRACT
  because jsdom runs no layout.
* The count was drawn in `--color-text-subtle` (~3.3:1 here, under AA), which
  was fine for decoration. A sentence carrying the record of the work is
  content, so it takes the label's `--color-text-muted` (~7:1).

The count survives as `data-steps` for anything that needs the number, and as
the fallback label, so a caller that cannot describe its members still never
folds them silently. `Worked for` stays the group's prefix — the TR-16 journey
pins it.

Verified on the real screen, before and after, not only in jsdom.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…gn and our own splash do

"The fonts look off" has four candidate causes in TYPOGRAPHY-PROPOSAL.md, and
this is the one that is certain, face-independent and two lines.

Every design mock sets `-webkit-font-smoothing: antialiased` on body
(design-kit/app-v3/copilot.css:110). So does our own boot splash, on `.boot`
(apps/desktop/renderer/BootProgress.css:34). The app body did not — so text
rendering visibly CHANGED the moment the splash unmounted, inside one React
root: stems went from the splash's thin grayscale to Blink's heavier default.
On a dark theme that default reads as slightly bold and slightly smeared, which
is a fair description of the complaint and has nothing to do with the typeface.

The proposal marked this option's visual payoff "NOT verified — I did not run
the app and did not measure a pixel." It is verified now: a like-for-like 1:1
crop of the same Focus transcript, before and after, shows light-on-dark stems
thinning to their designed weight. Captured from the real Electron window, not
headless Chromium and not jsdom.

macOS-only properties; inert on Windows, where ClearType owns this. Holds
whichever way the proposal's A-vs-C typeface decision goes, which is why it can
land ahead of that decision. Revert = delete two lines.

Deliberately NOT included: B2, the whole-pixel size ladder. It moves 450 call
sites by 0.2-0.48px, the design uses BOTH 12px and 12.5px so retuning `xs`
swaps which group is wrong, and the proposal itself says not to merge it
without a parity run. That wants its own change and its own look at the app.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… a reasoning model

`transcript_rendering` was green only by accident of which model ran it. Every
phase was written against a model that emits no reasoning, where a tool or
fleet card is a top-level transcript row. On a reasoning model — what a real
user on Claude Opus has — `absorbThoughtActivity` folds every card into the
"Thought process" block, which closes when the run settles. Pointing the suite
at Anthropic turned TR-1, TR-2, TR-7 and TR-8 red with messages that all read
like product bugs:

    no required new fleet card appeared … 'text': '', 'fleetStatus': 'done'
    page.click: Timeout … element is not visible
    fleet card missing 'Dispatched' copy: ''

`innerText` is layout-aware, so a card inside a closed disclosure reads `''`,
and Playwright will not click what has no visible box. Same root cause as the
TR-3 fix two PRs ago, now handled once instead of per phase:

* `JS_WITH_REVEALED` — open every disclosure above a node, read, RESTORE, all in
  one evaluate. Reading `innerText` forces the layout flush, so the text is real
  and nothing repaints. Restoring matters: TR-16 asserts every settled group IS
  collapsed, and a snapshot that left older groups open would fail it three
  phases later with no trail back. `card_snapshots` and both Focus readers use
  it.
* `reveal()` — the mutating sibling, for phases that go on to CLICK. A user who
  wants to operate a folded card opens the fold first. TR-1's first click and
  `ensure_native_disclosure` use it; `ensure_fleet_card_interaction` drops its
  private copy.

Two more stale assumptions, both in the Focus readers, both from when each was
its own journey with its own boot:

* **They took the FIRST card in the DOM.** Grouped phases share a conversation,
  so by TR-8 the first fleet card is TR-2's, six turns old. They now skip the
  cards that existed before the phase sent and take the newest — selecting card
  HOSTS (`[data-tool-status]`), because the bare testid prefix also matches a
  card's own `…-details` child and "the last match" was a fragment of a card.
* **TR-7 learned which tool ran by reading copy.** The title is model-authored
  now ("math.isqrt docs"), so the raw name is on no pixel. `ToolCallCard` stamps
  `data-tool-name` on its root, beside `data-tool-status` and for the same
  reason, on both render arms.

Result on Anthropic (a reasoning model): TR-1, 2, 3, 6, 7, 8, 9 all pass.

Not fixed here, and worth knowing before trusting a run: the DEFAULT journey
provider is Virtuals, and its catalog currently preselects Gemini 2.5 Pro,
which fails every run with `external_service_error` ("Service unavailable").
That is a gateway/default-model problem, not a transcript one — run with
`JOURNEY_PROVIDER=anthropic` until it is looked at.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
0x-copilot-dev and others added 3 commits September 20, 2026 16:42
…e 2026-08-27

`build-and-audit` is red on dev, and has been since something that is not a
code change: `ci-frontend` was GREEN on 2026-08-27 (ed72f5c) and is red on
today's dev HEAD (7193fbc), with no dependency edit in between. New advisories
were published against versions already pinned. The gate is
`npm audit --audit-level=high`, so the 6 moderates do not matter and these 8 do:

  critical  astro             <=7.2.7    RCE via AVIF image optimization
  high      sharp             <0.35.4    libheif GHSA-g89c-p67h-r497 / -2jg2-4ch7-h545
  high      svgo              4.0.0-4.0.2  removeScripts sanitiser bypasses (x2)
  high      smol-toml         <=1.7.0    DoS via malformed TOML
  high      js-yaml           4.0.0-4.3.1  maxTotalMergeKeys does not bound CPU
  high      browserslist      <=4.28.6   unbounded cache growth -> OOM
  high      fast-uri          3.0.0-3.1.5  host confusion via skipped IDN canonicalisation
  high      @xmldom/xmldom    <=0.8.14   XML fragment injection via EntityReference

Every move is a patch/minor INSIDE the existing ranges, so no package.json
changes and no new direct dependency: astro 7.2.0 -> 7.3.3 (satisfies ^7.2.0),
sharp 0.35.3 -> 0.35.4, svgo 4.0.2 -> 4.1.0, smol-toml 1.7.0 -> 1.8.0, js-yaml
4.3.1 -> 4.3.2, browserslist 4.28.6 -> 4.29.0, fast-uri 3.1.5 -> 3.1.8, xmldom
0.8.13 -> 0.8.15. All of it is the `apps/website` (Astro) tree — nothing in
`apps/desktop` or `packages/*` moves.

WHY `--package-lock-only` AND NOT `npm audit fix`. This repo's lockfile does not
survive a full re-resolve (peer conflict in the electron-builder tree plus
website hoisting), and a regenerated lockfile reds the SBOM gate. A targeted
`audit fix --package-lock-only` touches only the advisory subtrees and leaves
node_modules alone. Verified stable: `npm ci` succeeds against the result and
does NOT mutate the lockfile further, which is the property a regeneration
breaks.

On the SBOM step that runs right after the audit: `cyclonedx-npm` fails LOCALLY
on both this lockfile and the one before it, identically — `npm ls` reports 26
pre-existing peer problems (@electron/asar, @electron/universal, ejs, all from
app-builder-lib) on a clean install of either. So it is not something this
change introduces, and the step passed in CI as recently as 2026-08-27, which
is the run this commit restores the job to.

Unrelated to the rest of this PR; it is here because the PR cannot go green
without it while dev is red.

Verified: audit gate passes, `npm ci` clean, no manifest changes, chat-surface
and desktop typecheck at 0 errors, desktop builds, and the website — the only
tree that moved — builds on the bumped Astro.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…e goes green

`build-and-audit` is red at "Generate SBOM". `cyclonedx-npm` shells out to
`npm ls --json --long --all`, which exits ELSPROBLEMS with 7 extraneous and 3
invalid packages:

  extraneous  @img/sharp-wasm32 + its @emnapi/*, @napi-rs/wasm-runtime,
              @tybys/wasm-util subtree — orphaned, nothing declares an edge
  invalid     app-builder-lib/node_modules/{@electron/asar@3.4.1,
              @electron/universal@2.0.3, ejs@3.1.10} — the pre-override
              versions the root `overrides` block exists to eliminate

CORRECTION TO 15c297c. That commit reports the SBOM breakage as pre-existing
and "not something this change introduces". Measured on CI's own toolchain that
is not so, and the difference is the npm version:

                        dev lockfile          15c297c lockfile
  npm 10.9.8 (CI)       npm ls exit 0         exit 1 — 7 extraneous, 3 invalid
  npm 11.9.0 (local)    exit 1 — 11 probs     exit 1 — 10 probs

Under npm 11.9 both trees look broken, which is what made it read as
pre-existing. Under the npm CI actually runs — 10.9.8, bundled with the
node 22.23.2 the workflow pins — dev is spotless and only the new lockfile
fails, reproducing the CI error package-for-package. The last green SBOM run
(2026-08-27, ed72f5c) uploaded its artifact off that same dev lockfile, which
is the authoritative check the local one contradicted.

The mechanism: 15c297c used `npm audit fix --package-lock-only` on npm 11.9,
whose resolver nests un-overridden copies under app-builder-lib and leaves the
stale sharp wasm32 subtree behind when sharp 0.35.4 drops it. npm 10.9.8 then
calls both out.

Note that npm 11.9 was not a gratuitous choice: npm 10.9.8 CANNOT run
`npm audit fix` in this repo at all — it dies on `"playwright": "$playwright"`
with `Unable to resolve reference $playwright`. Plain
`npm install --package-lock-only` and `npm update` are unaffected, so the
advisory work is done with those instead. Worth replacing that `$playwright`
reference with its literal version if we want `audit fix` back on npm 10.

WHAT THIS DOES. Starts from dev's clean lockfile and re-applies the same 8
bumps on npm 10.9.8 — `npm update` for seven, and a re-resolve inside the
unchanged `^7.2.0` range for astro. Same end state, no nested override
violations, no orphan subtree: astro 7.3.3, sharp 0.35.4, svgo 4.1.0,
smol-toml 1.8.0, js-yaml 4.3.2, browserslist 4.29.0, fast-uri 3.1.8,
@xmldom/xmldom 0.8.15. It is 303 lines SMALLER than the lockfile it replaces,
because it drops that subtree rather than adding to it.

Only package-lock.json moves; every package.json is byte-identical. astro now
resolves at apps/website/node_modules/astro rather than the root hoist — npm
10's own placement, not a hand edit.

Verified on node 22.23.2 / npm 10.9.8, the workflow's pinned pair: `npm ci`
clean, all six package typechecks at 0 errors, frontend builds, `npm audit
--audit-level=high` passes, `npm ls --json --long --all` exits 0, the SBOM step
produces a valid CycloneDX 1.6 document with 869 components, and the website —
the only tree that moved — builds on the bumped Astro.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@parthpahwa1
parthpahwa1 merged commit fa0b9eb into dev Sep 20, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants