Skip to content

Show live component lifecycle and generated source in the REPL (#881, 2/3) - #886

Merged
taras merged 19 commits into
mainfrom
agent/issue-881-pr2
Oct 10, 2026
Merged

taras merged 19 commits into
mainfrom
agent/issue-881-pr2

Conversation

@taras

@taras taras commented Oct 10, 2026 •

Copy link
Copy Markdown
Owner

Stack 2/3 — depends on #885. Merge bottom-up.

Why

The original Terminal Interface distinguishes an element that is executing or waiting from retained history, and makes generated source and output readable. Slice 2 of #881 provides those readings.

What changes

Before: lifecycle presentation cannot reliably describe the complete invocation, and generated source/output have a flatter or clipped reading.

After: the session observes ENTER, ACTIVE, EXIT and terminal settlement through middleware, including failure and cancellation. Completion follows teardown. Generated source appears where its producer stood; each entry reads as Output then Source, with fitted rails, badges and reachable question previews.

How it works

Core’s canonical expansion observation surrounds the complete dispatch. The REPL owns live occurrence/wait observations for its session; the Journal supplies retained facts. Static source inspection reports element boundaries without resolving, compiling or executing them. The renderer fits semantic rows through the existing measured layout.

Review guide

Start with packages/core/src/component-expansion.ts and component-api.ts; then review expand.ts structural paths, source-inspection.ts, and the REPL’s lifecycle.ts, source-reading.ts and fitting.ts. The executable-MDX spec, REPL spec and architecture inventory state the contracts.

What must stay true

  • Resolution finishing does not mean invocation completion; middleware teardown precedes terminal settlement.
  • Observations are immutable and session-owned. A slow/cancelled reader cannot delay or cancel execution.
  • Historical inspection carries retained prefix facts without live lifecycle authority. Cold reads resolve/compile/execute nothing.
  • No persistence protocol or Journal schema change is introduced.

How to verify it

Accepted feedback at acadbbdb4fa7d8c8c86fa85f6657d29c5d95bc3e:

deno task test packages/core/tests/component-expansion.test.ts packages/core/tests/source-inspection.test.ts packages/core/tests/component-api.test.ts packages/core/tests/invocation-scope.test.ts packages/core/tests/invocation-failures.test.ts packages/core/tests/invocation-identity.test.ts packages/core/tests/expansion-metadata.test.ts packages/core/tests/scanner.test.ts packages/core/tests/document-targets.test.ts packages/core/tests/document-target-execution.test.ts packages/core/tests/evaluate-component.test.ts
deno task test packages/cli/tests/repl-source-reading.test.ts packages/cli/tests/repl-execution.test.ts packages/cli/tests/repl-entries.test.ts packages/cli/tests/repl-layout-admission.test.ts packages/cli/tests/repl-presentation-style.test.ts packages/cli/tests/repl-forms.test.ts packages/cli/tests/repl-journey.test.ts packages/cli/tests/repl-agent-journey.test.ts
deno task test packages/cli/tests/repl-lifecycle.test.ts packages/cli/tests/repl-agent-execution.test.ts
  • Core focused command: 53 tests / 414 steps.
  • CLI focused command: 94 tests / 316 steps.
  • Supplemental lifecycle/Agent execution command: 14 tests / 52 steps.
  • Whitespace gate: exit 0.
  • All 44 production/reference pairs independently opened; 17 supplied controls fail discriminating assertions and restore source bytes.
  • The independent eight-step Plan-to-README journey covers all three profiles, complete narrow previews, resize, offscreen advancement and warm/cold prefix inspection.

Exact commands/results are retained in .reviewer/issue-881/pr2-planner-acceptance-evidence/focused-commands.json and pr2-implementation-handback.md locally.

Scope

Includes the canonical expansion observation and static source-inspection APIs, session-owned live views, source/output fitting and focused regressions. The recorded-order History rail is the dependent final slice.

New abstractions

  • Canonical expansion observation supplies complete invocation boundaries to middleware and the REPL’s live lifecycle view.
  • Static source inspection supplies lexical element ranges to presentation without evaluating a document.
  • Private occurrence/fitting values connect session observations to measured screen rows.

Each has a concrete accepted consumer; no new dependency is added.

Risks and limitations

Structural expansion paths share the same observation boundary; review failure, cancellation and teardown alongside the ordinary component path. Narrow readings use the existing vertical-window controls rather than truncating their underlying content.

Base: the published PR 1 integration branch at c7d54769d56661bef38cee8549b148f3bf87a904. Merge bottom-up. Draft publication only: delivery verification and required merge gates remain outstanding. CI’s pull-request trigger targets main, so this stacked PR does not receive that workflow until it targets main; the focused local evidence above is implementation feedback, not a substitute for those gates. Visual packets and Planner review are local, not public attachments.

Scope confirmation

  • Every changed file supports this slice’s purpose.
  • Unrelated cleanup and formatting changes are excluded.
  • Mechanical fixture/assertion changes are identified.
  • The description distinguishes focused implementation evidence from delivery verification.

@taras
taras added this pull request to stack #888 October 10, 2026 02:32
Base automatically changed from agent/issue-881-pr1-integrated to main October 10, 2026 03:16
taras added 17 commits October 9, 2026 23:16
…ayer 1a)

Core issues a branded request before an element is resolved, the public
`Component.expand` chain composes around it, and exactly one delegation of that
exact request runs the work. The terminal keeps the canonical outcome in
private state, so a handler that catches what canonical expansion raised does
not rescue it, and one that returns without delegating refuses the work.

Phases are published to each subscriber as its latest observation plus ordered
changes, so a late reader is told where the element is rather than walked
through a history it missed, and a slow or absent reader delays nothing. The
invocation boundary now tells an observer when the body's own work ended,
before its ordered teardown runs, which is the EXIT a reader watching a
destructor needs.

Work in progress: ACTIVE publication and the structural paths follow.
ACTIVE where accepted work starts, in all three body paths, and EXIT with its
reason when that work ends — before the invocation's ordered teardown, which is
the stretch a reader watching a destructor is interested in. COMPLETE still
waits for the whole dispatch to unwind.

The first observation tier drives real expansions through the public chain: the
ordered sequence, a failed body's detached report, an element resolution
refused, two readers, a reader that never reads, a handler that returns without
delegating, a counterfeit request, a repeated delegation, a caught canonical
failure that stays failed, and a late reader told where the element is rather
than replayed.

Writing it found the hazard the contract warns about: a subscription spawned on
the handler's own frame dies when dispatch unwinds, which is before COMPLETE is
published. Subscriptions belong to an owner that outlives the dispatch, and the
tier's harness now shows that.
…881 PR 2, layer 2)

`inspectSource(text, kind)` answers where each recognized element's delimiters
are and what the author called it. It is the scanner's own reading — the same
walk that decides what a fence, an inline code span, a quoted `>` and a
tag-like expression are — so there is no second grammar to disagree with the
first, and a nested collector is committed only with the parse that found it.

A document's body boundary is read lexically, matching the installed
extractor's rule without calling its value parser: a byte-order mark, `---` on
its own line, `----` that is not one, an optional header language, the first
closing line, and the newline after it. `parseSource` asserts the two agree, so
a divergence is loud rather than a quiet offset drift. Probed against the
engine over BOM, CRLF, a header language, an absent closing delimiter and a
`---` inside a list: every case agrees, and the one the value parser refuses is
the one inspection reads without it.

The tier cuts every range out of the exact input it was given, and checks that
what inspection reports is what the engine's own scan recognized — including
an incomplete passive tag, which is prose, and leaves no guessed children.
`Component.expand` surrounds an element's whole expansion, which makes it a
step in the walk that reached it rather than a walk of its own. Counting it as
a walk would make every element a separate expansion to pause and release, and
a structural element crossing the seam would release the hold its enclosing
walk is still standing in.

The boundary inventory test caught the new member the moment it existed, which
is what it is for.
…r 1c)

`<If>`, `<Each>`, `<Loop>`, `<All>` and `<Switch>` each cross the seam once,
through one shared helper rather than five copies of the reconciliation. The
selected `<Case>`, the `<Else>` that actually runs and every `<Spawn>` an
`<All>` really starts are elements of their own; an unselected branch expands
nothing and is told nothing.

Structural syntax has no separate body to enter, so ACTIVE is published where
its region begins and EXIT when the region ends, with cancellation assumed
until the work says otherwise.

The tier now counts boundaries per construct, checks that exactly one terminal
observation reaches each, and shows an unchosen `<Case>` and the component
inside it were never elements at all. Two cleanup cases pin EXIT to the
stretch a destructor runs in, and a failing destructor to a failed completion.
The whole Core suite passes: 440 tests, 3253 steps.
…er 3b)

A private lifecycle module holds the session's reading. Each actual call of an
element gets a key of this entry's own, so a document that writes one element
twice is two calls and a repeated logical id never overwrites anything; the
enclosing key is bound *around* the delegation, following the controller's own
`CurrentWalk` composition, so descendants know which call they are in while
their capability environment stays as narrow as it was.

Consumers run on the session's owner as siblings of the dispatch. That is not
a preference: Core publishes the terminal phase after the dispatch unwinds, so
a consumer owned by the handler would be gone before it was told, and one the
handler joined would wait for itself.

A question counts a wait against the element that is asking it, read from the
lineage where the wait is taken, so two concurrent questions are two waiting
elements rather than one session that is "busy". Ordinary suspension counts
nothing.

ACTIVE is now published where resolution and validation accept the element,
which is where the accepted work starts on every body path — a refusal above
it is an element that never became active at all. The session's liveness
release moved to the entry's own finalizer, so `live` turns false after the
execution, the output consumer and the Agent attachment have all released, and
the completion wake is announced even when nothing was appended.

Tier L drives real sessions: repeated calls, nested lineage, a counted
question wait that stops when answered, ordinary work counting nothing, an
entry that takes its reading with it, and observed and unobserved runs
rendering the same thing.
…, layers 3c/4a)

The live reading reaches the view through the path the Agent reading already
takes, and a frame frozen at a recorded position leaves it out — stated once in
`viewFor`, beside the other live halves it already suppresses, rather than
decided again at each surface.

A held expansion counts a wait of its own, so an element standing at a boundary
reads as waiting rather than as slow.

Both window controls are spelled `[↑ earlier]` and `[↓ later]` everywhere a
reading has a window. The semantic actions and activation are unchanged and no
shortcut is introduced; the three suites that read those labels off a real
terminal pass unchanged otherwise.

architecture.md gains the two Core contracts and their modules;
specs/executable-mdx-spec.md gains §5.7 expansion observation and §5.8 source
inspection; specs/repl-spec.md gains what this process is doing element by
element — entered, active, waiting, exited, settled — and that none of it is
written down, reconstructed cold or shown at a frozen position.
…a/4b

The reading itself: `source-reading.ts` turns one model entry into Output then
Source, and `fitting.ts` cuts that reading to a measured width by asking the
engine how wide text actually is.

Static syntax acquires no phase. An element reads as observed only where an
observation was attributed to it — this entry, its recorded source position and
the fragment that position belongs to — so a retained prefix shows the rails
that mean nothing was observed, which is the truth about it.

Generated source replaces its producer inside the enclosure bytes the author
wrote, and keeps its own offsets as a child region rather than being spliced:
splicing would renumber the positions its own elements were recorded against. A
refused, absent or ambiguous admission keeps the producer.

Wrapping is engine answers all the way down. A UTF-16 code unit is not a cell,
so every width is measured and every cut is at a grapheme boundary; the rail,
the gap and the status column are reserved from measurements rather than from
character counts, with the cleanup wait's two readings setting the column so a
phase change cannot reflow somebody's document. A region with no room for one
grapheme is a frame refusal, never a clip.

S1–S4 in `repl-source-reading.test.ts`: 33 steps, every width from a real
engine. Two defects it found and this fixes — a self-closing element lost the
badge it had settled with, and badges align at the right inner edge rather than
at a shared start column.

The archive's `L` and `RAIL` constants join the independent reference fixture.
…yer 4

The transcript pane now holds one entry's reading: Output, then Source, sharing
one vertical window with its own earlier/later controls outside the moving
content. Every row is fitted by `prepareFrame` from engine answers, carries its
rail and, where an observation was attributed to it, the reading of the element
that owns it against the pane's right inner edge.

The defect this found in itself: `ReplMeasured.boundsOf` closes over engine
state, not over a copy, so reading it after the reading's probe render answered
from the probe tree. Every window's capacity came out wrong — the question
drawer was admitted shorter than the one on screen, a form field lost its Tab
stop, and the journey typed two answers into one field. Every bound this pass
gives is now copied into a map before anything else is measured, which is what
`planning-measurement.md` said to do and what the existing passes already did
by accident of ordering.

Also found: a self-closing element lost the badge it settled with; badges align
at the right inner edge, not at a shared start column; and the reading's two
controls have to be reserved in the pass that measures the window, or the one
it grows is drawn over the reading's last row.

The replaced presentation premise, itemized. The transcript's record rows are
gone, so fourteen assertions across five suites were re-anchored onto what the
product now says — each with the claim it is still making and why its old
anchor cannot carry it:

  - a recorded result, an outcome, a reason and live text read off the reading's
    Output half rather than off a row keyed by the record's position;
  - a wrapped reason is recovered from its rows instead of ending in an ellipsis,
    which is the stronger claim;
  - a fact about an event is read where one still is, and metadata that was a
    transcript record is the catalog's or the Bindings column's;
  - a nested scope's retained answers are read by selecting that scope, which is
    where the product says they were asked;
  - two journey probes that searched the screen for a question's words now
    assert the absence of a question and click inside the drawer that is asking
    — the reading shows the schema those words were written in, so the words
    were never the claim.

`repl-source-reading.test.ts` carries S1–S4, 33 steps, every width from a real
engine. The frozen focused command: Core 53 tests / 404 steps, CLI 94 / 310,
both exit 0. `deno check --frozen` clean, `git diff --check` clean.

Specs and architecture say what exists: the reading, its four output readings,
observation versus syntax, generated replacement, and that nothing here is
clipped or estimated.
**A pending permission is a waiting element.** The authority holds a
`"permission"` wait on the element the request was asked inside, released in
`ensure` — so a request abandoned at teardown releases its wait without
inventing a decision nobody made. L2's third reason is now counted like the
other two, and two cases in `repl-agent-execution.test.ts` assert it on the
`<Prompt>` whose turn owns the request, not on the session.

**Each call of one element is its own reading.** Several occurrences at one
authored position — what a `<Loop>` produces — are read in the order they were
observed, each in its own source group with its own badge and its own rail.
Showing only the last said the earlier calls never happened.

**A continuation keeps its own column.** A wrapped line's later rows are set in
to the whitespace its first row was written at. Without it a continuation sat
further left than the row it continues and read as a new, shallower line —
which is a claim about structure the source does not make. Display only: the
line is still recovered exactly from its rows.

**C3 and C4 are completed.** An observed run is compared with an unobserved one
byte for byte, not screen for screen: same elements, same identities, same
positions, same rendered result. A reader that takes one phase and stops, and a
reader whose scope is cancelled mid-expansion, each delay and cancel nothing.

Three new controls join the five: don't hold the permission wait; read only the
last call; start a continuation at the margin. All eight break their frozen
suites and restore byte-identically.

Core 53 tests / 407 steps, CLI 108 / 363, both exit 0. `deno check --frozen`
clean, `git diff --check` clean. The spec says all three of the new behaviours.
**C2's remaining rows.** An import that returned is not an element that
finished: while the body is held the element is entered and has no terminal
phase at all. A printed failure — a return the schema refuses, or a body whose
component prints its own errors — completes `Ok` and puts the problem in the
document, which is the frozen "preserve printed semantic failure" and exactly
why a reading takes its outcome from the phase and never from the prose.
Recovering a child's failure does not fail the parent, with the unrecovered
case beside it so the assertion is not about a parent that can never fail.
COMPLETE follows the whole dispatch, stated as the thing that ordering decides:
a subscription on the handler's own frame never sees it and one on an owner
that outlives the dispatch does.

Three of those four assertions were first written as guesses and two were
backwards. A probe against the real engine replaced them: the product completes
a printed failure `Ok`, and the drain order of an asynchronous consumer cannot
witness a publication order.

**L3's ordering.** An installation of the entry's own reads, at the moment it
is released, what the session was still saying about itself — so the claim that
everything the entry held comes down before `live = false` and the ready wake is
observed rather than read off the comment that states it. Plus the wake arriving
with no input, and a late update failing to revive a reading the entry took.

**A correction.** The last handback reported that a fence-admitted generated
fragment's source is not substituted. There is no such fragment:
`components/Evaluate.ts` is the only caller of the generated-XMD admission, so
`<Evaluate>` is the only element that makes one and the frozen rule covers every
admission there is. The claim is withdrawn, not explained.

K9 joins the controls: say the entry is not live before releasing what it held,
and L3's ordering case fails. All nine break and restore byte-identically.

The frozen focused command, run exactly as written: Core 53 / 411 steps, CLI
94 / 313, both exit 0. `repl-lifecycle` and `repl-agent-execution` are reported
separately at 14 / 52 rather than folded into it.
The Entries surface is the catalog *and* the reading of the entry it has
selected. A narrow frame has one column for both, so both are in it: the
catalog takes its own height while it fits and its share when it does not —
exactly the rule the sidebar already uses — and the reading takes what is left.

This is not a routing change. Which surface a narrow frame shows is still the
route's, and the footer still keeps its seven rows. What changes is that the
Entries outlet shows what the Entries surface *is*: routing a reader to the
surface their entry is on and then showing them nothing of it was the defect,
and every 72×20 capture in the gallery showed it.

`UI13` now asserts the reading is placed at all three profiles, in whichever
region the profile has for it — `content` at narrow, `transcript` above it. Its
old narrow leg asserted the opposite, that no row is placed there at all; that
premise is what this replaces, and the claim it was making (the reason is drawn
inside its own region, footer untouched) is unchanged and now made at every
size rather than two.

`TL6` skips a covered line that is only a window-control label. Both the
outlet and the drawer spell `[^ earlier]` the same way, so finding one inside
the drawer's rectangle says the drawer has its own rather than that something
behind it showed through — a false positive this slice created by giving narrow
a second pair of controls, not a coverage defect.

K10 joins the controls: route a narrow frame to Entries and show it no reading.
All ten break and restore byte-identically.

Frozen focused command as written: Core 53 / 411, CLI 94 / 313, both exit 0;
`repl-lifecycle` + `repl-agent-execution` beside it at 14 / 52.
**R1 — a terminal carried a request its invocation did not issue.** The brand
check proved a request was Core's and unspent; nothing proved it was *this*
element's. An authentic, unconsumed request from another live invocation,
delegated through this one's `next`, reached the claim — it spent that
element's one claim and ran this body, and the refusal that followed came after
the effect. Each terminal now compares the delegated request with the one it
issued, by identity, before the claim and before any body. The Planner's probe
went from `bodies:["A"]` to `bodies:[]`.

**R2 — the lexical envelope disagreed with the installed extractor**, and the
assertion in `parseSource` turned that into a refusal of documents the engine
accepts. I wrote the helper from the delimiter prose; the real rule came from
comparing against `matter()` over twenty-four envelope shapes. A BOM is never
body even with no header; an opener with no closing line consumes the whole
text; a closing line *starts with* `---` and the body begins immediately after
those three characters, plus one line ending if one is there.

`SI2: an unterminated header is not a header` asserted the opposite of all of
this and is corrected: that text is a header with no body, and its value parser
refuses it outright, so inspection claiming an executable element inside a
document nothing will run was the worse of the two answers. A new case checks
offset *and* suffix against `matter()` across every shape.

**R3 — published observations were mutable.** One phase object reaches every
subscriber and is retained as the latest, so a reader could rewrite a terminal
phase to `active` and a late reader would be told that. Phases are frozen
before publication, and an AggregateError's `errors` array is frozen as well as
the error around it.

Also corrected, as the review requires: `renderer.measure` already returns a
frozen copy of each bound, so the handback's claim that `boundsOf` closes over
live engine state was wrong. The geometry snapshot it justified is removed —
the journeys pass without it, which is the test that says it was never the fix.

K11–K13 join the controls, one per finding. All thirteen break and restore
byte-identically. Frozen commands: Core 53 / 414, CLI 94 / 313, supplemental
14 / 52, all exit 0.
The Planner's journey review found three things beyond R1–R3, all mine.

**A 2.8 MB test log was committed into the branch.** `.cliall.log` entered at
`108fa00c` and `.controls.json` at `a210e4eb`; `git diff --check c7d5476..HEAD`
exits 2 on the log's trailing whitespace, so the branch-level whitespace gate
has been failing while the working-tree check I was running said clean. An
empty working tree does not check the committed change. Both are gone and
`.gitignore` now names them: a test log and a control record are evidence, and
evidence belongs under `.reviewer/`.

**K1 was not the control it said it was.** It claimed "remove the rails" and
recoloured one — and when I made it actually remove them, it came back
UNBROKEN. Nothing asserted a rail reaches a row: every case read `line.rail`,
which is the model's answer, so the reading could have lost its rails entirely
and the suite would have passed. `repl-source-reading.test.ts` now checks that
every composed row opens with the rail glyph, under the rail role naming the
region the row is in, with an observed rail among them so painting every row
alike does not satisfy it. K1 removes the rails and fails; the recolouring
mutation stays as K1b, named for what it does.

**A control record kept only the matcher line.** `expect(received).toBe(
expected)` says nothing about what broke. The runner now preserves the whole
first failure, so K1's record reads `- "│"  + "  "` against the row key.

Fourteen controls, all breaking and restoring byte-identically. Frozen
commands: Core 53 / 414, CLI 94 / 314, supplemental 14 / 52, all exit 0.
The re-review closed R1–R3 and found one blocking defect left. A question's
message was split at its newlines and each logical line became one drawer row.
A drawer row is one row — so a long line had nowhere to put its tail, and
scrolling could only ever reveal rows that were prepared. At 72×20 the README
confirmation showed `"summary": "A lightweight workspace for coordinati` and
no continuation. Approve was still reachable, which is why reaching the
controls never caught it.

The preview is now fitted by the same engine-measured wrapping the reading
uses, at the drawer's own width, in `prepareFrame` beside it. Its rows are what
the description places, what the window counts and what a scroll moves through.
`fitPlain` is the reading's cutting without the rail and status reservation:
read-only text in a box uses all of the box.

`drawerContentRows` is deleted. The review cited it as counting logical lines,
which it did — and nothing called it. A dead function stating a rule the
product no longer follows is worse than no function.

The regression took two corrections worth keeping. Asserting on description
labels proved nothing: the label always held the whole line, and the loss
happened at the row. Asserting that the decisions stay reachable proved nothing
either — the reviewed build satisfied that. What discriminates is that no
message row is wider than the drawer it was measured for, together with every
logical line being recoverable from the rows it was cut into. K14 restores one
row per logical line and fails on `drawer:message:4` overflowing.

Fifteen controls, all breaking and restoring byte-identically. Core 53 / 414,
CLI 94 / 315, supplemental 14 / 52, all exit 0.
Self-audit, prompted by the shape of what the reviews keep finding. Twice now a
control was vacuous because the assertion behind it read the model where the
claim was about the screen: K1 read `line.rail` and never asked whether a rail
reached a row; K14's first draft read a description's label, which always held
the whole line while the drawn row lost its tail.

Badge alignment had the same shape. "All lifecycle glyphs/words align at the
measured source pane's right inner edge" is a frozen requirement, and
everything asserting it read the composed run strings — this terminal's own
answer about itself. Nothing checked a cell.

P1-T3 now commits a frame with one observed element and finds the row that
actually says `● ACTIVE`: its last cell is the transcript interior's last cell,
it is drawn in the archive's own literal for that phase, and the source beside
it keeps its own colours. K15 shortens the pad by three columns and it fails at
121 instead of 124.

Two fixture details the check needed, both real product behaviour: a fence is
not an element the scanner reports, so observing at one matches nothing; and an
occurrence whose position names no path is not matched to an entry's own
source, which is what keeps another document's position from claiming a row
here.

The control runner now writes itself and its record into the review packet, so
a stale copy cannot describe a different list than the one that ran — it did
twice while I was adding to it.

Sixteen controls, all breaking and restoring byte-identically. Core 53 / 414,
CLI 94 / 316, supplemental 14 / 52, all exit 0.
Finishing the audit the last commit started. Three frozen visual claims were
asserted against the model rather than the screen; one is now cells, and these
are the other two.

**Every reading row opens with its rail, in cells**, in the rail's own ink,
with the observed region drawn in the archive's `active` literal so painting
every row alike does not satisfy it. K1 checks the composed runs; nothing
checked what the terminal drew.

**Exactly one row says the badge.** "Only the first visual row carries each
badge" is frozen: a second would say the phase changed between two halves of
one element.

K16 puts the badge on every row of its delimiter — and came back **UNBROKEN**,
because the fixture's tag fitted on one row. A rule about continuation rows
cannot be tested by an element that has none. The observed element is now long
enough to wrap, and K16 fails.

That is the third vacuous control this slice has produced, all the same shape:
an assertion that cannot reach the thing it claims. K1 read the model where the
claim was the screen, K14's first draft read a label where the loss was at the
row, and K16 read a fixture that could not exhibit the defect. Each was found
by making the control honest, never by reading the code.

Seventeen controls, all breaking and restoring byte-identically. Core 53 / 414,
CLI 94 / 316, supplemental 14 / 52, all exit 0.
@taras
taras force-pushed the agent/issue-881-pr2 branch from acadbbd to 3815aef Compare October 10, 2026 03:16
@taras
taras marked this pull request as ready for review October 10, 2026 03:16
taras added 2 commits October 10, 2026 05:17
An ordinary body failure hid a `DurablePersistenceError` or a Files fatal
raised while expansion middleware unwound: `observeExpansion` reported and
threw the canonical settlement before looking at what the dispatch raised on
its way out, so `durabilityFailure()` found nothing and enclosing
reconciliation read the run as merely incorrect rather than unable to persist.

Reconcile the two halves after the whole dispatch has unwound, in the order
`expand.ts` already ranks bound failures: durability, then a Files fatal, then
the canonical failure over a later ordinary middleware one. The earlier
canonical failure is still preferred within each fatal kind, and middleware
still cannot rescue a canonical failure. The selected failure is both what is
thrown and what the one detached terminal report conveys.
@taras
taras merged commit 9fe494d into main Oct 10, 2026
43 of 44 checks passed
@taras
taras deleted the agent/issue-881-pr2 branch October 10, 2026 10:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant