Repository navigation
Show live component lifecycle and generated source in the REPL (#881, 2/3) - #886
Merged
Merged
Conversation
…ayer 1a) Core issues a branded request before an element is resolved, the public `Component.expand` chain composes around it, and exactly one delegation of that exact request runs the work. The terminal keeps the canonical outcome in private state, so a handler that catches what canonical expansion raised does not rescue it, and one that returns without delegating refuses the work. Phases are published to each subscriber as its latest observation plus ordered changes, so a late reader is told where the element is rather than walked through a history it missed, and a slow or absent reader delays nothing. The invocation boundary now tells an observer when the body's own work ended, before its ordered teardown runs, which is the EXIT a reader watching a destructor needs. Work in progress: ACTIVE publication and the structural paths follow.
ACTIVE where accepted work starts, in all three body paths, and EXIT with its reason when that work ends — before the invocation's ordered teardown, which is the stretch a reader watching a destructor is interested in. COMPLETE still waits for the whole dispatch to unwind. The first observation tier drives real expansions through the public chain: the ordered sequence, a failed body's detached report, an element resolution refused, two readers, a reader that never reads, a handler that returns without delegating, a counterfeit request, a repeated delegation, a caught canonical failure that stays failed, and a late reader told where the element is rather than replayed. Writing it found the hazard the contract warns about: a subscription spawned on the handler's own frame dies when dispatch unwinds, which is before COMPLETE is published. Subscriptions belong to an owner that outlives the dispatch, and the tier's harness now shows that.
…881 PR 2, layer 2) `inspectSource(text, kind)` answers where each recognized element's delimiters are and what the author called it. It is the scanner's own reading — the same walk that decides what a fence, an inline code span, a quoted `>` and a tag-like expression are — so there is no second grammar to disagree with the first, and a nested collector is committed only with the parse that found it. A document's body boundary is read lexically, matching the installed extractor's rule without calling its value parser: a byte-order mark, `---` on its own line, `----` that is not one, an optional header language, the first closing line, and the newline after it. `parseSource` asserts the two agree, so a divergence is loud rather than a quiet offset drift. Probed against the engine over BOM, CRLF, a header language, an absent closing delimiter and a `---` inside a list: every case agrees, and the one the value parser refuses is the one inspection reads without it. The tier cuts every range out of the exact input it was given, and checks that what inspection reports is what the engine's own scan recognized — including an incomplete passive tag, which is prose, and leaves no guessed children.
`Component.expand` surrounds an element's whole expansion, which makes it a step in the walk that reached it rather than a walk of its own. Counting it as a walk would make every element a separate expansion to pause and release, and a structural element crossing the seam would release the hold its enclosing walk is still standing in. The boundary inventory test caught the new member the moment it existed, which is what it is for.
…r 1c) `<If>`, `<Each>`, `<Loop>`, `<All>` and `<Switch>` each cross the seam once, through one shared helper rather than five copies of the reconciliation. The selected `<Case>`, the `<Else>` that actually runs and every `<Spawn>` an `<All>` really starts are elements of their own; an unselected branch expands nothing and is told nothing. Structural syntax has no separate body to enter, so ACTIVE is published where its region begins and EXIT when the region ends, with cancellation assumed until the work says otherwise. The tier now counts boundaries per construct, checks that exactly one terminal observation reaches each, and shows an unchosen `<Case>` and the component inside it were never elements at all. Two cleanup cases pin EXIT to the stretch a destructor runs in, and a failing destructor to a failed completion. The whole Core suite passes: 440 tests, 3253 steps.
…er 3b) A private lifecycle module holds the session's reading. Each actual call of an element gets a key of this entry's own, so a document that writes one element twice is two calls and a repeated logical id never overwrites anything; the enclosing key is bound *around* the delegation, following the controller's own `CurrentWalk` composition, so descendants know which call they are in while their capability environment stays as narrow as it was. Consumers run on the session's owner as siblings of the dispatch. That is not a preference: Core publishes the terminal phase after the dispatch unwinds, so a consumer owned by the handler would be gone before it was told, and one the handler joined would wait for itself. A question counts a wait against the element that is asking it, read from the lineage where the wait is taken, so two concurrent questions are two waiting elements rather than one session that is "busy". Ordinary suspension counts nothing. ACTIVE is now published where resolution and validation accept the element, which is where the accepted work starts on every body path — a refusal above it is an element that never became active at all. The session's liveness release moved to the entry's own finalizer, so `live` turns false after the execution, the output consumer and the Agent attachment have all released, and the completion wake is announced even when nothing was appended. Tier L drives real sessions: repeated calls, nested lineage, a counted question wait that stops when answered, ordinary work counting nothing, an entry that takes its reading with it, and observed and unobserved runs rendering the same thing.
…, layers 3c/4a) The live reading reaches the view through the path the Agent reading already takes, and a frame frozen at a recorded position leaves it out — stated once in `viewFor`, beside the other live halves it already suppresses, rather than decided again at each surface. A held expansion counts a wait of its own, so an element standing at a boundary reads as waiting rather than as slow. Both window controls are spelled `[↑ earlier]` and `[↓ later]` everywhere a reading has a window. The semantic actions and activation are unchanged and no shortcut is introduced; the three suites that read those labels off a real terminal pass unchanged otherwise. architecture.md gains the two Core contracts and their modules; specs/executable-mdx-spec.md gains §5.7 expansion observation and §5.8 source inspection; specs/repl-spec.md gains what this process is doing element by element — entered, active, waiting, exited, settled — and that none of it is written down, reconstructed cold or shown at a frozen position.
…a/4b The reading itself: `source-reading.ts` turns one model entry into Output then Source, and `fitting.ts` cuts that reading to a measured width by asking the engine how wide text actually is. Static syntax acquires no phase. An element reads as observed only where an observation was attributed to it — this entry, its recorded source position and the fragment that position belongs to — so a retained prefix shows the rails that mean nothing was observed, which is the truth about it. Generated source replaces its producer inside the enclosure bytes the author wrote, and keeps its own offsets as a child region rather than being spliced: splicing would renumber the positions its own elements were recorded against. A refused, absent or ambiguous admission keeps the producer. Wrapping is engine answers all the way down. A UTF-16 code unit is not a cell, so every width is measured and every cut is at a grapheme boundary; the rail, the gap and the status column are reserved from measurements rather than from character counts, with the cleanup wait's two readings setting the column so a phase change cannot reflow somebody's document. A region with no room for one grapheme is a frame refusal, never a clip. S1–S4 in `repl-source-reading.test.ts`: 33 steps, every width from a real engine. Two defects it found and this fixes — a self-closing element lost the badge it had settled with, and badges align at the right inner edge rather than at a shared start column. The archive's `L` and `RAIL` constants join the independent reference fixture.
…yer 4
The transcript pane now holds one entry's reading: Output, then Source, sharing
one vertical window with its own earlier/later controls outside the moving
content. Every row is fitted by `prepareFrame` from engine answers, carries its
rail and, where an observation was attributed to it, the reading of the element
that owns it against the pane's right inner edge.
The defect this found in itself: `ReplMeasured.boundsOf` closes over engine
state, not over a copy, so reading it after the reading's probe render answered
from the probe tree. Every window's capacity came out wrong — the question
drawer was admitted shorter than the one on screen, a form field lost its Tab
stop, and the journey typed two answers into one field. Every bound this pass
gives is now copied into a map before anything else is measured, which is what
`planning-measurement.md` said to do and what the existing passes already did
by accident of ordering.
Also found: a self-closing element lost the badge it settled with; badges align
at the right inner edge, not at a shared start column; and the reading's two
controls have to be reserved in the pass that measures the window, or the one
it grows is drawn over the reading's last row.
The replaced presentation premise, itemized. The transcript's record rows are
gone, so fourteen assertions across five suites were re-anchored onto what the
product now says — each with the claim it is still making and why its old
anchor cannot carry it:
- a recorded result, an outcome, a reason and live text read off the reading's
Output half rather than off a row keyed by the record's position;
- a wrapped reason is recovered from its rows instead of ending in an ellipsis,
which is the stronger claim;
- a fact about an event is read where one still is, and metadata that was a
transcript record is the catalog's or the Bindings column's;
- a nested scope's retained answers are read by selecting that scope, which is
where the product says they were asked;
- two journey probes that searched the screen for a question's words now
assert the absence of a question and click inside the drawer that is asking
— the reading shows the schema those words were written in, so the words
were never the claim.
`repl-source-reading.test.ts` carries S1–S4, 33 steps, every width from a real
engine. The frozen focused command: Core 53 tests / 404 steps, CLI 94 / 310,
both exit 0. `deno check --frozen` clean, `git diff --check` clean.
Specs and architecture say what exists: the reading, its four output readings,
observation versus syntax, generated replacement, and that nothing here is
clipped or estimated.
**A pending permission is a waiting element.** The authority holds a `"permission"` wait on the element the request was asked inside, released in `ensure` — so a request abandoned at teardown releases its wait without inventing a decision nobody made. L2's third reason is now counted like the other two, and two cases in `repl-agent-execution.test.ts` assert it on the `<Prompt>` whose turn owns the request, not on the session. **Each call of one element is its own reading.** Several occurrences at one authored position — what a `<Loop>` produces — are read in the order they were observed, each in its own source group with its own badge and its own rail. Showing only the last said the earlier calls never happened. **A continuation keeps its own column.** A wrapped line's later rows are set in to the whitespace its first row was written at. Without it a continuation sat further left than the row it continues and read as a new, shallower line — which is a claim about structure the source does not make. Display only: the line is still recovered exactly from its rows. **C3 and C4 are completed.** An observed run is compared with an unobserved one byte for byte, not screen for screen: same elements, same identities, same positions, same rendered result. A reader that takes one phase and stops, and a reader whose scope is cancelled mid-expansion, each delay and cancel nothing. Three new controls join the five: don't hold the permission wait; read only the last call; start a continuation at the margin. All eight break their frozen suites and restore byte-identically. Core 53 tests / 407 steps, CLI 108 / 363, both exit 0. `deno check --frozen` clean, `git diff --check` clean. The spec says all three of the new behaviours.
**C2's remaining rows.** An import that returned is not an element that finished: while the body is held the element is entered and has no terminal phase at all. A printed failure — a return the schema refuses, or a body whose component prints its own errors — completes `Ok` and puts the problem in the document, which is the frozen "preserve printed semantic failure" and exactly why a reading takes its outcome from the phase and never from the prose. Recovering a child's failure does not fail the parent, with the unrecovered case beside it so the assertion is not about a parent that can never fail. COMPLETE follows the whole dispatch, stated as the thing that ordering decides: a subscription on the handler's own frame never sees it and one on an owner that outlives the dispatch does. Three of those four assertions were first written as guesses and two were backwards. A probe against the real engine replaced them: the product completes a printed failure `Ok`, and the drain order of an asynchronous consumer cannot witness a publication order. **L3's ordering.** An installation of the entry's own reads, at the moment it is released, what the session was still saying about itself — so the claim that everything the entry held comes down before `live = false` and the ready wake is observed rather than read off the comment that states it. Plus the wake arriving with no input, and a late update failing to revive a reading the entry took. **A correction.** The last handback reported that a fence-admitted generated fragment's source is not substituted. There is no such fragment: `components/Evaluate.ts` is the only caller of the generated-XMD admission, so `<Evaluate>` is the only element that makes one and the frozen rule covers every admission there is. The claim is withdrawn, not explained. K9 joins the controls: say the entry is not live before releasing what it held, and L3's ordering case fails. All nine break and restore byte-identically. The frozen focused command, run exactly as written: Core 53 / 411 steps, CLI 94 / 313, both exit 0. `repl-lifecycle` and `repl-agent-execution` are reported separately at 14 / 52 rather than folded into it.
The Entries surface is the catalog *and* the reading of the entry it has selected. A narrow frame has one column for both, so both are in it: the catalog takes its own height while it fits and its share when it does not — exactly the rule the sidebar already uses — and the reading takes what is left. This is not a routing change. Which surface a narrow frame shows is still the route's, and the footer still keeps its seven rows. What changes is that the Entries outlet shows what the Entries surface *is*: routing a reader to the surface their entry is on and then showing them nothing of it was the defect, and every 72×20 capture in the gallery showed it. `UI13` now asserts the reading is placed at all three profiles, in whichever region the profile has for it — `content` at narrow, `transcript` above it. Its old narrow leg asserted the opposite, that no row is placed there at all; that premise is what this replaces, and the claim it was making (the reason is drawn inside its own region, footer untouched) is unchanged and now made at every size rather than two. `TL6` skips a covered line that is only a window-control label. Both the outlet and the drawer spell `[^ earlier]` the same way, so finding one inside the drawer's rectangle says the drawer has its own rather than that something behind it showed through — a false positive this slice created by giving narrow a second pair of controls, not a coverage defect. K10 joins the controls: route a narrow frame to Entries and show it no reading. All ten break and restore byte-identically. Frozen focused command as written: Core 53 / 411, CLI 94 / 313, both exit 0; `repl-lifecycle` + `repl-agent-execution` beside it at 14 / 52.
**R1 — a terminal carried a request its invocation did not issue.** The brand check proved a request was Core's and unspent; nothing proved it was *this* element's. An authentic, unconsumed request from another live invocation, delegated through this one's `next`, reached the claim — it spent that element's one claim and ran this body, and the refusal that followed came after the effect. Each terminal now compares the delegated request with the one it issued, by identity, before the claim and before any body. The Planner's probe went from `bodies:["A"]` to `bodies:[]`. **R2 — the lexical envelope disagreed with the installed extractor**, and the assertion in `parseSource` turned that into a refusal of documents the engine accepts. I wrote the helper from the delimiter prose; the real rule came from comparing against `matter()` over twenty-four envelope shapes. A BOM is never body even with no header; an opener with no closing line consumes the whole text; a closing line *starts with* `---` and the body begins immediately after those three characters, plus one line ending if one is there. `SI2: an unterminated header is not a header` asserted the opposite of all of this and is corrected: that text is a header with no body, and its value parser refuses it outright, so inspection claiming an executable element inside a document nothing will run was the worse of the two answers. A new case checks offset *and* suffix against `matter()` across every shape. **R3 — published observations were mutable.** One phase object reaches every subscriber and is retained as the latest, so a reader could rewrite a terminal phase to `active` and a late reader would be told that. Phases are frozen before publication, and an AggregateError's `errors` array is frozen as well as the error around it. Also corrected, as the review requires: `renderer.measure` already returns a frozen copy of each bound, so the handback's claim that `boundsOf` closes over live engine state was wrong. The geometry snapshot it justified is removed — the journeys pass without it, which is the test that says it was never the fix. K11–K13 join the controls, one per finding. All thirteen break and restore byte-identically. Frozen commands: Core 53 / 414, CLI 94 / 313, supplemental 14 / 52, all exit 0.
The Planner's journey review found three things beyond R1–R3, all mine. **A 2.8 MB test log was committed into the branch.** `.cliall.log` entered at `108fa00c` and `.controls.json` at `a210e4eb`; `git diff --check c7d5476..HEAD` exits 2 on the log's trailing whitespace, so the branch-level whitespace gate has been failing while the working-tree check I was running said clean. An empty working tree does not check the committed change. Both are gone and `.gitignore` now names them: a test log and a control record are evidence, and evidence belongs under `.reviewer/`. **K1 was not the control it said it was.** It claimed "remove the rails" and recoloured one — and when I made it actually remove them, it came back UNBROKEN. Nothing asserted a rail reaches a row: every case read `line.rail`, which is the model's answer, so the reading could have lost its rails entirely and the suite would have passed. `repl-source-reading.test.ts` now checks that every composed row opens with the rail glyph, under the rail role naming the region the row is in, with an observed rail among them so painting every row alike does not satisfy it. K1 removes the rails and fails; the recolouring mutation stays as K1b, named for what it does. **A control record kept only the matcher line.** `expect(received).toBe( expected)` says nothing about what broke. The runner now preserves the whole first failure, so K1's record reads `- "│" + " "` against the row key. Fourteen controls, all breaking and restoring byte-identically. Frozen commands: Core 53 / 414, CLI 94 / 314, supplemental 14 / 52, all exit 0.
The re-review closed R1–R3 and found one blocking defect left. A question's message was split at its newlines and each logical line became one drawer row. A drawer row is one row — so a long line had nowhere to put its tail, and scrolling could only ever reveal rows that were prepared. At 72×20 the README confirmation showed `"summary": "A lightweight workspace for coordinati` and no continuation. Approve was still reachable, which is why reaching the controls never caught it. The preview is now fitted by the same engine-measured wrapping the reading uses, at the drawer's own width, in `prepareFrame` beside it. Its rows are what the description places, what the window counts and what a scroll moves through. `fitPlain` is the reading's cutting without the rail and status reservation: read-only text in a box uses all of the box. `drawerContentRows` is deleted. The review cited it as counting logical lines, which it did — and nothing called it. A dead function stating a rule the product no longer follows is worse than no function. The regression took two corrections worth keeping. Asserting on description labels proved nothing: the label always held the whole line, and the loss happened at the row. Asserting that the decisions stay reachable proved nothing either — the reviewed build satisfied that. What discriminates is that no message row is wider than the drawer it was measured for, together with every logical line being recoverable from the rows it was cut into. K14 restores one row per logical line and fails on `drawer:message:4` overflowing. Fifteen controls, all breaking and restoring byte-identically. Core 53 / 414, CLI 94 / 315, supplemental 14 / 52, all exit 0.
Self-audit, prompted by the shape of what the reviews keep finding. Twice now a control was vacuous because the assertion behind it read the model where the claim was about the screen: K1 read `line.rail` and never asked whether a rail reached a row; K14's first draft read a description's label, which always held the whole line while the drawn row lost its tail. Badge alignment had the same shape. "All lifecycle glyphs/words align at the measured source pane's right inner edge" is a frozen requirement, and everything asserting it read the composed run strings — this terminal's own answer about itself. Nothing checked a cell. P1-T3 now commits a frame with one observed element and finds the row that actually says `● ACTIVE`: its last cell is the transcript interior's last cell, it is drawn in the archive's own literal for that phase, and the source beside it keeps its own colours. K15 shortens the pad by three columns and it fails at 121 instead of 124. Two fixture details the check needed, both real product behaviour: a fence is not an element the scanner reports, so observing at one matches nothing; and an occurrence whose position names no path is not matched to an entry's own source, which is what keeps another document's position from claiming a row here. The control runner now writes itself and its record into the review packet, so a stale copy cannot describe a different list than the one that ran — it did twice while I was adding to it. Sixteen controls, all breaking and restoring byte-identically. Core 53 / 414, CLI 94 / 316, supplemental 14 / 52, all exit 0.
Finishing the audit the last commit started. Three frozen visual claims were asserted against the model rather than the screen; one is now cells, and these are the other two. **Every reading row opens with its rail, in cells**, in the rail's own ink, with the observed region drawn in the archive's `active` literal so painting every row alike does not satisfy it. K1 checks the composed runs; nothing checked what the terminal drew. **Exactly one row says the badge.** "Only the first visual row carries each badge" is frozen: a second would say the phase changed between two halves of one element. K16 puts the badge on every row of its delimiter — and came back **UNBROKEN**, because the fixture's tag fitted on one row. A rule about continuation rows cannot be tested by an element that has none. The observed element is now long enough to wrap, and K16 fails. That is the third vacuous control this slice has produced, all the same shape: an assertion that cannot reach the thing it claims. K1 read the model where the claim was the screen, K14's first draft read a label where the loss was at the row, and K16 read a fixture that could not exhibit the defect. Each was found by making the control honest, never by reading the code. Seventeen controls, all breaking and restoring byte-identically. Core 53 / 414, CLI 94 / 316, supplemental 14 / 52, all exit 0.
taras
force-pushed
the
agent/issue-881-pr2
branch
from
October 10, 2026 03:16
acadbbd to
3815aef
Compare
taras
marked this pull request as ready for review
October 10, 2026 03:16
An ordinary body failure hid a `DurablePersistenceError` or a Files fatal raised while expansion middleware unwound: `observeExpansion` reported and threw the canonical settlement before looking at what the dispatch raised on its way out, so `durabilityFailure()` found nothing and enclosing reconciliation read the run as merely incorrect rather than unable to persist. Reconcile the two halves after the whole dispatch has unwound, in the order `expand.ts` already ranks bound failures: durability, then a Files fatal, then the canonical failure over a later ordinary middleware one. The earlier canonical failure is still preferred within each fatal kind, and middleware still cannot rescue a canonical failure. The selected failure is both what is thrown and what the one detached terminal report conveys.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stack 2/3 — depends on #885. Merge bottom-up.
Why
The original Terminal Interface distinguishes an element that is executing or waiting from retained history, and makes generated source and output readable. Slice 2 of #881 provides those readings.
What changes
Before: lifecycle presentation cannot reliably describe the complete invocation, and generated source/output have a flatter or clipped reading.
After: the session observes ENTER, ACTIVE, EXIT and terminal settlement through middleware, including failure and cancellation. Completion follows teardown. Generated source appears where its producer stood; each entry reads as Output then Source, with fitted rails, badges and reachable question previews.
How it works
Core’s canonical expansion observation surrounds the complete dispatch. The REPL owns live occurrence/wait observations for its session; the Journal supplies retained facts. Static source inspection reports element boundaries without resolving, compiling or executing them. The renderer fits semantic rows through the existing measured layout.
Review guide
Start with
packages/core/src/component-expansion.tsandcomponent-api.ts; then reviewexpand.tsstructural paths,source-inspection.ts, and the REPL’slifecycle.ts,source-reading.tsandfitting.ts. The executable-MDX spec, REPL spec and architecture inventory state the contracts.What must stay true
How to verify it
Accepted feedback at
acadbbdb4fa7d8c8c86fa85f6657d29c5d95bc3e:Exact commands/results are retained in
.reviewer/issue-881/pr2-planner-acceptance-evidence/focused-commands.jsonandpr2-implementation-handback.mdlocally.Scope
Includes the canonical expansion observation and static source-inspection APIs, session-owned live views, source/output fitting and focused regressions. The recorded-order History rail is the dependent final slice.
New abstractions
Each has a concrete accepted consumer; no new dependency is added.
Risks and limitations
Structural expansion paths share the same observation boundary; review failure, cancellation and teardown alongside the ordinary component path. Narrow readings use the existing vertical-window controls rather than truncating their underlying content.
Base: the published PR 1 integration branch at
c7d54769d56661bef38cee8549b148f3bf87a904. Merge bottom-up. Draft publication only: delivery verification and required merge gates remain outstanding. CI’s pull-request trigger targetsmain, so this stacked PR does not receive that workflow until it targetsmain; the focused local evidence above is implementation feedback, not a substitute for those gates. Visual packets and Planner review are local, not public attachments.Scope confirmation