Show a room what the agent is doing while it does it - #175
Conversation
A terminal prints tool calls as they run. A room got the typing indicator and then, some minutes later, a wall of prose -- so a turn that was working and a turn that had hung looked identical. rooms.R registered no observers at all; the whole asymmetry was that. The trail is one collapsible line -- "Ran 3 commands, read 2 files +102 -33" -- posted on the first tool call and edited as the turn goes. <details> and <summary> are in Matrix's allowed subset and Element renders them natively, so it is a real disclosure widget rather than a heading over a wall of text. The plain body carries the same summary in one line, because that is what a client without HTML shows and what the push notification says. Two costs shape the cadence, and neither exists in a terminal. A Matrix edit is an ordinary timeline event, so it competes with real messages for the rate limit -- Synapse allows a burst of ten and then one every five seconds -- and it is permanent: only the latest renders, but every frame stays in the room and comes back out of chat_history() on the next restart. So the trail is pushed when its text actually changes, with a five-second floor, rather than on a timer. A fixed one-second tick would spend six hundred permanent events on a ten-minute turn and lose the frames at the end, where the work finished. Liveness stays the typing indicator's job: an ephemeral EDU that costs nothing and was already being sent. The observer is attached per turn and removed after. A session outlives the turn, and a leftover one would keep appending to an accumulator whose message was already finalized -- the next turn's first tool call would silently edit the previous turn's trail. Restored on error too, or one failed turn doubles the trail on every turn after it. Tool names are model-supplied and escaped before they reach the markup: an unescaped one could close the summary element and write into the room. Diffs report exact counts because .diff_summary_counts() now rides along on the payload -- `lines` is clipped to a budget, so anything recounting from it undercounts a large diff.
The wrapper in bot_poll() had no coverage: removing it left the suite green, because every other test of the module builds the accumulator itself. This drives a real poll and asserts the observer is on the session the turn runs on, and that what it produces reaches the room. It also pins the post-then-edit shape end to end -- one send on the first tool call, one edit with the finished state -- and that the edit targets the message the send created. The second tool call landing inside the five-second floor is the reason there are two frames and not three.
Mutation testing found dead code and a wasted event behind it. The final flush set last_text to NULL first, which looked like it was bypassing the interval floor -- but flush() never consulted the floor, so the line did nothing except hide what it was actually doing: re-sending the turn's own text as a permanent event, on every turn whose last tool call happened more than five seconds before it finished. The skip belongs in flush() rather than only in the observer's gate. The floor is about spending the rate-limit budget during a turn; a duplicate frame is never useful, whoever asks for it.
An encrypted Matrix room refuses edits: the replacement text would ride in an ordinary event, so chat_edit() throws there by design. The observer catches those errors, which meant an encrypted room got the first frame -- "Ran a command" -- and then kept it for the rest of the turn while every update was rejected. Nothing logged, nothing looked wrong, and the room said something untrue for ten minutes. That is worse than showing nothing, and it is the exact shape of failure this migration has been finding all along: an error treated as an answer. chat_capabilities()$edits is now asked once, before the first frame. A client that can edit gets the live trail. One that cannot gets no intermediate frames at all and a single final message, written after everything happened and therefore accurate. Unknowable counts as cannot. Found in review, not by me -- I had reported it as a missing disclosure widget, which is the cosmetic half of it.
The capability check went in between rooms_with_activity()'s comment and rooms_with_activity(), so the file read as though the paragraph about attaching an observer per turn described the one-line predicate underneath it.
|
Fixed, and you were right that I'd described it as the cosmetic half. It isn't a missing widget — the room gets a stale one. First frame posts, every Went with the capability gate rather than documenting it, because the third option is better than both: rooms_activity_live(chat) # chat_capabilities()$edits, once, before the first frame
The one-accurate-message path is strictly better than what an encrypted room gets today, which is nothing. chat.api#11 is merged ( 2926/2926 locally. One thing I'd still call unsettled, unchanged from before: the five-second floor is a starting point, not a measurement, and I quoted Synapse's |
Two changes from review, both about what the trail costs a room. The interval is what bounds the cost, not the content check. Almost every completed tool call changes the summary -- "Ran 3 commands" becomes "Ran 4 commands" -- so the change check filters very little and five seconds was close to one permanent event per call. Fifteen puts a ten-minute turn at forty events instead of hundreds, and `activity_interval` in the config moves it. Still not a timer. Nothing runs between tool calls, so a frame goes out on the first completed call after the interval elapses rather than at the instant it does, and the final flush covers whatever the last interval did not. And the trail is m.notice while replies stay m.text. That is what the msgtype is for: automated output another bot should not answer, which matters most in exactly the rooms this runs in. It is not a mute -- the spec asks clients not to auto-respond and says nothing about push gateways, so a room that wants silence still needs a push rule. Claiming otherwise in a comment would be the more expensive mistake. Needs chat.api 0.0.1.20: chat_edit() hardcoded m.text in m.new_content, so the first edit turned the notice into an ordinary message.
…e one corteza 0.7.1.21 needs chat_edit() with rich and kind -- 0.0.1.20 -- and declared 0.0.1.17. A host on 0.0.1.17 would have passed the runtime guard and then found no chat_edit() at all, inside the observer's tryCatch, leaving the room a stale first frame. Which is the exact failure the capability gate was just added to prevent, arriving by a different route. Two bumps did it. Both were sed with a pattern that no longer matched, so both silently changed nothing -- and the assertions pinning the constant were edited by the same non-matching patterns, so they stayed consistent with the stale value and went on passing. CI then verified 0.0.1.19 against a floor of 0.0.1.17 and went green: a verification step reporting success for the thing it exists to catch. The guard against a repeat reads the Suggests bound out of DESCRIPTION and compares it to the constant, rather than restating a literal that can only ever agree with itself.
|
Correction to the green tick you reviewed — it was green on a stale floor, and I only caught it while merging.
A verification step reporting success for exactly the thing it exists to catch. The consequence was real, not bookkeeping. corteza 0.7.1.21 needs Fixed, with a guard that reads the Suggests bound out of DESCRIPTION and compares it to the constant rather than restating a literal that can only ever agree with itself. Verified by mutation: bumping DESCRIPTION alone turns the suite red. 2939/2939. CI is rerunning against the corrected floor with chat.api 0.0.1.20 on main. The general lesson I'd take from this one: a |
Depends on cornball-ai/chat.api#11.
A terminal prints tool calls as they run. A room got the typing indicator and then, some minutes later, a wall of prose — so a turn that was working and a turn that had hung looked identical.
The cause turned out to be one line long:
rooms.Rregistered no observers at all. The terminal registers two.add_observer()is generic and session-scoped, the events already carrycall$tool,outcome,successanddiff, and none of it was reaching a room.What a room sees now
One collapsible line, posted on the first tool call and edited as the turn goes:
<details>and<summary>are in Matrix's allowed HTML subset and Element renders them natively, so it is a real disclosure widget rather than a heading over a wall of text. The plain body carries the same summary in one line — that is what a client without HTML shows and what the push notification says, so it has to stand alone.The cadence, and why it is not a timer
Two costs shape this and neither exists in a terminal.
A Matrix edit is an ordinary timeline event, so it competes with real messages for the rate limit — Synapse's default allows a burst of 10 and then one every 5 seconds. And it is permanent: only the latest renders, but every frame stays in the room forever and comes back out of
chat_history()on the next restart.So the trail is pushed when its text actually changes, with a five-second floor. A fixed 1s tick would spend 600 permanent events on a ten-minute turn, and the frames it lost to throttling would be the ones at the end, where the work finished.
Liveness stays the typing indicator's job — an ephemeral EDU that costs nothing and was already being sent. The edits are for what the agent is doing, which changes a dozen times in a long turn, not six hundred.
Lifetime
The observer is attached per turn and removed after. A session outlives the turn, and a leftover one would keep appending to an accumulator whose message was already finalized — the next turn's first tool call would silently edit the previous turn's trail. Restored on error too, or one failed turn doubles the trail on every turn after it.
Verification
2919/2919. 10 mutations, all caught.
Three found real problems rather than test gaps:
acc$last_text <- NULLbefore the final flush was dead code —flush()never consulted the interval floor. It was hiding a wasted event: every turn whose last tool call was more than five seconds before it finished ended by re-sending its own text as a permanent event. The skip now lives inflush(), where a duplicate frame is never useful whoever asks for it.<summary>element and write into the room. Escaped, with a test that tries..diff_summary_counts()now rides along on the diff payload.linesis clipped to a budget, so anything recounting from it undercounts a large diff, and the only other place the true numbers survived was inside an English sentence.Not settled
The five-second floor is a starting point, not a measurement. Worth watching in a real room and comparing against a fixed tick — it is two lines different.