Skip to content

Show a room what the agent is doing while it does it - #175

Merged
TroyHernandez merged 7 commits into
mainfrom
room-activity
Aug 8, 2026
Merged

TroyHernandez merged 7 commits into
mainfrom
room-activity

Conversation

@TroyHernandez

Copy link
Copy Markdown
Contributor

Depends on cornball-ai/chat.api#11.

A terminal prints tool calls as they run. A room got the typing indicator and then, some minutes later, a wall of prose — so a turn that was working and a turn that had hung looked identical.

The cause turned out to be one line long: rooms.R registered no observers at all. The terminal registers two. add_observer() is generic and session-scoped, the events already carry call$tool, outcome, success and diff, and none of it was reaching a room.

What a room sees now

One collapsible line, posted on the first tool call and edited as the turn goes:

Ran 3 commands, read 2 files  +102 -33     ▸

<details> and <summary> are in Matrix's allowed HTML subset and Element renders them natively, so it is a real disclosure widget rather than a heading over a wall of text. The plain body carries the same summary in one line — that is what a client without HTML shows and what the push notification says, so it has to stand alone.

The cadence, and why it is not a timer

Two costs shape this and neither exists in a terminal.

A Matrix edit is an ordinary timeline event, so it competes with real messages for the rate limit — Synapse's default allows a burst of 10 and then one every 5 seconds. And it is permanent: only the latest renders, but every frame stays in the room forever and comes back out of chat_history() on the next restart.

So the trail is pushed when its text actually changes, with a five-second floor. A fixed 1s tick would spend 600 permanent events on a ten-minute turn, and the frames it lost to throttling would be the ones at the end, where the work finished.

Liveness stays the typing indicator's job — an ephemeral EDU that costs nothing and was already being sent. The edits are for what the agent is doing, which changes a dozen times in a long turn, not six hundred.

Lifetime

The observer is attached per turn and removed after. A session outlives the turn, and a leftover one would keep appending to an accumulator whose message was already finalized — the next turn's first tool call would silently edit the previous turn's trail. Restored on error too, or one failed turn doubles the trail on every turn after it.

Verification

2919/2919. 10 mutations, all caught.

Three found real problems rather than test gaps:

  • The poll-loop wrapper had no coverage. Removing it left the suite green, because every other test builds the accumulator directly. There is now an end-to-end test that drives a real poll and asserts the observer is on the session the turn runs on.
  • acc$last_text <- NULL before the final flush was dead code — flush() never consulted the interval floor. It was hiding a wasted event: every turn whose last tool call was more than five seconds before it finished ended by re-sending its own text as a permanent event. The skip now lives in flush(), where a duplicate frame is never useful whoever asks for it.
  • Tool names are model-supplied and reach the markup. Unescaped, one could close the <summary> element and write into the room. Escaped, with a test that tries.

.diff_summary_counts() now rides along on the diff payload. lines is clipped to a budget, so anything recounting from it undercounts a large diff, and the only other place the true numbers survived was inside an English sentence.

Not settled

The five-second floor is a starting point, not a measurement. Worth watching in a real room and comparing against a fixed tick — it is two lines different.

A terminal prints tool calls as they run. A room got the typing
indicator and then, some minutes later, a wall of prose -- so a turn
that was working and a turn that had hung looked identical. rooms.R
registered no observers at all; the whole asymmetry was that.

The trail is one collapsible line -- "Ran 3 commands, read 2 files
+102 -33" -- posted on the first tool call and edited as the turn goes.
<details> and <summary> are in Matrix's allowed subset and Element
renders them natively, so it is a real disclosure widget rather than a
heading over a wall of text. The plain body carries the same summary in
one line, because that is what a client without HTML shows and what the
push notification says.

Two costs shape the cadence, and neither exists in a terminal. A Matrix
edit is an ordinary timeline event, so it competes with real messages
for the rate limit -- Synapse allows a burst of ten and then one every
five seconds -- and it is permanent: only the latest renders, but every
frame stays in the room and comes back out of chat_history() on the next
restart. So the trail is pushed when its text actually changes, with a
five-second floor, rather than on a timer. A fixed one-second tick would
spend six hundred permanent events on a ten-minute turn and lose the
frames at the end, where the work finished. Liveness stays the typing
indicator's job: an ephemeral EDU that costs nothing and was already
being sent.

The observer is attached per turn and removed after. A session outlives
the turn, and a leftover one would keep appending to an accumulator
whose message was already finalized -- the next turn's first tool call
would silently edit the previous turn's trail. Restored on error too,
or one failed turn doubles the trail on every turn after it.

Tool names are model-supplied and escaped before they reach the markup:
an unescaped one could close the summary element and write into the
room. Diffs report exact counts because .diff_summary_counts() now rides
along on the payload -- `lines` is clipped to a budget, so anything
recounting from it undercounts a large diff.
The wrapper in bot_poll() had no coverage: removing it left the suite
green, because every other test of the module builds the accumulator
itself. This drives a real poll and asserts the observer is on the
session the turn runs on, and that what it produces reaches the room.

It also pins the post-then-edit shape end to end -- one send on the
first tool call, one edit with the finished state -- and that the edit
targets the message the send created. The second tool call landing
inside the five-second floor is the reason there are two frames and not
three.
Mutation testing found dead code and a wasted event behind it. The final
flush set last_text to NULL first, which looked like it was bypassing
the interval floor -- but flush() never consulted the floor, so the line
did nothing except hide what it was actually doing: re-sending the
turn's own text as a permanent event, on every turn whose last tool call
happened more than five seconds before it finished.

The skip belongs in flush() rather than only in the observer's gate. The
floor is about spending the rate-limit budget during a turn; a duplicate
frame is never useful, whoever asks for it.
An encrypted Matrix room refuses edits: the replacement text would ride
in an ordinary event, so chat_edit() throws there by design. The
observer catches those errors, which meant an encrypted room got the
first frame -- "Ran a command" -- and then kept it for the rest of the
turn while every update was rejected. Nothing logged, nothing looked
wrong, and the room said something untrue for ten minutes. That is worse
than showing nothing, and it is the exact shape of failure this
migration has been finding all along: an error treated as an answer.

chat_capabilities()$edits is now asked once, before the first frame. A
client that can edit gets the live trail. One that cannot gets no
intermediate frames at all and a single final message, written after
everything happened and therefore accurate. Unknowable counts as cannot.

Found in review, not by me -- I had reported it as a missing disclosure
widget, which is the cosmetic half of it.
The capability check went in between rooms_with_activity()'s comment
and rooms_with_activity(), so the file read as though the paragraph
about attaching an observer per turn described the one-line predicate
underneath it.
@TroyHernandez

Copy link
Copy Markdown
Contributor Author

Fixed, and you were right that I'd described it as the cosmetic half. It isn't a missing widget — the room gets a stale one. First frame posts, every chat_edit() after it throws, the observer swallows the error, and "Ran a command" sits there for the rest of a ten-minute turn. Nothing logged, nothing looks wrong, and the room is asserting something untrue. Same shape as the fails-quiet bugs this whole migration kept turning up: an error treated as an answer.

Went with the capability gate rather than documenting it, because the third option is better than both:

rooms_activity_live(chat)   # chat_capabilities()$edits, once, before the first frame
  • Can edit → the live trail, unchanged.
  • Cannot edit → no intermediate frames at all, and a single final message written after everything happened. So an encrypted room still learns what the agent did; it just learns it once, and what it learns is true.
  • Cannot tell → treated as cannot. Posting frames that may never be updatable is precisely the failure being prevented, so an unreadable capability list fails to the safe side.

The one-accurate-message path is strictly better than what an encrypted room gets today, which is nothing.

chat.api#11 is merged (c10189a), so this PR's dependency is satisfied — CI should go green on the rerun.

2926/2926 locally.

One thing I'd still call unsettled, unchanged from before: the five-second floor is a starting point, not a measurement, and I quoted Synapse's rc_message defaults while corteza's own comments say the fleet runs Conduit. Worth watching in a real room before treating the number as load-bearing.

Two changes from review, both about what the trail costs a room.

The interval is what bounds the cost, not the content check. Almost
every completed tool call changes the summary -- "Ran 3 commands"
becomes "Ran 4 commands" -- so the change check filters very little
and five seconds was close to one permanent event per call. Fifteen
puts a ten-minute turn at forty events instead of hundreds, and
`activity_interval` in the config moves it.

Still not a timer. Nothing runs between tool calls, so a frame goes out
on the first completed call after the interval elapses rather than at
the instant it does, and the final flush covers whatever the last
interval did not.

And the trail is m.notice while replies stay m.text. That is what the
msgtype is for: automated output another bot should not answer, which
matters most in exactly the rooms this runs in. It is not a mute -- the
spec asks clients not to auto-respond and says nothing about push
gateways, so a room that wants silence still needs a push rule. Claiming
otherwise in a comment would be the more expensive mistake.

Needs chat.api 0.0.1.20: chat_edit() hardcoded m.text in m.new_content,
so the first edit turned the notice into an ordinary message.
…e one

corteza 0.7.1.21 needs chat_edit() with rich and kind -- 0.0.1.20 --
and declared 0.0.1.17. A host on 0.0.1.17 would have passed the runtime
guard and then found no chat_edit() at all, inside the observer's
tryCatch, leaving the room a stale first frame. Which is the exact
failure the capability gate was just added to prevent, arriving by a
different route.

Two bumps did it. Both were sed with a pattern that no longer matched,
so both silently changed nothing -- and the assertions pinning the
constant were edited by the same non-matching patterns, so they stayed
consistent with the stale value and went on passing. CI then verified
0.0.1.19 against a floor of 0.0.1.17 and went green: a verification step
reporting success for the thing it exists to catch.

The guard against a repeat reads the Suggests bound out of DESCRIPTION
and compares it to the constant, rather than restating a literal that
can only ever agree with itself.
@TroyHernandez

Copy link
Copy Markdown
Contributor Author

Correction to the green tick you reviewed — it was green on a stale floor, and I only caught it while merging.

.CHAT_API_MIN was 0.0.1.17, not 0.0.1.20. Two bumps I made were sed commands whose pattern no longer matched, so both silently changed nothing. The assertions pinning the constant were edited by the same non-matching patterns, so they stayed consistent with the stale value and went on passing. CI then reported:

chat.api 0.0.1.19 (>= 0.0.1.17)

A verification step reporting success for exactly the thing it exists to catch.

The consequence was real, not bookkeeping. corteza 0.7.1.21 needs chat_edit() with rich and kind — a host on 0.0.1.17 would have passed the runtime guard and then found no chat_edit() at all, inside the observer's tryCatch, leaving the room a stale first frame. Same failure the capability gate was just added to prevent, arriving through a different door.

Fixed, with a guard that reads the Suggests bound out of DESCRIPTION and compares it to the constant rather than restating a literal that can only ever agree with itself. Verified by mutation: bumping DESCRIPTION alone turns the suite red.

2939/2939. CI is rerunning against the corrected floor with chat.api 0.0.1.20 on main.

The general lesson I'd take from this one: a sed that does not match is indistinguishable from a sed that did its job, and I used it on version constants at least four times this session. The two places it bit were both cases where the same non-matching pattern was applied to the code and its test.

@TroyHernandez
TroyHernandez merged commit 9889f7e into main Aug 8, 2026
2 checks passed
@TroyHernandez
TroyHernandez deleted the room-activity branch August 8, 2026 00:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant