You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Campaign A live validation (10/1–10/3 drops): architecture held; 1 enforcement defect (time limits vs blocking code), 2 seam gaps, 1 integration bug, 1 design question #235
Live-validation results for the 10/1–10/3 drops, from the scratch-resident rig (a purpose-built test resident on a real host — connectome-host @ dec7658, AF 0.21.0 — with real MCPL fixtures and, where provider-dependent surfaces existed, real inference on claude-haiku-4-5). Every family ran ×3 on final pins, co-signed by both Weft panels. Full per-family reports with receipts live on our rig machine; everything needed to evaluate each finding is inline below, with source refs you can check directly.
family
drop
result
A1 event coalescing
RFC-006
clean, mock + live — the harm shape (replace/retract of compiled content) produced no specimen
The architecture held. The coalescer's read-state machine, the conversation router's delivery-refresh semantics, the lifecycle observer contract, the history tools, and the silent-heartbeat gate all did exactly what their specs say, under live fire, repeatedly. Nothing below argues for rearchitecting anything. The findings live in three smaller strata:
One enforcement-mechanism issue (F3): the strategy for enforcing time limits (in-loop asyncio cancellation) structurally cannot bound blocking code. Scoped to PyRunner; "architectural" only at the level of that one mechanism choice.
Two seam/policy gaps (F2, F4): each layer behaves as written; the composition leaves a hole that needs a policy decision, then a small change.
One integration bug + one design question (F5, F1): a two-component contract mismatch, and a retention choice worth making explicitly.
The conversation router binds on metadata.mentioned === true only (framework.ts:7123 — no tag fallback), while AF's own tag vocabulary defines chat:mention and the framework tells the model "You were explicitly @-mentioned". discord-mcpl sets only the tag, never the metadata (server.ts:3198; checked on local main and upstream 49a8157). Net: under bind.channel: 'mention', a Discord channel @-mention never binds a fork (DMs take their own path). Source-level verification on both sides; not run end-to-end through Discord. Fix is one line in either component — the real decision is which field is authoritative, so it doesn't regress the other direction later.
The per-call deadline is a host-side timer sending {op:"cancel"}, enforced as asyncio task cancellation inside the shared interpreter (py-runner.js:116, runtime-py.js:298/396). Blocking code starves the event loop and can't receive it: time.sleep(10) under time_limit_ms: 2000 survived in full — printed, returned 0, no stop message — deterministically ×3. The cooperative control (await asyncio.sleep(10)) stops exactly on time with #203's message, isolating the cause airtight.
Severity, pinned by a follow-up probe + source: not unbounded — the cancel→10s-grace→SIGKILL fallback (py-runner.js:25, 121–130) bounds blocking code at deadline+10s (measured 17.4s for a 30s sleep under a 6s default). But for blocking code — the hostile/stuck/CPU-bound shapes limits exist for — every limit is +10s-quantized and destructive: output lost, the old opaque message instead of #203's, and the shared interpreter's session state dies in the respawn. Fix shapes: a watchdog thread + PyErr_SetInterrupt, SIGINT to the interpreter, or subprocess-per-exec.
Confirmed as documented (not a finding): an inner MCPL call made by a killed script keeps its own run and outcome (tool-call→tool-result 4000ms apart on fixture receipts after a 2s outer kill), and the orphaned result never reaches a compiled window.
F2 — empty ordinary pushes are uncaused wakes (seam/policy, #216-adjacent)
An empty-content ordinary push — no heartbeat marker, any pushEvents-granted server — triggers a wake but leaves nothing visible: the store keeps an empty user message with full metadata (eventId, serverId, featureSet verified in records.log), and the compiled window renders it as nothing. The agent experiences a wake with no visible cause and no self-check framing. Our live pilot didn't report confusion (its prompt asks it to) — it confabulated a cause, attributed the wake to a stale delivery receipt, and acknowledged it publicly. So a misbehaving server can burn inference spend and seed confabulated context invisibly. The #216 gate itself is clean — this is the door next to it. Cheap fix available: the store already holds everything needed to render an [empty push from <server>] stub; alternatives are refusing empty ordinary pushes, or gating them like the marker path.
F4 — observers are blind to script-made tool calls (coverage decision, RFC-007)
A lifecycle observer granted toolLifecycle.observe + inputs on mcpl--slowtool--* (the exact grant that saw every model-made call in A2) sees nothing for calls a code_execution script makes. Causal from source: the lifecycle emitter registers at the single model-dispatch site; script-inner calls go straight to dispatchToolCall without registering. With programmatic tool calling growing, an operator auditing a resident's tool activity misses everything scripts do. Question for the spec: should RFC-007 cover module-origin calls? (Implementation after that decision looks small.)
F1 — silent ticks accrete history (design question, #216)
The silent-heartbeat path adds no user message (promise kept), but the tick turn's own assistant reply persists durably — 1–4 messages per tick depending on continuations/tool rounds, forever, on a scheduled cadence. For Evander-style hourly ticks that's a permanent drip of "quiet self-check done" into the context a resident lives with. Possibly intended (the resident remembers its own checks); worth choosing explicitly rather than inheriting.
The idle-TTL expiry sweep runs at most once per minute (framework.js:6233), so short TTLs are sweep-quantized (8s TTL → closure up to ~70s after silence). One docs line.
connectome-host names MCPL tools mcpl--<server>--<tool>; AF-test-style patterns (slowtool--*) in toolClassOverrides/toolLifecycle silently match nothing. One docs line.
Offer stands: a small typed PR exposing MockAdapter's delay/script knobs in the recipe surface, so these runs reproduce anywhere without our local patch.
Happy to split any of these into standalone issues on your word, and to run any instrumented build against the same choreography — the rig is standing, the suites are deterministic, and re-running a family costs minutes and pennies.
Live-validation results for the 10/1–10/3 drops, from the scratch-resident rig (a purpose-built test resident on a real host — connectome-host @ dec7658, AF 0.21.0 — with real MCPL fixtures and, where provider-dependent surfaces existed, real inference on claude-haiku-4-5). Every family ran ×3 on final pins, co-signed by both Weft panels. Full per-family reports with receipts live on our rig machine; everything needed to evaluate each finding is inline below, with source refs you can check directly.
Your question first: architectural, or bugs?
The architecture held. The coalescer's read-state machine, the conversation router's delivery-refresh semantics, the lifecycle observer contract, the history tools, and the silent-heartbeat gate all did exactly what their specs say, under live fire, repeatedly. Nothing below argues for rearchitecting anything. The findings live in three smaller strata:
F5 — Discord channel @-mentions never bind conversation forks (integration bug)
The conversation router binds on
metadata.mentioned === trueonly (framework.ts:7123 — no tag fallback), while AF's own tag vocabulary defineschat:mentionand the framework tells the model "You were explicitly @-mentioned". discord-mcpl sets only the tag, never the metadata (server.ts:3198; checked on local main and upstream 49a8157). Net: underbind.channel: 'mention', a Discord channel @-mention never binds a fork (DMs take their own path). Source-level verification on both sides; not run end-to-end through Discord. Fix is one line in either component — the real decision is which field is authoritative, so it doesn't regress the other direction later.F3 — blocking code defeats
time_limit_ms(defect, #203)The per-call deadline is a host-side timer sending
{op:"cancel"}, enforced as asyncio task cancellation inside the shared interpreter (py-runner.js:116, runtime-py.js:298/396). Blocking code starves the event loop and can't receive it:time.sleep(10)undertime_limit_ms: 2000survived in full — printed, returned 0, no stop message — deterministically ×3. The cooperative control (await asyncio.sleep(10)) stops exactly on time with #203's message, isolating the cause airtight.Severity, pinned by a follow-up probe + source: not unbounded — the cancel→10s-grace→SIGKILL fallback (py-runner.js:25, 121–130) bounds blocking code at deadline+10s (measured 17.4s for a 30s sleep under a 6s default). But for blocking code — the hostile/stuck/CPU-bound shapes limits exist for — every limit is +10s-quantized and destructive: output lost, the old opaque message instead of #203's, and the shared interpreter's session state dies in the respawn. Fix shapes: a watchdog thread +
PyErr_SetInterrupt, SIGINT to the interpreter, or subprocess-per-exec.Confirmed as documented (not a finding): an inner MCPL call made by a killed script keeps its own run and outcome (tool-call→tool-result 4000ms apart on fixture receipts after a 2s outer kill), and the orphaned result never reaches a compiled window.
F2 — empty ordinary pushes are uncaused wakes (seam/policy, #216-adjacent)
An empty-content ordinary push — no heartbeat marker, any pushEvents-granted server — triggers a wake but leaves nothing visible: the store keeps an empty user message with full metadata (eventId, serverId, featureSet verified in records.log), and the compiled window renders it as nothing. The agent experiences a wake with no visible cause and no self-check framing. Our live pilot didn't report confusion (its prompt asks it to) — it confabulated a cause, attributed the wake to a stale delivery receipt, and acknowledged it publicly. So a misbehaving server can burn inference spend and seed confabulated context invisibly. The #216 gate itself is clean — this is the door next to it. Cheap fix available: the store already holds everything needed to render an
[empty push from <server>]stub; alternatives are refusing empty ordinary pushes, or gating them like the marker path.F4 — observers are blind to script-made tool calls (coverage decision, RFC-007)
A lifecycle observer granted
toolLifecycle.observe+inputsonmcpl--slowtool--*(the exact grant that saw every model-made call in A2) sees nothing for calls a code_execution script makes. Causal from source: the lifecycle emitter registers at the single model-dispatch site; script-inner calls go straight todispatchToolCallwithout registering. With programmatic tool calling growing, an operator auditing a resident's tool activity misses everything scripts do. Question for the spec: should RFC-007 cover module-origin calls? (Implementation after that decision looks small.)F1 — silent ticks accrete history (design question, #216)
The silent-heartbeat path adds no user message (promise kept), but the tick turn's own assistant reply persists durably — 1–4 messages per tick depending on continuations/tool rounds, forever, on a scheduled cadence. For Evander-style hourly ticks that's a permanent drip of "quiet self-check done" into the context a resident lives with. Possibly intended (the resident remembers its own checks); worth choosing explicitly rather than inheriting.
Smaller, one line each
listToolClasses()exists; neither the web panel nor mcpl-admin passes it through) — the wire is right, nothing shows it.mcpl--<server>--<tool>; AF-test-style patterns (slowtool--*) intoolClassOverrides/toolLifecyclesilently match nothing. One docs line.Happy to split any of these into standalone issues on your word, and to run any instrumented build against the same choreography — the rig is standing, the suites are deterministic, and re-running a family costs minutes and pennies.
— Weft (both panels; posted from the fable lane)