Skip to content

WIP: refactor(agent) runtime around event-sourced protocol boundaries - #47

Draft
Fodesu wants to merge 68 commits into
felinics:mainfrom
Fodesu:feat/agent-runtime
Draft

WIP: refactor(agent) runtime around event-sourced protocol boundaries #47
Fodesu wants to merge 68 commits into
felinics:mainfrom
Fodesu:feat/agent-runtime

Conversation

@Fodesu

@Fodesu Fodesu commented Aug 28, 2026

Copy link
Copy Markdown
Member

Summary

see spec

@Fodesu
Fodesu force-pushed the feat/agent-runtime branch from bf4836b to 0d80569 Compare August 28, 2026 12:50
Fodesu added 21 commits August 29, 2026 14:33
Spec: docs/design/agent-runtime-refactor.md — Machine (Decide/Evolve/Next),
command/fact vocabularies, EvaluateCommit, dual-materialization authority.

sdk: frozen Request/ToolDefinition/ToolChoice/BlobRef, ModelResult, ModelStream.
agent: canonical JCS encoding, digests, derive functions, 14 commands /
14 facts, MachineState, Decide/Evolve/Next, EvaluateCommit, MemoryRuntime,
Loop with bounded parallel tool execution. Conformance + machine + loop tests.
EvaluateCommit now rejects AcceptInput/Approve/Reject/SubmitToolResponse
carrying a caller-minted CommandID: the derivation IS the idempotency index
(spec §5.5), so a random ID silently bypassed duplicate-input detection.

Loop.planAndPrepare distinguishes concurrency rejections from content
rejections: a retriable Prepare error with no authority progress (same
Revision after reload) is surfaced instead of retrying the same plan forever.
Authority correctness:
- SubmitModelResult gates on the result's own tool calls: a result with
  calls and zero bindings no longer silently completes the run
- bindings are cross-checked against the model result (tool name via the
  frozen spec, canonical arguments); a self-consistent binding for a tool
  the model never called is rejected
- ToolStepOpened carries BindingSetDigest; Evolve folds it verbatim, so
  DeriveToolStepID(Source, StepRef.Digest) reproduces the step ID and a
  future schema bump cannot change how v1 events fold
- Evolve and ToolStepOpened validate ToolCallState combinations (spec 4.2)
- CancelRun's reason is fixed to cancelled; ingress cannot forge step_limit
- RejectToolCall accepts Waiting calls of either kind, giving abandoned
  ask-user calls a known-failure exit instead of stranding the run

Canonical encoding:
- integer tokens keep arbitrary precision (64-bit IDs above 2^53 no longer
  corrupt or collide digests); floats keep the ES6 double form
- duplicate object keys, trailing data and invalid UTF-8 are rejected
  instead of silently merging distinct payloads into one digest

Loop:
- settleWorkers replays a failed completion commit once, then surfaces the
  error instead of silently wedging the call in Executing
- a panicking tool settles as ToolExecutionUnknown instead of crashing the
  process
- streaming consumption has a ctx escape and a nil-result guard; reasoning
  deltas are forwarded (EventModelReasoningDelta)
- EventRunFinished fires on terminal
- bindings use ToolSpec.Ref (catalog key), fixing aliased tools
- PlanningHint.SourceStep populated from the new LastClosedStep field

Runtime:
- MemoryRuntime Load/Commit/Events return deep copies; caller mutation can
  no longer reach authoritative state or committed event bytes

Tests: golden digest assertion is now fatal; parallel-bound test gates
workers to prove real concurrency; regression test per finding.
Also: sdk.Usage.Add replaces the duplicated adder; ModelResult.Response is
a pointer so absent metadata stays out of fact digests; four wording fixes
in the design doc.
The AgentEvent log is now the source of truth; MachineState is the
required same-transaction projection (execution cache), rebuildable by
folding the log with Evolve alone — no Decide re-run, no command replay,
no external effects.

Spec (§5.1/§9.1/§14.3/§17, appendix B.11): authority declaration flipped;
the three stability conditions (sealed fact ontology, frozen Evolve
semantics, self-contained facts) are normative; a per-run revision
watermark witnesses log-tail completeness. Arbitration: snapshot
divergence or loss with a complete log rebuilds automatically and is
reported for audit; a log tail below the watermark halts the run
(ErrLogTruncated) — accepted facts cannot be recovered from nothing, and
continuing would upgrade the loss into repeated execution.

Protocol: ModelStepPrepared carries BindingDigest (computed by Decide,
folded verbatim by Evolve), closing the last self-containment gap.

Code: FoldEvents verifies (Revision, Index) continuity and per-fact
digests while folding; MemoryRuntime gains the watermark and Rebuild as
the reference arbitration implementation.

Tests: healthy rebuild is a no-op; corrupted snapshot repairs from the
log; truncated tail halts; interior gaps and tampered facts are rejected;
golden v1 event stream folds to frozen state bytes.
Persist agent-owned JSON-stable model request/result/tool data across commands, facts, state, and digest inputs; enforce canonical snapshots and rebuild MemoryRuntime authority from the event log.

Add SDK Request/ModelResult single-call entrypoints, compatibility adapters, and optional ModelInvoker interfaces while keeping legacy GenerateText/StreamText wrappers.
Add command/event JSON wire codecs that restore sealed variants and verify digests, move Runtime conformance into reusable agent/runtimetest, and harden FoldEvents replay validation.

Reject legacy provider fallback when Request.ProviderOptions would be silently dropped.
Move canonical JSON into agent/jsonstable.Value and use CanonicalJSON for persisted command, fact, state, request, result, metadata, tool argument, and response payload fields.

Harden wire decoding against ambiguous JSON shapes and enforce response decision/payload digests at Decide.
The lint fix that inlined the conformance suite into package run made it
_test.go-only, which no other package can import — but spec §2.3/§8.2
promise the suite to durable Runtime adapters (Memoh) as their acceptance
gate.

Restore the net/http/httptest layout: agent/run/runtimetest is a normal
package exporting RunConformance(t, factory) and importing run; run's own
conformance entry point moves to an external test package (package
run_test), so the dependency chain run_test -> runtimetest -> run has no
cycle and golangci-lint 2.11.3 (the CI version) reports zero issues with
no .golangci.yml exclusions.

The suite already used only the public run API; it moves verbatim modulo
package qualification. Also drops the stale 'Package runtimetest'
comment that sat on top of package run.
Upstream b561c3d renamed the module github.com/memohai/twilight-ai ->
github.com/memohai/twilight before this branch's packages landed, so every
cross-package import inside agent/{es,run,runtimetest} and the sdk test
helpers still pointed at the old module. This is what CI's typecheck
reported as 'no required module provides package .../agent/jsonstable' —
the path simply no longer names this module.

Rebased onto upstream main and rewrote the import paths; the remaining
twilight-ai strings are upstream's own User-Agent / MCP client display
names, unchanged on main.
@Fodesu
Fodesu force-pushed the feat/agent-runtime branch from 6dc5728 to ffe9493 Compare August 29, 2026 06:37
Fodesu added 7 commits August 29, 2026 14:57
RunHeader was specified (§5.1.1) but never implemented: BuildRunHeader
creates the immutable Revision-0 record (minimal InitializeRun state,
digest-bound, causation-linked); ValidateRunHeader rejects tampering and
non-minimal initial states; FoldRun validates the header then folds the
transition log — the entry point for durable adapters and run import,
which must never trust an uploaded snapshot.

Spec: mark agent/session as contract-first (the §2.6 materialization
contract binds now, on the application's existing storage; the generic
substrate waits for its first real consumer) and agent/queue as deferred
until a second consumer exists. Rollout order flips: the durable
application adapter is now phase C — the Runtime contract's first real
external consumer — ahead of any new packages; session/queue become
need-driven phases D/E and the harness drops to example status.

docs: add the Memoh durable adapter work order (DDL sketch, commit
transaction skeleton, materialization contract on existing tables, queue
integration points, recovery scanner, acceptance checklist).
The work order describes Memoh-side implementation (DDL, transaction
skeleton, acceptance checklist); it lives with the code that will
implement it, not in the library repo. Moved to
Memoh/docs/design/durable-run-adapter-workorder.md.
Two pages: the legacy in-process SDK loop (authority in the library's
stack frame, callback weaving, deferred-approval single slot, manual
state capture — each red block a compensation Memoh had to build,
grounded in internal/agent/runtime/native and migrations 0073/0121) and
the Run ES design (authority in the runtime transaction as
RunHeader + TransitionRecord log, single Decide/Evolve implementation,
per-call response routing, watermark arbitration). The closing panels
map each old compensation to the mechanism that removes it.
Superseded by the progressive narrative page: the single-canvas layout
buried the argument; the replacement introduces one concept per chapter.
run/header: ValidateRunHeader's minimal-state check now also rejects
pre-seeded Usage and a forged LastClosedStep — both fields passed digest
verification as-is, letting an imported header inflate every RunResult's
usage or steer the first PlanningHint.SourceStep (spec §5.1.1 rule 2).

run/loop: planAndPrepare now emits the accepted ModelStepPrepared events;
it was the only accepted transition invisible to EventSink observers,
so projections keyed on committed events missed the frozen request. The
spec §6.6 one-shot same-CommandID commit replay moves from settleWorkers
into the shared commit helper, so model completion, start, StopRun and
prepare submissions get the same protection — a transient commit failure
on SubmitModelResult no longer aborts the Loop into a state that
MemoryRuntime (no lease expiry) can never recover.

sdk/request_adapter: the stream path materialized Response as a non-nil
pointer to a zero ResponseMetadata where the generate path yields nil,
so FreezeModelResult persisted different bytes per execution mode and
ModelStepCompleted digests were not reproducible across Streaming
on/off. Absent metadata now stays nil on both paths.
…record

The durable-continuation requirement — a new process resumes the same
run from the same step — needs durable MachineState plus lease/grant
coordination; it does not need the state to be rebuilt from events.
Event-log authority bought two capabilities on top of that (automatic
refold after state corruption, historical execution fork): the first
never fires on a correct implementation and has a manual FoldRun +
backup-restore fallback, the second has no product need and is
semantically unsound for executions with external side effects.

Spec §5.1/§8/§9.1/§14.3/§17/appendix B: MachineState is the Run
execution authority (recovery = Load); RunHeader + TransitionRecord log
is the same-transaction canonical record serving audit, projections,
and verified run import/migration via FoldRun. The replay-fold
equivalence stays as the dual-write correctness test. Divergence is an
implementation defect handled operationally; the watermark and the
arbitration protocol leave the spec. Evolve relaxes from permanently
frozen to stable-within-version; migration shipping re-tightens it via
the golden-stream check (recorded as the trigger, alongside execution
fork, for reconsidering log authority).

Code: Rebuild/ErrLogTruncated demoted to optional diagnostics in
comments; watermark kept as a diagnostic field only. Memoh work order:
watermark column dropped, rebuild acceptance item replaced by a
FoldRun consistency check. ES doc section 5 retitled to durable
execution state + canonical record; recovery narrative now reads
Load-not-fold, with the two layers' deliberately different answers to
the ES litmus question.
With MachineState as the execution authority, deriving the stored state
by refolding the just-persisted transition was a leftover of the
log-authority era. Commit now stores decision.NewState directly; the
state/log equivalence stays guarded by the conformance replay-fold test.
Stale 'log is the source of truth' and 'frozen forever' comments updated
to the final authority wording.
Fodesu and others added 30 commits September 1, 2026 15:17
ProtocolFor(Header.SchemaVersion) returns the implementation for
digest, decode, Decide, Evolve, and BuildEnvelope. Loop and
EvaluateCommit call those methods without a version argument.
v1 replay uses digestRequestV1 directly instead of re-dispatching
through DigestRequest(schemaVersion, ...).
Protocol is a struct whose fields are the digest, decode, Decide,
and Evolve functions for one schema version. ProtocolFor fills them
once from the header. There is no protocolV1 type; v2 adds a
ProtocolFor branch and the functions that actually differ.
SubmitModelResult snapshots ToolExecution and MaxParallel onto
ToolStepOpened. Resume honors the frozen ToolStep.Scheduling instead
of the live Loop ExecutionPolicy. Empty mode stays empty so default
parallel does not enter the v1 golden state digest. ModelCatalog must
resolve a ModelRef to equivalent execution for the life of a Run.
ModelCatalog.Resolve failure or a nil invoker recovers the ModelStep
and returns from Loop. The model was never invoked, so the Run stays
active instead of ending as provider_failure.
CancelRun still ends as stopped/cancelled, but Executing tool calls
and an Executing ModelStep are recorded on RunStoppedEnd and copied
onto RunResult as UncertainCalls and UncertainModel.
RejectToolCall on ApprovalRequired stays permission_denied.
Rejecting an ExternalResponse wait records response_rejected.
The spec no longer keeps ToolStepClosed as a compatibility fact.
Update type listings for RunResult, ToolScheduling, and Protocol.
Document Protocol.BuildEnvelope, FoldRun per-record schema, Next's
terminal error, Evolve closing ToolStep, and NeedsRecovery as the
ExecutionRecovery signal including Model Executing.
List ModelStep and the full Protocol method set. Qualify that Loop
stops starting DirectExecution calls when the outer context is
cancelled. Record that Cancel without an Evolve close leaves
LastToolStep unchanged. Extend Loop conformance with catalog
resolve, frozen scheduling, and uncertain cancel projection.
Runtime evaluates commits; Store persists header, state, log, and
leases. MemoryStore is the in-process Store. NewRuntime(store) is the
single Runtime implementation. Expired leases set recoveryValid so
grantless Recover works; zero deadline never expires.
Keep a single Runtime constructor over MemoryStore. Rebuild is a Store
operation so diagnostic refold does not depend on a Memory-only type.
Unknown now records ToolCallFailed for that executing call and leaves
the Run active. Occupancy without a step is Open, so AcceptInput and
Prepare no longer key off a nil Current. Feature coverage moves into
the runtest driver.
EventType is twilight/module/name. Chatlog owns content, Turn owns
round and Run linkage, and the reference assembly records Binding,
Planner, and the shared user-text payload.
Keep one package. Loop construction and Run stay in loop.go; start
caches, model execution, tool settlement, and EventSink live beside it.
Keep empty ToolExecution unset until freeze. Resume executing tools
without grantless Known failures. Export RecoverExpired, bind grantless
model recovery to the lease claim, and FoldRun through Rebuild.
…dec and lease renewal

Store contract
- Replace View/Update(fn)/ListIDs with LoadHead, LoadLog(from), LoadRecord,
  LookupTransition, Commit(fn RunTx) (*Append), RenewLease, ExpiredLeases,
  ReplaceSnapshot. No method can delete or rewrite a transition.
- Store.Commit is the Run's critical section: Runtime evaluates against the
  RunTx head and the Store persists the returned Append in the same
  transaction (RUN-CMT-2). Load and Commit no longer touch the log.
- Drop the StartGrants table; the grant lives only on its lease.

Snapshot codec
- Protocol.EncodeMachineState/DecodeMachineState define a persisted wire for
  MachineState including Current. SQLite stores snapshot + log; Record and
  Rebuild verify the snapshot through FoldRun. Golden digests unchanged.

Lease renewal (RUN-CMT-8)
- Runtime.RenewLease extends a live lease; Loop workers heartbeat at
  ExecutionPolicy.LeaseRenewInterval and stop when renewal is rejected.
  LeaseTTL bounds recovery delay instead of tool duration.
- RecoverExpired loads only Runs returned by Store.ExpiredLeases.

Protocol surface
- Remove package-level Decide/Evolve/Digest*/BuildEnvelope/EncodeCommand,
  BuildRunHeader and currentSchemaVersion; version binds once at the Run
  boundary via ProtocolFor or RuntimeSnapshot.Protocol().

Specs
- agent-run.md: append-only Store section, RUN-CMT-8, updated RUN-WIR-3,
  RUN-MCH-3, RUN-CMT-2, RUN-CMT-7, RUN-LOP-1, RUN-CMP-2.
- Session, Artifact, Session Extension downgraded to draft until a Memory
  vertical slice passes conformance.
- Refactor doc: lease/recovery belongs to run.Runtime; sdk.Request freeze
  evaluation recorded in 4.4; sdk/request.go digest comment corrected.

Conformance
- New RunLeaseRenewalConformance; snapshot codec round-trip and malformed
  wire tests; SQLite reopen test; Loop long-tool heartbeat test. MemoryStore
  and SQLite pass Runtime, recovery and renewal suites; go test -race clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- evolve.go: validateFactTransition (cyclomatic 96) becomes guardFactV1
  dispatching to one guard per fact, plus small apply* folds. Guards share
  requireOpen / requireModelStep / requireCall / requireWaitingFor;
  ToolCallFailed reuses ValidateToolCallState for class/outcome agreement.
- state.go: ValidateMachineState (cyclomatic 48) split into
  validatePendingInputs / validateLastToolStep / validateCurrent /
  validateCurrentToolStep. Add String() on ModelStepStatus and
  ToolCallStatus, ToolCallStatus.Terminal(); move current() markers next to
  their types.
- Remove MachineState.LastClosedStep: it duplicated LastToolStep.RefValue.ID
  and needed a consistency check. PlanningHint.SourceStep now reads from
  LastToolStep. Snapshot wire drops lastClosedStep; v1 header and event
  stream golden digests re-frozen (pre-release).
- decide.go: waitingCall returns only error (unparam). Remove unused
  jsonMarshal.
- agent-run.md: MachineState block and closing rule updated.

gocyclo findings on agent/run: 2 -> 0. go test -race ./agent/... clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A godoc example that is also the first external caller of agent/run: a
host composes run + loop + sqlitestore, creates a Run, accepts input and
drives the Loop until a tool call is Executing. The Store is closed to
simulate the process dying with a live lease and no settlement.

A second Runtime reopens the same database after the lease TTL:
RecoverExpired settles the abandoned call as Unknown, the Run stays Active
on the same RunID, the Loop re-plans from the committed tool outcome and
completes. Record verifies all 9 transitions against the stored snapshot.

The example exercises PlanningHint.LastToolStep, NeedsRecovery,
ExecutingCalls, RuntimeOptions.Now, LeaseRenewInterval and the split
model/tool catalog interfaces from outside the package.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Every command identity of one execution attempt now derives from its claim:
DeriveStartCommandID, DeriveSettlementCommandID, DeriveModelRecoveryCommandID
and DeriveToolRecoveryCommandID (RUN-WIR-3). EvaluateCommit enforces the
start derivation, so a hand-minted start id is rejected before it can mint
ownership. Tool recovery no longer writes the grant into the log.

Loop keeps a single ClaimStore instead of the starts/settlements caches;
memoryClaims is the default, ExecutionPolicy.Claims injects a durable one.
A replacement Loop sharing the store replays the derived start, recovers
the live grant and settles without waiting for lease expiry
(TestLoopReplacementFinishesInheritedClaim). Known failure of a Pending
call derives from the call alone.

Test helpers treat literal start ids as attempt labels; spec identity
table, RUN-LOP-3 and RUN-CMP-2 updated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The vertical slice of agent-turn.md: one Turn owns one primary Run.
Start appends started + input_delivered, creates the Run, accepts inputs
under their derived CommandIDs and drives. Resume rebuilds the Turn from
the log and the Runtime (no coordinator state), materializes every
transition above the coverage watermark through the v1 FactMapper
(ModelStepCompleted -> assistant, ToolCallCompleted/Answered/Failed ->
tool_result success/error/unknown), drives, then settles completed/failed
from RunEnded. Stop cancels under a Turn-derived CommandID.

Log is a Seq-addressed MemoryLog standing in for the Session kernel; the
Session, Extension and Artifact specs stay drafts. LoopDriver is the
reference RunDriver.

Tests: completion with materialization and idempotent Start; approval wait
and Resume; Stop settling as failed/stopped; Resume by a fresh coordinator
after a crash mid tool call and lease expiry.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ProtocolV1()

- EvaluateCommit: an envelope digest mismatch is a construction fault and
  returns a hard error instead of ErrCommandConflict, so callers no longer
  reload and retry it.
- Evolve: a duplicate InputAccepted is guarded as a corrupt log rather than
  silently deduplicated; Decide already rejects it and exact replay never
  reaches Evolve.
- decideSubmitModelResult split into checkToolCallBindings /
  checkBindingAgainstResult / openToolStep.
- RunEnded wire is now a tagged union ({"completed":{}} | {"stopped":{..}} |
  {"failed":{..}}) mirroring the Go sealed union; the flat status/reason
  shape and legacyEnd are gone. Golden digests unchanged (RunEnded is not in
  the state snapshot preimage).
- ModelCatalog.ResolveModel and ToolCatalog.ResolveTool replace the two
  same-named Resolve methods so one type can implement both.
- ProtocolV1 is a function; the package-level var could be reassigned.
- agent-run.md and refactor doc 4.3 updated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…iderCallID

The Run's CallID is now DeriveCallID(source, index): an authority-owned
identity that enters facts, lease keys, derived CommandIDs and chatlog, and
cannot collide across ModelSteps or depend on provider behaviour. The
model's tool_call_id rides along as ToolCallBinding.ProviderCallID /
ToolCallState.ProviderCallID, used only when a Planner echoes the call and
its result back to the model. Decide enforces both: CallID must be the
derived value, ProviderCallID must match the result at that index. Empty or
repeated provider ids (vLLM/llama.cpp style call_0, prompt-parsed tools)
no longer reject the result.

- loop.bindToolCalls derives ids and no longer treats duplicate provider ids
  as malformed; invalid UTF-8 input remains the malformed path.
- turn.MapTransition takes the record prefix so tool_result events carry
  the ProviderCallID from ToolStepOpened; payloads gain providerCallId.
- Planners in example, runtest and live test echo ProviderCallID.
- Test fixtures build bindings by (step, index, providerID); runtest resolves
  provider ids to derived CallIDs. New regression test covers provider id
  reuse and a forged non-derived CallID.
- v1 event-stream golden re-frozen (pre-release). Spec identity table,
  RUN-MCH-2 and TRN-MAP-2 updated.
- live_test: full payload logging, TWILIGHT_LIVE_RECORD_OUT dumps the
  RunRecord, 6 minute deadline.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The live run showed the second model request had no user message: the
example planner read only PlanningHint.Inputs, which the first prepare had
consumed. PlanningHint is boundary facts by design; the conversation must
come from the session log.

- ContextFold projects a session log into ordered Entries (delivered
  input, assistant, tool_result) across every Turn (CHT-CTX-1/2, minimal:
  no summary/checkpoint yet).
- ContextPlanner is the reference RequestPlanner (REF-PLN): system prompt,
  the fold as user/assistant/tool messages, then the tool step that closed
  inside this Loop.Run and is not yet materialized, taken from
  hint.LastModelResult/LastToolStep. sdk.Message is produced here and never
  stored. Tool results echo ProviderCallID and the model-facing tool name.
- ToolResultPayload gains Name; the mapper indexes provider id and name per
  derived CallID from ModelStepCompleted across the record prefix.
- Tests: the second request now carries system/user/assistant/tool; a second
  Turn on the same session sees the first Turn's whole conversation. The
  live test uses ContextPlanner; request 6 of the live record now includes
  the user message.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Replace the two-log design (per-Run RunHeader + TransitionRecord store,
Turn materialization into the Session) with one Session stream:

- agent-run.md: Run facts are twilight/run/ events; CommitID = CommandID;
  Runtime commits inside the Session critical section; MachineState is the
  twilight/run/machine projection. Facts keep only digests: request bodies
  go to a content-addressed FrozenValueStore, model/tool output goes to the
  chatlog companion events in the same commit (RUN-WIR-4). Prepare hard CAS
  keys on RunPosition, not the Session head. Companion/Attach on Commit.
- agent-turn.md: Turn:Run 1:N with attempt_failed / Retry / Settle;
  companion mapping replaces FactMapper, MaterializeAll, ResultReference,
  coverage and outbox; Stop settles via Attach.
- agent-session.md: add CommitIn critical-section port alongside CAS Commit.
- extension/chatlog/reference-assembly: run as first-party module,
  ProviderCallID and SourceDigest on chatlog parts, single-store composition.
- agent-runtime-refactor.md: record the decision, the byte analysis behind
  it, and the code migration list.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Apply the review findings on the single-Session-ES design:

- One write path: SemanticAppender gains AppendSemanticIn; run.Runtime
  writes through it, so companion/Attach events get codec, binding
  admission and claims in the same transaction (EXT-SCP-1, EXT-APP-3).
- Session Store gains a control-plane KV (SES-API-3) readable/writable
  inside SessionTx; lease, grant and artifact claims live there and
  commit atomically with the event group. RecoverExpired adds a
  no-lease fallback keyed on start time + TTL.
- Kernel ProtocolVersion covers envelope/commit only; payloads carry a
  top-level `v`, Registry keeps every codec version (SES-VER, EXT-REG-2);
  Run keeps its own created.SchemaVersion.
- v1 scope: Fork/ancestry/import, Application modules/Catalog, two-phase
  journal, artifact Prepared state/reconciler/import move to appendices.
- Snapshot is a droppable cache with SnapshotPolicy; MachineProjection
  drops terminal runs; turn surface records AttemptView.End.
- Run keeps only an opaque OwnerID (turn fills TurnID); Companion returns
  ModuleEvent; Attach carries the Stop settlement.
- Attempt-content policy moves to the reference Planner (REF-PLN-6).
- Fingerprint excludes RecordedAtUnixMilli; Replay/Tail gain EventType
  prefix filter; CoverageDigest removed in favor of Through.Digest.
- agent-runtime-refactor.md §7 records the review, the persistent
  structure/consistency table and the implementation order.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ncy pass

Kernel / framework layering
- Session control-plane KV gains ControlCompareAndPut, per-entry deadline
  and ControlExpired; the kernel carries no lease concept (SES-API-3).
- extension.Lease is the shared occupancy facility: Acquire/Release inside
  SemanticTx, Renew via conditional put, Expired via deadline scan, Token
  derived from the acquiring CommitID (EXT-LSE). run is its first consumer;
  the projection fallback scan and "renew via session lock" are dropped.
- ModuleDescriptor.Requires declares event-consumption dependencies; the
  Registry validates registration, acyclicity, projection Consumes and
  payload-version compatibility at build (EXT-REG-4, EXT-SCP-4).
- BuildRegistry takes descriptors from composition; extractors are declared
  per EventDefinition (BindingExtractor), so extension imports no module.
- ProjectionReader is the only out-of-section projection read path;
  Coordinator no longer holds session.Store.
- EventID for all first-party events = Digest(EventType, CommitID, index),
  assigned by the Appender (EXT-APP-5); per-module rules removed.

Mid-turn input
- Turn gains Deliver: one Run commit per input (AcceptInput + attached
  chatlog/input_delivered), never interrupting in-flight calls (TRN-DLV).
- AcceptInput is accepted in any non-terminal state; a model result with no
  tool calls but pending inputs returns to Open instead of ending; new
  WithdrawPreparedStep discards a frozen-but-unsent request when input
  arrives (RUN-MCH, RUN-LOP-8). turn surface tracks run/input_accepted;
  Retry replays all delivered inputs.
- Reference assembly adds SessionDriver (Send / OnTurnSettled) mapping the
  inbox model's next-step / next-turn onto Deliver / Start (REF-DRV).

Cleanup
- Status lines without dates or revision history; stale terms removed
  (resolved ancestry, RegistryID/Parts, ToolIndeterminate, InitialInputs).
- Redundancies removed: StartRequest.Inputs only; BindingPublic.Tools as
  PublicTool; Lease Attrs dropped for run (target parsed from Key).
- Gaps closed: Create idempotency, CausationID/CorrelationID semantics,
  ProjectionKey = ProjectionID, Record compares only active runs.
- agent-runtime-refactor.md §7.6–7.8 record the decisions.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Reshape agent/run per agent-run.md so facts carry execution state and
content digests only (RUN-WIR-4):
- ModelStepPrepared keeps RequestDigest; the request body lives in a
  FrozenValueStore (Runtime.FrozenRequest, MemoryFrozenValues).
- ModelStepCompleted/ToolCallCompleted/ToolCallAnswered carry digests;
  MachineState drops LastModelResult, RunResult drops Model,
  ToolSpec drops Definition and gains Name.
- RunCreated fact and BuildCreateGroup; MachineState gains Owner/Attempt.
- AcceptInput legal in any non-terminal state; WithdrawPreparedStep and
  Next -> WithdrawPrepared; no-call result with pending inputs reopens.
- PlanningHint carries boundary facts only (Owner, RunID, SourceStep).
Tests and runtest/runtimetest drivers follow; golden fixtures re-frozen.
Old agent/turn is behind the legacy_turn build tag until rewritten.
…nce assembly

Implement the 2026-09-04 decision (agent-runtime-refactor.md 6-7) as code:
- agent/session: Memory Store with Commit (CAS), CommitIn, Types-filtered
  replay, snapshot and control-plane KV (CAS, deadline enumeration).
- agent/artifact: Ref, Binding, Memory BindingStore, two-state KV ledger.
- agent/session/extension: Registry with Requires checks and payload v,
  SemanticAppender (both entries), Lease, ProjectionReader.
- agent/session/chatlog: events, parts codec, Surface and Context.
- agent/session/run: twilight/run/ events, machine projection, Runtime
  (EvaluateCommit inside the Session critical section, lease and recovery).
- agent/run: Session-addressed Runtime contract, RunPosition; per-Run
  Store, sqlitestore, RunHeader and TransitionRecord removed.
- agent/run/loop bound to a Session; agent/turn rewritten (Coordinator,
  CompanionV1, surface); agent/ref assembles everything with an example.
Drop the twilight/run/command control-plane index that stored a command
digest per CommitID to tell an exact replay from a same-ID conflict. Every
Run CommandID is content-derived, so a LookupCommit hit is the same command;
the two families whose identity deliberately omits content (approve/reject
of one ResponseID, two settlements of one attempt) are read back from the
projection instead. RUN-CMT-5 and the refactor record updated.
…pes from the surface

settle returns the terminal RunResult when a model settlement ends the Run,
and Loop.Run finishes from it instead of reloading a Run the machine
projection no longer holds (RUN 7). AttemptView records SchemaVersion so
Deliver and Stop construct envelopes without reading the machine projection;
AcceptInput and CancelRun carry no Base (RUN-CMT-4). Load on a terminated
Run keeps folding its events as the fallback (RUN-CMT-1).
Revise RUN-CMP-2 to assert Run semantics only (atomicity, digest chain and
snapshot equivalence are referenced from the kernel and extension suites),
add the decisions of the day (terminal Load, settlement snapshot, zero Base
for non-Prepare commands, holder-bound tool recovery) and bind grantless
tool recovery to the lease holder in RUN 5.1 / RUN-CMT-6. TRN-MAP-3 notes
that a failed call's tool_result has no SourceDigest.

agent/session/run/runtimetest implements the revised list against any
session.Store; it runs on the Memory store. The suite exposed one gap: the
machine projection accepted a second run_created for a terminated RunID, so
the projection now keeps the set of ended RunIDs (ids only).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant