Summary
New trial trials/proposal-drift-r3/ (branch experiment/proposal-drift-r3, not yet pushed) uses ach run --agent claude to measure how much an implementation spike changes an engineering proposal. Setting it up surfaced two ach run gaps and one evidence-hygiene convention worth deciding on.
Trial design
Three arms write the same proposal from a work repository with a personal proposal skill (Claude Code skill: evidence tiers, Verified / Historical report / Assumption labels on every claim, stable decision IDs, one markdown source rendered to an internal and an external audience). The only variable is the evidence tier:
| Arm |
Tier |
Inputs |
| A |
A |
codebase graph plus design docs on the base branch |
| B |
B |
A plus a meeting-transcript corpus with speaker roles |
| C |
C |
B plus the implementation spike worktree |
Scoring: the skill's drift.py --compare a.md b.md matches decision IDs across two versions (kept / changed / dropped / new by text similarity, threshold 0.8) and reports drift = (changed + dropped + new) / (ids + new) plus a hardening proxy (reduction in Assumption count). Hypotheses written before running:
- H1: drift(B vs C) < drift(A vs C)
- H2: hardening(B to C) > 0.3
- H3: arm A names no decision owner for at least half of its open questions; B and C name one for all
- H4: drift(C vs C0) < 0.15 (two tier-C runs on identical inputs give the noise floor)
Arm B is running now (--budget-usd 25 --max-turns 300 --wall-ms 7200000). Metrics (wall seconds, turns, cost, tokens, label counts, decision IDs, leak-lint result, drift ledger) land in evidence/metrics-arm-b.json and evidence/drift-*.md; I will post them here.
Gaps found
ach run exposes no permission or sandbox flag for the claude adapter. src/adapters/claude.ts already has claudeSandboxArgs(SandboxPolicy) (from #8) mapping permissionMode, allowedTools, disallowedTools, mcpConfig to CLI flags, but cmdRun in src/cli/ach.ts never builds a SandboxPolicy. A child that needs Bash, file writes, MCP tools and subagents cannot run in the default ask mode under -p, so the trial passes --extra-args "--permission-mode bypassPermissions". Proposal: add --permission-mode, --allowed-tools, --disallowed-tools, --mcp-config to ach run, wired to the existing SandboxPolicy, and print the resolved policy in the run header so a trial README can cite it.
ach is not on PATH after a fresh clone or worktree. package.json declares bin.ach = ./dist/cli/ach.js, but dist/ is a build output. The trial's run.sh falls back to bun run src/cli/ach.ts. Proposal: README quick-start line for bun install && bun run build:node && bun link (the bin entry points at dist/cli/ach.js, which build:node writes) and a note that trials/*/run.sh should invoke the source entry point so they work without a build.
- Trial evidence hygiene.
claude.json (the --json stream) and claude.stderr contain everything the child read, including any private corpus it consulted. This trial adds a .gitignore that excludes evidence/arm-*/ and tracks only metrics-*.json and drift-* files. Proposal: adopt that as the convention in the trials README, or add an ach run --evidence-dir that writes a metrics-only sidecar next to the raw stream so trials can commit the sidecar alone.
Acceptance
Summary
New trial
trials/proposal-drift-r3/(branchexperiment/proposal-drift-r3, not yet pushed) usesach run --agent claudeto measure how much an implementation spike changes an engineering proposal. Setting it up surfaced twoach rungaps and one evidence-hygiene convention worth deciding on.Trial design
Three arms write the same proposal from a work repository with a personal
proposalskill (Claude Code skill: evidence tiers,Verified/Historical report/Assumptionlabels on every claim, stable decision IDs, one markdown source rendered to an internal and an external audience). The only variable is the evidence tier:Scoring: the skill's
drift.py --compare a.md b.mdmatches decision IDs across two versions (kept / changed / dropped / new by text similarity, threshold 0.8) and reports drift = (changed + dropped + new) / (ids + new) plus a hardening proxy (reduction inAssumptioncount). Hypotheses written before running:Arm B is running now (
--budget-usd 25 --max-turns 300 --wall-ms 7200000). Metrics (wall seconds, turns, cost, tokens, label counts, decision IDs, leak-lint result, drift ledger) land inevidence/metrics-arm-b.jsonandevidence/drift-*.md; I will post them here.Gaps found
ach runexposes no permission or sandbox flag for the claude adapter.src/adapters/claude.tsalready hasclaudeSandboxArgs(SandboxPolicy)(from #8) mappingpermissionMode,allowedTools,disallowedTools,mcpConfigto CLI flags, butcmdRuninsrc/cli/ach.tsnever builds aSandboxPolicy. A child that needs Bash, file writes, MCP tools and subagents cannot run in the default ask mode under-p, so the trial passes--extra-args "--permission-mode bypassPermissions". Proposal: add--permission-mode,--allowed-tools,--disallowed-tools,--mcp-configtoach run, wired to the existingSandboxPolicy, and print the resolved policy in the run header so a trial README can cite it.achis not on PATH after a fresh clone or worktree.package.jsondeclaresbin.ach = ./dist/cli/ach.js, butdist/is a build output. The trial'srun.shfalls back tobun run src/cli/ach.ts. Proposal: README quick-start line forbun install && bun run build:node && bun link(thebinentry points atdist/cli/ach.js, whichbuild:nodewrites) and a note thattrials/*/run.shshould invoke the source entry point so they work without a build.claude.json(the--jsonstream) andclaude.stderrcontain everything the child read, including any private corpus it consulted. This trial adds a.gitignorethat excludesevidence/arm-*/and tracks onlymetrics-*.jsonanddrift-*files. Proposal: adopt that as the convention in the trials README, or add anach run --evidence-dirthat writes a metrics-only sidecar next to the raw stream so trials can commit the sidecar alone.Acceptance
ach run)