Skip to content

trials: proposal-drift experiment under ach run, plus permission-mode and install gaps #13

Description

@sblattj

Summary

New trial trials/proposal-drift-r3/ (branch experiment/proposal-drift-r3, not yet pushed) uses ach run --agent claude to measure how much an implementation spike changes an engineering proposal. Setting it up surfaced two ach run gaps and one evidence-hygiene convention worth deciding on.

Trial design

Three arms write the same proposal from a work repository with a personal proposal skill (Claude Code skill: evidence tiers, Verified / Historical report / Assumption labels on every claim, stable decision IDs, one markdown source rendered to an internal and an external audience). The only variable is the evidence tier:

Arm Tier Inputs
A A codebase graph plus design docs on the base branch
B B A plus a meeting-transcript corpus with speaker roles
C C B plus the implementation spike worktree

Scoring: the skill's drift.py --compare a.md b.md matches decision IDs across two versions (kept / changed / dropped / new by text similarity, threshold 0.8) and reports drift = (changed + dropped + new) / (ids + new) plus a hardening proxy (reduction in Assumption count). Hypotheses written before running:

  • H1: drift(B vs C) < drift(A vs C)
  • H2: hardening(B to C) > 0.3
  • H3: arm A names no decision owner for at least half of its open questions; B and C name one for all
  • H4: drift(C vs C0) < 0.15 (two tier-C runs on identical inputs give the noise floor)

Arm B is running now (--budget-usd 25 --max-turns 300 --wall-ms 7200000). Metrics (wall seconds, turns, cost, tokens, label counts, decision IDs, leak-lint result, drift ledger) land in evidence/metrics-arm-b.json and evidence/drift-*.md; I will post them here.

Gaps found

  1. ach run exposes no permission or sandbox flag for the claude adapter. src/adapters/claude.ts already has claudeSandboxArgs(SandboxPolicy) (from #8) mapping permissionMode, allowedTools, disallowedTools, mcpConfig to CLI flags, but cmdRun in src/cli/ach.ts never builds a SandboxPolicy. A child that needs Bash, file writes, MCP tools and subagents cannot run in the default ask mode under -p, so the trial passes --extra-args "--permission-mode bypassPermissions". Proposal: add --permission-mode, --allowed-tools, --disallowed-tools, --mcp-config to ach run, wired to the existing SandboxPolicy, and print the resolved policy in the run header so a trial README can cite it.
  2. ach is not on PATH after a fresh clone or worktree. package.json declares bin.ach = ./dist/cli/ach.js, but dist/ is a build output. The trial's run.sh falls back to bun run src/cli/ach.ts. Proposal: README quick-start line for bun install && bun run build:node && bun link (the bin entry points at dist/cli/ach.js, which build:node writes) and a note that trials/*/run.sh should invoke the source entry point so they work without a build.
  3. Trial evidence hygiene. claude.json (the --json stream) and claude.stderr contain everything the child read, including any private corpus it consulted. This trial adds a .gitignore that excludes evidence/arm-*/ and tracks only metrics-*.json and drift-* files. Proposal: adopt that as the convention in the trials README, or add an ach run --evidence-dir that writes a metrics-only sidecar next to the raw stream so trials can commit the sidecar alone.

Acceptance

  • arm B metrics and B vs C0 drift ledger posted here
  • decision on gap 1 (flag surface on ach run)
  • decision on gap 2 (install docs vs source entry point convention)
  • decision on gap 3 (evidence convention or metrics sidecar)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions