Finding
Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement (arXiv:2609.01481, submitted 2026-09-01) reports continual planning, coding, and testing loops over existing coding harnesses, with small verifiable increments, independent evaluation, progressive capability exposure, reuse, and versioned histories. Across GameCraft-Bench, FrontierSWE, and ProgramBench with three harness/model pairs, the originating team reports an average relative gain of 52.25% and a maximum gain of 82.86% after three iterations.
Evidence status: originating-team measured with linked GitHub project; not independently reproduced by RuV.
Why this is not a new architecture by default
Dream Machine, MetaHarness, witness, Core Memory, RVM, and the trace-driven optimizer already implement much of the same shape: bounded candidates, independent gates, histories, evidence, rollback, and authority separation. The highest-information task is therefore a differential benchmark, not another framework.
Differential benchmark
Freeze:
- model and exact version
- coding harness and version
- task set and seeds
- total token/compute/time budget
- independent evaluator
- safety and authority policy
Compare:
- current Dream Machine plus MetaHarness loop
- HoH-inspired loop with progressive artifact disclosure and explicit repair/capability-growth scheduling
Report absolute task success, iteration-to-success, regression rate, token cost, wall time, evaluator cost, artifact reuse, duplicated work, security findings, and historical retention.
Promotion gate
Only import a new primitive if it improves held-out task success by at least 5 absolute points, or preserves success while lowering total compute/cost by at least 20%. If current RuV performs within variance, record that as a negative result and reuse the existing substrate.
Governance
No self-merge, no evaluator mutation, no authority expansion. Statistical claims are subject to the persistent validity gate and independent MetaHarness reproduction.
Finding
Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement(arXiv:2609.01481, submitted 2026-09-01) reports continual planning, coding, and testing loops over existing coding harnesses, with small verifiable increments, independent evaluation, progressive capability exposure, reuse, and versioned histories. Across GameCraft-Bench, FrontierSWE, and ProgramBench with three harness/model pairs, the originating team reports an average relative gain of 52.25% and a maximum gain of 82.86% after three iterations.Evidence status: originating-team measured with linked GitHub project; not independently reproduced by RuV.
Why this is not a new architecture by default
Dream Machine, MetaHarness, witness, Core Memory, RVM, and the trace-driven optimizer already implement much of the same shape: bounded candidates, independent gates, histories, evidence, rollback, and authority separation. The highest-information task is therefore a differential benchmark, not another framework.
Differential benchmark
Freeze:
Compare:
Report absolute task success, iteration-to-success, regression rate, token cost, wall time, evaluator cost, artifact reuse, duplicated work, security findings, and historical retention.
Promotion gate
Only import a new primitive if it improves held-out task success by at least 5 absolute points, or preserves success while lowering total compute/cost by at least 20%. If current RuV performs within variance, record that as a negative result and reuse the existing substrate.
Governance
No self-merge, no evaluator mutation, no authority expansion. Statistical claims are subject to the persistent validity gate and independent MetaHarness reproduction.