Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

What evidence is needed to accept a lesson?

Prepared 19 September 2026 as an independent Construct-2 ancillary study. Status: bounded experimental phase complete; results and reproduction published. Read AGENTS.md and sources/README.md, then proceed through bounded workload discovery, development and fresh experiments. You own methods, protocols, implementation, resource sizing, diagnosis and publication.

Current experimental evidence

In two constructed reporting histories, source SQL passed its original slice but double-counted under independent child-table multiplicity. Across six later cases per arm, cheap schema review and paid probing both achieved 6/6 initial and final correct outputs; raw experience achieved 2/6 initially and 6/6 after ordinary repair. Independent behavioral checks passed 5/6 raw final programs and all cheap/ paid programs. Paid probing added no measured benefit over competent cheap review. Direct reuse of the cheaply repaired SQL also passed 6/6, requiring no later model calls. The phase closes on this schema-solvable workload limit, with all weak-reviewer and serving failures preserved. Original LA1–LA3 below remain unchanged; their assessments and native-unit costs are in the report.

See the study report, initial protocol, fresh protocol amendment, serving diagnosis, and public-method ledger. Offline source/probe verification: uv run scripts/verify.py. Results/cost reproduction: uv run scripts/pilot.py summarize and uv run scripts/costs.py. The full report gives live replication commands, the measured three-reuse-per-history horizon, limitations, and the assessment of a bounded next step. Full source episodes, proposals, review prompts, responses and execution costs are preserved in evidence/pilot-v1/.

Question

When does acquiring additional evidence about a proposed reusable lesson improve future complete work enough to justify its cost, beyond competent cheap checks, reflection and reuse of retained source experience?

An episode's outcome, its explanation, a lesson's validity and that lesson's future usefulness are separate claims. A successful task can support an overly broad rule; a failed task can contain a useful procedure. Acceptance here means choosing a candidate for future guidance within a declared scope. It is not a declaration of universal truth or authority over a later user correction.

The broader program asks how agents accumulate useful experience and where it should live. This study investigates the evidence for accepting experience before asking whether weights or another representation should carry it. Explicit notes, source examples and acquired code are all legitimate memory. A learned selector, skill format, general verification framework or adapter advantage is not required.

Starting workload and development

Start with recurring data transformation or database reporting, where executable consequences and complete outputs can be checked. Source work might produce lessons about join cardinality, identifier normalization, missing values or conditional aggregation. For example, joining on a display name may work in the first observed slice while failing when names repeat. This is an illustrative possibility, not an observed natural error or a prescribed benchmark trap.

Use grounded public artifacts when they offer a clear comparison. Motivated constructed fixtures are also permitted; disclose their origin and limits. Keep public specifications, schema constraints, competent parsing/arithmetic, ordinary tests and acquired programs available. If those settle acceptance cheaply, that is an explanatory result. Do not hide an obvious key, corrupt otherwise valid records, or weaken the reader to manufacture a benefit for verification.

Collect actual task episodes and have a declared proposer extract candidate lessons from eligible observations. Keep successes, failures, proposals and their provenance. Authored or teacher corrections are allowed and charged; distinguish them from model-acquired lessons. Deliberately injected bad lessons may diagnose a mechanism but cannot establish their natural prevalence. Freeze proposal generation and preserve the original wording before comparing acceptance policies.

First establish useful source reuse, a competent executor given appropriate evidence, candidate lessons whose acceptance matters, and a functioning affordable check. Diagnose poor task/tool behavior before attributing a null to memory. Select or revise the workload on development material. If this lead lacks an unresolved useful decision, investigate a better lead within the same question; do not repeatedly expand a solved finite catalog.

Starting comparison

Fork the same source history and acquired artifacts. Keep the primary, tools, context budget and competent ordinary delivery policy fixed initially.

Condition Information used to decide what becomes reusable guidance
Raw-experience reuse Retain source episodes and acquired code for competent later access/reconstruction, without pre-accepting abstract lessons.
Inexpensive grounded acceptance Evaluate shared candidates using source traces, available requirements, provenance, ordinary checks and developed reflection.
Acceptance with paid probing The same inputs and candidates, plus the option to buy a targeted environment observation or isolated execution check.

All conditions retain the original source experience and can reuse successful programs. A rejected lesson does not erase its episode. The raw baseline need not regenerate or discard working code. The two acceptance policies see identical initial proposals; their additional acquired observations can differ and must be counted. Record the proposer, acceptance decision, evidence available at decision time, check performed, cost and eventual use. Accept, decline or defer are valid outcomes; declining everything still has to compete on later complete work.

Distinguish replaying the original episode from challenging an unobserved condition or using an independent source case. A check may reveal that a rule needs a narrower scope. Preserve that as a new version and distinguish rewriting from acceptance of the original candidate. A small fixed-candidate comparison can identify the acceptance effect; a later whole-curation comparison can test the benefit of rewriting. You choose the useful sequence and exact protocol.

If claiming that selective checking earns its cost, compare with a developed fixed checking policy under a comparable allowance; equal-budget random allocation or check-all can be diagnostics when they distinguish a material explanation. Avoid a large mandatory matrix. Extra thought tokens, additional task practice and genuinely new evidence are different treatments and costs.

Evidence and evaluation

Checks may inspect the task-visible environment, query public source data or execute a procedure on owned copies. They must not consult final evaluation answers, future tasks or an investigator-only truth table disguised as a tool. Independent grading is allowed for evaluation and declared source feedback; make any use as teaching explicit. Probe side effects must not silently improve the later task environment, and reused probe artifacts count as acquired evidence.

Evaluate later complete tasks with fresh material after acceptance policies and budgets are fixed. Include intended computations or operations, cases outside the proposed scope and still-valid obligations. Plan more than one independent source history where feasible; many paraphrases of one history do not establish history transfer. Source, policy development, probing and final evaluation need distinct boundaries. If final outcomes guide diagnosis, use new material for a subsequent confirmation and preserve the earlier result.

Measure complete outcomes and their paired changes, harmful generalizations, useful lessons rejected, actual exposure/use and all known costs. Verifier agreement or acceptance precision alone is insufficient. Include source collection, proposal generation, cheap checks, paid checks, failed probes, teacher work, selection, later inference/tools and repair. Separate deployment work from experimental search and avoid counting shared costs twice. Report tokens, calls, wall time and money where measured in their native units; unknown engineering costs remain unknown. Report cumulative quality and costs over an explicit reuse horizon before claiming repayment.

Start with a stable ordinary reader. The reviewed runtime's learned reader is sensitive to catalog representation; its frozen 1.5B worker is not a competent default comparator for this question. Learning an acceptance or consolidation policy is within the remit if evidence makes it useful. Establish acquisition, comparable source opportunity and fresh-history behavior before interpreting a learning null or claiming a neural advantage. No additional routine root approval is needed for that justified bounded development.

Prospective expectations — LA1–LA3

Preserve this wording and append later assessments, including adverse outcomes.

  • LA1: When plausible candidate lessons exceed what their source episodes support, additional environment evidence can improve later complete work over inexpensive grounded checks and reflection. No incremental benefit when cheap checks already settle acceptance would narrow the useful role of probing.
  • LA2: A successful replay of the source episode need not establish a lesson's scope. Checks that challenge an unobserved condition should better predict useful transfer when that condition matters; no advantage over source replay would limit this expectation in the selected workload.
  • LA3: Better acceptance accuracy need not repay checking cost. Selective probing should be most useful when the lesson recurs, its errors are costly, and inexpensive checks leave a consequential uncertainty. Raw reuse or a developed fixed checking rule may remain preferable over the measured horizon.

Runtime and publication

Use Construct Runtime as a pinned instrument where helpful. The reviewed pin is 09f66837ca76db1f674ee3372d734219ad0f3e27. Preparation supplies an ignored, owned checkout in .deps/construct-runtime; no service or model is started. At preparation the pin was local-only, so do not assume a GitHub clone has it. From this lab checkout, recreate the dependency with:

mkdir -p .deps
git clone --no-hardlinks ../../derivatives/construct-runtime .deps/construct-runtime
git -C .deps/construct-runtime checkout --detach 09f66837ca76db1f674ee3372d734219ad0f3e27

Inspect its instructions, docs/INVESTIGATOR.md, docs/MEMORY_WORKER.md and reports/MILESTONE-3.md. It provides external environments, state forks, source/event records and separate model roles. The built-in scope retriever is account-specific, and authority/checked fields are trusted declarations. Keep the study's proposals, conditional guidance, checker and acceptance records in an investigator-owned controller. Do not relabel an empirical hypothesis as an external authority merely to pass the runtime's delivery guard. Use a disclosed context path or owned adaptation when needed, keeping it identical across arms.

You may adapt this owned dependency or use another instrument when justified; record the patch and exact source identity. Leave the derivative's active checkout, root and sibling evidence unchanged. Do not wait for another runtime build.

Publish an abstract here and linked methods, evidence, failed attempts, costs, reproduction commands and an assessment of LA1–LA3. Assess the most useful bounded next step before standing down. Close a phase on explanatory progress, a demonstrated limitation or a concrete resource constraint. A first failed recipe or publication alone is not closure; a verification or neural win is not required.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages