An independent Construct-2 ancillary investigation, commissioned 28 September 2026. Status: phases 1–7 and their bounded diagnoses published locally. Root scientific review remains separate.
Phase 7 findings add shared scratch-test access and two post-freeze edits in python-slugify. All final patches pass executable tests, but only empirical reuse fulfills both requested documentation updates; its first verification loop nevertheless exhausts the budget. Learned energy selects compact then broad evidence, costing 59,738 logical tokens versus ordinary dependency access's 55,466 at equal measured requirement coverage. Own checks sometimes contain wrong expectations or miss the known bug. A separate public-scope review repairs ordinary's documentation omission in one call, without changing frozen scores. All twelve attempt grades, 22 scratch outcomes, 32 frozen hashes and 26 repository tests verify. Reproduction, final costs and next step distinguish test completion, full requested deliverables, explicit finish and reviewed repair.
Phase 6 findings report source-native python-dotenv edits with executable tests and a fixed local reader. A six-weight set energy learns earlier executed preferences, but on two later edits selects exactly the same evidence as ordinary empirical policy reuse. All four policies complete the API extension and fail the line-ending change despite passing its public example. Ordinary dependency access misses one held-out case; compact reuse/energy miss three. An isolated investigator repair explains the ordinary failure; it is not counted as reader success. The two earlier obligations survive the one accepted later revision. All 24 saved outcomes replay, three fits reproduce bit-for-bit, and 22 repository boundary tests pass. Reproduction, costs and limitations/next step include authored-task provenance, the chronology correction, failed attempts, and the common reader-interface limit. No energy-specific advantage is established.
Phase 5 findings add an authored versioned-configuration workload with native CPython execution and fresh longer dependency chains. Earlier executed feedback teaches compact selection: trained exact energy and ordinary closure both complete 48/48 noncached later tasks (72/72 including reuse). Ordinary closure remains slightly more compact and cheaper. The same energy's frozen gradient recipe completes only 34/48; a separate soft-penalty energy completes 29/48. Public repair closes all failures. Post-hoc extra steps/restarts repair the optimizer, while explicit hard constraints reproduce ordinary closure. These separate representation, objective and optimization limits—not an LLM or neural advantage. Reproduction, native costs and audit preserve five fits, failed attempts and all outcomes. All 18 repository tests pass; the seven new tests cover this phase.
Phase 2 findings extend the work to real MuSiQue prose, a small neural pairwise energy, and projected continuous selection. The learner improves development support recovery, while exact/beam search is more reliable and cheaper than the tested relaxation. The semantic diagnostic and frozen held-out selection and 24-case QA comparison are complete. Exact selection reaches 2/24 annotation-complete answers (3 with fallback), versus full-context access's 6/24. A source/label audit found mismatched-entity chains and exact-match penalties, so these are not validated counts of all genuinely correct tasks. Separate length-stop repair raises full-context completion to 7/24 but leaves learned compact completion at 2/24. This phase is explicitly benchmark-supervised, not continuing experience. The earlier development publication is retained unchanged. All eleven boundary tests and the final integrity audit pass. The native cost ledger includes 262 actual reader calls, representation construction, acquisition, search and repair. See reproduction and publication hashes.
Phase 3 scaling findings test the frozen energy with up to 160 candidates. Relaxation beats exhaustive enumeration at larger pools, but beam search matches the optimum with less measured work. These are optimizer diagnostics on seen material. Phase 2's frozen held-out selection result is 26/64 complete annotated evidence sets versus 15/64 for ordinary hybrid access, without an established useful complete-work advantage. The exposure audit found one previously inspected reconnaissance question in those 64: the 63-case sensitivity result remains 26 versus 15. Split-disjoint component IDs do not imply fully disjoint source text or absence of pretrained exposure.
Phase 4 findings show that continuous inference is sensitive to terms that vanish on every discrete evidence set: one such extension changes 41/48 rounded selections despite identical discrete energies. Discrete repair reduces the sensitivity. This is an inference-method diagnostic, not new learning or an additional fresh-task sample.
The phase 1 findings report an eight-coefficient set learner trained on earlier executed attempts in an explicitly authored quote/revision workload. After diagnosis and a claim freeze, ordinary dependency closure and exact learned selection each complete 30/30 fresh executable tasks. One-swap optimization of the same negative score completes 24/30 before ordinary repair. The fixed-reader diagnostic completes only 2/5 compact-context cases without tool repair. This is a functioning small learner and an explained ordinary-method advantage, not an EBM advantage or a natural-workload memory result.
Evidence: protocol, frozen claim,
confirmation outcomes,
reader outcomes, audit,
actual provenance, and
reproduction instructions.
Failed acquisition attempts and all raw reader responses are retained in runs/.
Phase 1's proposed source-native extension is now explored in phases 2–4. The
broader question remains open, especially validated complete behavior under
source revisions and learner-earned continuing feedback.
Can a small specialist learn from earlier experience to assemble compact, sufficient evidence for later complete work, and remain useful as that evidence changes, beyond competent ordinary retrieval and set selection?
The user selected this area after discussing energy-based models as memory and recognition components. A learned energy function is one treatment to develop. A neural advantage is not required. General-purpose EBM reasoning, replacing the primary LLM, and reproducing proprietary Kona are outside this study.
Start with AGENTS.md and START.md. The root selection preserves ES1–ES3, the scientific scope and its relation to Construct's enduring question. The source comparison records root's exact-version methods reading and inspection limits. Root's existing experience-guided investigation is independent and keeps its commission.
A short prompt can still omit the one fact needed to use another fact correctly. The selected opportunity concerns complementary, redundant, scoped or revised evidence in a recoverable archive. Competent ordinary retrieval and source reuse must remain available. Past Construct comparisons found learned delivery and score prediction without beating cheap explicit access; those results motivate a consequential workload, not weakening the alternatives.
There is close prior art. Lightweight set scoring and state-conditioned evidence assembly already have public methods. This study therefore concerns acquisition, transfer, maintenance and useful cost of selection experience, rather than a claim to invent joint retrieval. See the root ledger's P105–P109. In particular, include a serious set-aware policy; a win against pointwise ranking alone would leave a major alternative unexplained.
Own workload discovery and develop a small comparison with a fixed primary LLM, shared explicit source history and source-linked context delivery. Find real interactions among evidence and an independently scorable complete task. Inspect the public assets in FEASIBILITY.md before adopting them. Static retrieval/QA can support acquisition development; establishing experience across sessions requires a separately justified sequence, eligible earlier feedback and fresh later tasks. Do not label shuffled benchmark rows an acquired life history.
Specify which state accumulates: archive, embeddings, selection-policy parameters or several of these. Benchmark labels and externally authored demonstrations may help develop a learner, but are not feedback acquired through its own attempts. Identify the additional value of earlier selection experience with a suitable frozen-state or no-new-experience control. Existing artifacts and source answers are legitimate ordinary memory, not something to remove from the comparator.
The energy-based candidate should have a concrete operational definition. Since
energy = -score preserves every ranking, EBM and conventional set scoring are
not automatically disjoint method families. Separate interaction learning from
search or inference-time optimization when that distinction affects the claim.
You may compare different search methods on the same learned energy, or a
targeted interaction ablation. Choose the smallest informative contrast; root
has not prescribed a mandatory model, optimizer, dataset or factorial suite.
Measure complete primary-model outcomes alongside sufficient-evidence recovery, stale/inapplicable reuse and delivered context. Include acquisition, source encoding/indexing, labels, selection, optimization, fallback and repair costs. Preserve failures and diagnose unsuccessful acquisition purposefully. Freeze a developed claim before fresh confirmation, and report whether a simpler solution explains closure.
- ES1: interaction learning has the most opportunity when necessary complements remain missing after competent ordinary set selection; direct addressing or cheap scope rules may remove the advantage.
- ES2: useful access behavior may survive changed authoritative values while evidence relationships remain applicable, but can fail when those relationships or the quality of prior feedback change.
- ES3: iterative energy minimization must justify its work against ordinary small predictors and competent discrete set search; it may add no useful value.
The full original wording lives in the root selection note. Append assessments without rewriting those predictions. These are expectations, not observations.
Use the model resources and investigator policy in AGENTS.md. Start with the smallest useful local acquisition/behavioral comparison and verify mutable-state access. The first publication should establish a workable task and learner, explain a limitation through bounded diagnosis, or identify an actual resource constraint. Mere infrastructure setup is not the scientific endpoint. The study owns methods, execution and reporting; routine decisions need no new root approval.
Shared repository: https://github.com/alignment-farm/evidence-set-memory (private). Preparation provenance and asset boundaries are in PROVENANCE.md and FEASIBILITY.md.