Paper draft: Preserve, Don't Resolve (two-corpus sovereignty) - #168
Open
mavaali wants to merge 21 commits into
Open
Paper draft: Preserve, Don't Resolve (two-corpus sovereignty)#168mavaali wants to merge 21 commits into
mavaali wants to merge 21 commits into
Conversation
…overeignty) All 10 sections from the design skeleton, grounded in the actual results, with the corrected framing: lead the sovereignty contrast with the forced/architectural condition (CB6 17/18 model-independent), present abstain-offered fabrication as the honest model-dependent softness (CB4 panel 6-26/49, contracts 4/7), and confront the 'just use GPT-4o' objection head-on (it over-abstains, 25/33 missed; daftari's guarantee is structural). Two-corpus control(contracts)/treatment(Wikipedia) structure; central claim = measured invariance of the non-fabrication+provenance axis; honest assessment + single kill condition; evidence-map appendix. Related-work section flagged for a grounding pass.
…e claim honestly Deep-research (104 agents, 21 sources, 25 claims 3-vote-verified) found the thesis is NOT novel on its components: Graphiti (bi-temporal invalidation), ATMS (unresolved contradiction), ElephantBroker (supersession-vs-contradiction edges), Roynard (supersession-preserving provenance) all predate it. §9 now cites them and narrows the contribution to the STRUCTURAL CONJUNCTION (by-construction no-mint of a tension, vs ElephantBroker's LLM-extracted + confidence-decay = model-dependent — exactly what §6 measures) + the substrate + the empirical two-corpus invariance. Added a §8 bullet self-stating non-novelty (the EB reviewer-defense). Verified findings + caveats + open questions (MemGPT/Cognee/Engram/SmartVector unverified this pass; ATMS→de Kleer 1986; Roynard zero-LLM REFUTED) in drafts/2026-06-29-related-work-research.md.
…n vs tension) Split the prior art cleanly: supersession-preservation (Graphiti recency-resolves = foil behavior, Roynard resolves contradiction->supersession; daftari here too, no novelty) vs tension-preservation (ATMS structural-but-over-assumption-sets, ElephantBroker contradiction-edge but LLM-extracted+confidence-penalized). Makes the Graphiti line crisp (it preserves history, never holds a tension open) and pre-empts the 'Graphiti already does this' reviewer by showing which 'this'. Refined gap bullet 1 so ATMS doesn't refute it: the contribution is porting TMS's structural no-collapse to the agent-memory substrate as a by-construction invariant.
…es, fold into §9 Four parallel primary-source verifications (MemGPT/Letta, Cognee, Cartridges/Engram, SmartVector); all real, citations confirmed, folded into §9 preserving the two-axes structure: - Consolidation/overwrite pole: + MemGPT/Letta [2310.08560] (core_memory_replace overwrites in place), + Cognee [2505.24478] (memify updates/prunes; DB substrate not markdown). - Supersession-preservation axis: + SmartVector [2604.20598] — preserves archived/ superseded vectors but resolves contradictions by majority vote (votes tensions away = the keystone's forbidden move). - Inverse-substrate sentence is now a real citation: Cartridges [2506.06266] /Engram (trainable KV-cache, 38.6x less memory/26.4x throughput; NOT LoRA; use paper figures not company '100x'; severe name collision -> cite paper not product URL). - Substrate gap bullet now names the concrete DB/tensor/KG contrasts. Research-draft findings table updated; open-questions resolved (only the Recall eval harness remains, not a §9 competitor). Precision caveats recorded: cite Cognee paper only for existence/KG-framing not supersession; Engram optimizes token-cost not supersession (substrate-and-aim contrast, not head-to-head).
- add §2 'what the agent does with this' illustrative trace (agent consumes a preserved tension to abstain/escalate; explicitly not an evaluation) - deliver the Zep markdown rebuttal §9 previously only promised - purge all em-dashes (colons/commas per context; en-dash ranges untouched) - fix comma splices and a dangling participle introduced by the purge - surgical clarity edits to the abstract control sentence and §5/§6 methods
…heck - Cognee [2505.24478] is a tuning/eval study of the framework, not the system paper; recast so the memify mechanism reads as documentation-sourced - Roynard [2604.11364]: soften 'contradiction triggers a supersession' to the grounded claim (no unresolved state; supersession is evidence-gated) - record the re-verification in the §9 header note (ElephantBroker, SmartVector, TOKI, PAM, Graphiti quote, Cartridges 38.6×/26.4×, 57% all body-grounded)
- references.bib: 18 entries (16 arXiv + de Kleer 1986 + Zep blog), titles and first-authors verbatim from the 2026-06-30 primary-source re-verification; SmartVector/PAM confirmed solo-author, others flagged first-author + et al. - add a human-readable ## References section to the draft (reading copy) - SmartVector and Portable Agent Memory full titles fetched from arXiv; Zep blog URL resolved (blog.getzep.com/markdown-is-not-agent-memory) - cross-checked: all 16 inline arXiv cites map 1:1 to entries, no orphans
Two limitations bullets prompted by the Zep 'Markdown is not agent memory' critique: (1) transaction-time history is preserved by auto-commit but not first-class queryable (bounded as-of reconstruction verified to compose with the per-vault process lock; valid-time unmodeled, unpulled since contracts are recency-resolvable per §4); (2) access control is collection-grained, finer isolation gated on a shared-decision use case. No new references.
Review + correction plan (2026-07-01-moderator-review-correction-plan.md, all items closed): - Fix two factual errors: relabel the CB4 panel abstain-offered in §5 (was 'forced'); replace the corpus-B value-perturbation claim with what ran (post-cutoff 14/37, 12/12 stale) and state perturbation as a limitation. - Two new experiments: Mem0 v2.0.11 real write path on the 39 corpus-B items (default add() is additive-only; correction silently unregistered 26/33) and CB6 gate negative controls (2/8 settled rejected; prompt recovered verbatim from transcript, now preserved in scripts/cb6-gate-negative-controls.mjs). - Retire the second-rater gate: author pass showed the question ill-posed (0/6 textual-conflict reading); independent rater pass validates distillation 4/6 distinct + 4/6 fair, alternative-side defects disclosed, and shows human first-position bias (6/6 asserted winners, 1/3 chance on controls). - Reframe control/treatment as contrasting regimes with a confounds paragraph; rewrite the abstract; numeric kill-condition thresholds; item-clustered 17/18; daftari-vs-GPT-4o dominance comparison in §8. - Apparatus: verbatim prompts appendix, setup/attrition disclosures (no snapshot pins, CB4 re-run), code/data availability, ethics (usernames, CC BY-SA, EDGAR), front matter; strip meta notes; bold sweep per policy. - Citations: complete author lists; fix three false et-al entries; MemGPT working_context.replace; venues of record (UIST/ICLR/NeurIPS); AIS verified Computational Linguistics 49(4):777-840, 2023. - Title: 'Preserve, Don't Resolve: Non-Fabrication and Provenance as the Evaluation Axis for Agent Memory'. - LaTeX first pass under docs/paper/latex/ (xelatex, 21pp, TikZ figures); markdown remains canonical.
Add a §9 related-work paragraph positioning data-olympus as the nearest peer on substrate (markdown+YAML+git) with a reproducible accumulation-corpus retrieval benchmark. Frames its result as a correct measurement of the accumulation half, not a competitor to the two-corpus thesis: no accumulation-only benchmark can surface the tension axis, which lives between the corpora. Numbers taken from the project's WHY.md, not independently reproduced. Because data-olympus shares the markdown-in-git substrate, soften the 'The substrate' gap bullet from 'No cited system' to 'Almost no cited system' and name data-olympus as the exception, with the distinction being the two-corpus frame rather than the store. Mirrored into the LaTeX build target; bib entry added to both references.bib files.
mavaali
force-pushed
the
paper/draft-preserve-dont-resolve
branch
from
July 3, 2026 19:58
e1f7be1 to
fdd6027
Compare
…ate passed - verify data-olympus benchmark reproduces bit-for-bit at pinned SHA ccaffdb - feud_corpus.py: 10 co-active contradiction pairs, no supersession link - feud_queries.py: fifth 'feud' stratum, gold = both sides - test_feud_disjoint.py: 5 honesty guardrails green - retrieval sanity: 10/10 feud queries return both sides via their Index - record model decision (neutral third-party via OpenRouter) + build progress
…proven
- structured answer contract {answer, evidence_state, cited_docs} (decided)
- agent.py: MockLLM (offline) + OpenRouterLLM (billed), agent loop w/ 3a tool call
- substrate.py: four cell configs (data-olympus, no-tg, tg-3a, tg-3b)
- feud_metrics.py: deterministic surface/pick/fabricate/miss classifier + rates
- run_feud.py: runner, offline mock default, --live for OpenRouter
- 18/18 tests green; mock e2e shows 3b-3a delta mechanism (+1.0 lazy, 0 diligent)
- stops before billed live run + live-daftari-MCP fidelity (both need explicit go)
… confirmed (scoped) - divergent-regime corpus: side B authored in divergent vocab, buried by retrieval over the 250-concept base bed; tension link is the only path to it - neutral-query construction isolates retrieval (removes A-perspective confound) - faithful tension join keyed on retrieved ids, not gold - fix YAML colon bug that path-derived feud doc ids - live reads recorded: shared (confounded), divergent A-biased (small), divergent neutral gpt-5.4-mini (buried topics: no-tg 0/5 surface vs tg 3-4/5) - conclusion: tension-graph win is real but scoped to the recall-limited regime; 3a>=3b (evidence against rushing tensions-in-search); model capability matters
…nner - 15 more real engineering disputes (serialization, config lang, css, render, typing, error-signaling, DI, versioning, pagination, cache, delivery, secrets, session, schema-timing, test-doubles); all pass disjointness + query-leak - burial over 250-bed: 15 buried / 10 co-retrieved (non-tuned) - run_panel.py: model panel x reps, streams trials.jsonl, summary with 95% CIs + two-proportion test on buried-topic surfacing + per-model robustness
- 3 neutral models x 3 reps over 25 topics (15 buried); stand-in substrate
- BURIED topics: baseline surfaces 0.022, tg-3a 0.185 (p=1.1e-5), tg-3b 0.444
(p=2.2e-16); co-retrieved: all cells ~0.7-0.86 (substrate adds little)
- robust across gpt-5.4-mini/gemini-2.5-flash/gpt-5-mini
- inline (3b) > dedicated tool (3a) across all models -> flips the earlier lean
on tensions-in-vault_search toward building it
- artifact: results/{trials.jsonl,summary.md}
- phase2_export.py: emit 250-base bed + feud specs for the Node harness - phase2_build_contexts.ts: throwaway daftari vault, real reindexVault + hybridSearch + addTension/listTensions -> contexts.jsonl (retrieval signal in body since daftari indexes body+title, not frontmatter triggers) - run_panel.py: --contexts replay mode; agent/classifier/stats unchanged - daftari hybrid buries divergent side 15/25 (same as data-olympus FTS) - embeddings did NOT rescue the buried side; recall-limited regime is real on daftari's actual retrieval - smoke (gpt-5-mini): buried no-tg 0.067 -> tg-3b 0.400 (p=0.03), replicates
…ftari - 675 trials, real daftari hybridSearch + tensions, 3 models x 3 reps - BURIED: no-tg 0.081, tg-3a 0.422 (p=1.1e-10), tg-3b 0.459 (p=2.8e-12) - co-retrieved: no-tg 0.933, tg 0.97-1.00 (substrate adds little) - effect at least as strong as Phase 1 stand-in; burial 15/25 matched exactly - fidelity gate passed: not a stand-in artifact; stand-in caveat retired - inline (3b) robust across models; tool (3a) model-dependent
…de availability Convert the 'one corpus too narrow' assertion into a measurement: feud augmentation on data-olympus's own corpus, buried-side surfacing baseline 0.08 vs tension-graph 0.42-0.46 (p<1e-10, 3 neutral models), replicated on live daftari retrieval. Scope stated honestly (recall-limited regime; co-retrieved null). Reference the benchmarks/tension-graph harness in code availability.
Mirror the §9 'augmentation, measured' paragraph and the code-availability reference into the .tex build target; compiles clean under tectonic.
5 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
First full draft of the two-corpus sovereignty paper, plus the verified related-work research that grounds §9.
What's here
docs/paper/preserve-dont-resolve.md— full draft, all 10 sections, grounded in the actual experiment results.docs/superpowers/drafts/2026-06-29-related-work-research.md— the verified-findings table, caveats (ATMS → de Kleer 1986; Roynard "zero-LLM" refuted), and open questions.Honest status
Docs-only; no code changes. Test suite unaffected (101 green on main).