Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
34 changes: 34 additions & 0 deletions data/validation/pi_fprior_killtests_2026-07-07.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,34 @@
{
"date": "2026-07-07",
"context": "Two pre-registered cheap kill-tests for the DE-44 'survivors' (calibrated-UQ PI-recalibration; invivo-F-prior). Both foreclosed. Local macOS public-clone stack; deltas are stack-internal (baseline reproduces the ~2.74 band).",
"pi_adaptive_conformal_killtest": {
"question": "Does any per-drug difficulty signal predict |log10 meta fold-error|, so a normalized/adaptive conformal PI could be tighter than the flat /÷12.9 while holding coverage under Invariant #5 (calibrate on train, not holdout)?",
"current_pi": {"method": "split_conformal", "q90_meta_log10": 1.111, "fold": 12.9, "holdout_coverage_at_nominal_0.90": 0.9533, "calibration": "train n=67"},
"spearman_rho_error_vs_signal": {
"div_eng_ml": -0.035, "div_eng_meta": -0.002, "div_ml_meta": -0.057,
"n_ad_flags": -0.051, "abs_log_meta": -0.069, "meta_cmax": 0.056, "not_in_ad": -0.061,
"mc_parameter_uncertainty_halfwidth": 0.039
},
"mondrian_by_compound_type": {
"verdict": "DEAD — widens +86%",
"reason": "n=67 split into acid/base/neutral/zwitterion (14/23/27/3); finite-sample conformal quantile ceil((n+1)(1-a)) lands at ~class-max at small n, so per-class q exceeds the pooled 90th pct. Train-calibrated per-class q (base 1.356, neutral 1.545) > global 1.100.",
"train_q90_by_class": {"acid": 1.112, "base": 1.356, "neutral": 1.545, "zwitterion": "INF (n=3<9)"}
},
"verdict": "WALLED — max |rho| = 0.069 across all obs-free signals incl. MC parameter-uncertainty. The structural model-form error (~60pp of the spread, not in the MC) is per-drug non-discriminable (confirms DE-41). The flat /÷12.9 PI is near-optimal under Invariant #5; its 0.953 over-coverage is the safe direction and cannot be honestly tightened (would require holdout-tuning or a non-existent adaptive signal)."
},
"invivo_f_prior_killtest": {
"question": "Does an in-vivo F prior (xgboost_bioavailability.json, scaffold-CV R2=-0.09) applied via the measured-F routing at shrinkage w improve the holdout meta AAFE? (DE-44 pre-registered: rises +0.02-0.08 => kill.)",
"baseline_meta_aafe": 2.7428,
"meta_aafe_by_w": {"0.0": 2.7428, "0.25": 2.6852, "0.5": 2.6459, "1.0": 2.7360},
"delta_vs_baseline": {"0.25": -0.0576, "0.5": -0.0969, "1.0": -0.0068},
"surprise": "AAFE FELL (not rose) at w=0.25/0.5 — the pre-registered surprise branch. Investigated with placebo controls.",
"placebo_controls_at_w0.5": {
"real_fpred_delta": -0.0969,
"constant_geomean_delta": -0.1328,
"shuffled_fpred_delta": -0.1160,
"interpretation": "CONST and SHUFFLED reproduce AND EXCEED the real per-drug F improvement => zero per-drug F information (real is slightly WORSE than shuffled). The improvement is a pure flat-scalar median-bias-null."
},
"mechanism": "f_eng geomean 0.148 vs Fpred geomean 0.424; median(Fpred/f_eng)=2.57 (upscale). The engine systematically under-predicts F, so any upward F scalar nulls that median bias -> AAFE dips; a constant does it cleanest. This is DE-42 (flat scalar nulls median F, not dispersion) shown directly on the meta.",
"verdict": "KILLED (triple-dead): (1) the effect is the DE-42 flat-scalar artifact, placebo-proven, zero per-drug F info (DE-28 re-confirmed); (2) the scalar is a holdout-fit median-bias correction = Invariant #8 forbidden ('fudge to Cmax loss, any form'); (3) even the best (constant) delta -0.133 is within the bootstrap CI half-width ~0.42 and does not generalize (prospective F-under-call differs; DE-42 over-tail flip)."
}
}
20 changes: 20 additions & 0 deletions docs/research/dead-ends.md
Original file line number Diff line number Diff line change
Expand Up @@ -474,6 +474,26 @@ Artifacts: `scripts/probe_liver_zonation.py`, `tests/integration/test_liver_zona

---

### DE-55 — Adaptive / normalized conformal PI: no per-drug difficulty signal exists, so the flat /÷12.9 interval cannot be honestly tightened (2026-07-07)

**Date:** 2026-07-07

**The DE-44 "PI-recalibration" survivor, tested and foreclosed.** The user-facing 90% Cmax PI is a train-calibrated split-conformal interval (meta q90 = 1.111 → **/÷12.9**, holdout coverage **0.9533** at nominal 0.90). The over-coverage cannot be "fixed" by lowering q — that is tuning on the holdout (Invariant #5). The only honest tightening is **adaptive/normalized conformal**: give easy drugs a narrower interval via a per-drug difficulty signal σ(x). The pre-registered kill-test (Spearman ρ between |log10 meta fold-error| and σ) is **dead across every signal**: div(eng,ml)=−0.035, div(eng,meta)=−0.002, div(ml,meta)=−0.057, n_ad_flags=−0.051, |log meta|=−0.069, meta_cmax=+0.056, not_in_ad=−0.061, and — the specifically-named signal — **MC parameter-uncertainty half-width ρ=+0.039** (max |ρ|=0.069). Mechanistically airtight: the MC propagates *parameter* uncertainty only (~30% coverage at nominal 90%); the dominant ~60pp of the spread is **structural model-form error not in the MC**, and it is per-drug **non-discriminable** (confirms DE-41). **Mondrian-by-compound-type also DEAD** — it *widens* the mean interval **+86%**: splitting n=67 into acid/base/neutral/zwitterion (14/23/27/3) makes the finite-sample conformal quantile `ceil((n+1)(1−α))` land at ~the class maximum, so per-class train q (base 1.356, neutral 1.545) exceeds the pooled 1.100. **Verdict:** the flat /÷12.9 PI is near-optimal under Invariant #5; its 0.953 over-coverage is the *safe* direction. The PI is not broken (split-conformal already fixed the 30%-coverage MC PI) — it just cannot be adaptively narrowed because the error is unpredictable per-drug. Artifact `data/validation/pi_fprior_killtests_2026-07-07.json`.

**Telltale if it returns:** "a normalized / CQR / Mondrian conformal will tighten the 13-fold PI." It will not — no per-drug difficulty signal predicts the error (ρ≈0 for all, incl. MC-CV), and small-n Mondrian widens it. Ship PI work only if a *new* orthogonal difficulty signal clears |ρ|≥0.3 first.

---

### DE-56 — invivo-F-prior: the pre-registered "surprise" (AAFE falls) is the DE-42 flat-scalar median-bias artifact, placebo-proven (2026-07-07)

**Date:** 2026-07-07

**The DE-44 `invivo-F-prior` test, run with placebo controls — killed, richer than predicted.** Applied the structure→F predictor (`xgboost_bioavailability.json`, target log10(F/100), **scaffold-CV R²=−0.09** — noise, DE-28) to the holdout via the measured-F routing at shrinkage w∈{0.25,0.5,1.0}. DE-44 pre-registered "AAFE **rises** +0.02–0.08 → kill." Instead the meta AAFE **fell**: 2.7428 → 2.685 (w=0.25) → **2.646** (w=0.5) → 2.736 (w=1.0) — non-monotonic, the surprise branch. **Placebo controls at w=0.5 resolve it decisively:** REAL per-drug F Δ=**−0.097**, but a **CONSTANT** F (geomean) Δ=**−0.133** and a **SHUFFLED** F (drug↔F match destroyed) Δ=**−0.116** both *reproduce and exceed* it. Per-drug F information is **zero** (real is *worse* than shuffled). **Mechanism:** f_eng geomean 0.148 vs Fpred 0.424, median k=**2.57** (upscale) — the engine systematically under-predicts F, so *any* upward F scalar nulls that median bias → AAFE dips; a constant does it cleanest. This is **DE-42** (flat scalar nulls median F, not dispersion) shown directly on the meta headline. **Triple-dead:** (1) the effect is the placebo-proven flat-scalar artifact, no F signal (re-confirms DE-28/42); (2) the scalar is a holdout-fit median-bias correction = **Invariant #8 forbidden** ("fudge to Cmax loss, any form"); (3) even the best (constant) −0.133 is within the bootstrap CI half-width ~0.42 and does not generalize (the prospective F-under-call differs, DE-54; DE-42's over-tail flip). Artifact `data/validation/pi_fprior_killtests_2026-07-07.json`.

**Telltale if it returns:** "an in-vivo F prior improved the holdout AAFE by ~0.10." That drop is a **flat-scalar median-bias null**, not F information — a constant or shuffled F reproduces it, and the scalar is holdout-fit (Invariant #8) and non-generalizing. Any future F-prior must beat its own constant/shuffled placebo before it counts.

---

## 3. When to consult this list

- Before writing a design spec for any accuracy improvement.
Expand Down
24 changes: 24 additions & 0 deletions docs/research/experiment-log.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,30 @@ Reverse-chronological. The project README carries only the **current** headline

---

## 2026-07-07 — Two DE-44 kill-tests run: adaptive PI walled (DE-55), invivo-F-prior killed via placebo (DE-56)

Ran the two cheap pre-registered kill-tests DE-44 flagged as "survivors to test," on the 107-holdout
(local macOS public-clone; deltas are stack-internal). Both foreclosed.

- **Adaptive/normalized conformal PI (DE-55):** the user-facing 90% Cmax PI is train-calibrated
split-conformal, meta q90=1.111 (/÷12.9), holdout coverage 0.9533. The over-coverage can't be lowered
without tuning on the holdout (Inv #5); the only honest tightening is adaptive-by-difficulty. Spearman
ρ(|log meta fold-error|, σ) is dead for every obs-free signal — divergence(×3), AD flags, magnitude,
and the specifically-named **MC parameter-uncertainty half-width ρ=0.039** (max |ρ|=0.069). Mondrian
by compound-type *widens* +86% (n=67 split → per-class conformal q lands at ~class-max). The flat
/÷12.9 PI is near-optimal under Inv #5; the structural error is per-drug non-discriminable (DE-41).
- **invivo-F-prior (DE-56):** applied the structure→F predictor (R²=−0.09 noise) via measured-F routing
at shrinkage w. DE-44 pre-registered "AAFE rises → kill"; instead it **fell** (2.743 → 2.646 at w=0.5,
non-monotonic) — the surprise branch. **Placebo controls decided it:** a CONSTANT F (Δ−0.133) and a
SHUFFLED F (Δ−0.116) both reproduce/exceed the real per-drug F (Δ−0.097) → zero per-drug F info. It is
the DE-42 flat-scalar median-bias null (f_eng 0.148 vs Fpred 0.424, median k 2.57 upscale nulls the
engine's systematic F-under-call). Triple-dead: placebo-artifact + Invariant #8 (holdout-fit scalar =
fudge to Cmax loss) + within-CI + non-generalizing.

The pre-registration + placebo discipline caught what would otherwise read as a "broke the ceiling"
false positive (a −0.10 AAFE drop). No headline change (2.743). Artifact
`data/validation/pi_fprior_killtests_2026-07-07.json`; DE-55, DE-56.

## 2026-07-06 — OATP1B1 ECM: closed the FLUX-1 auto-ECM xfail; +BSA PSu,inf formally deferred (data-blocked)

A 3-agent investigation of "the one un-foreclosed accuracy lever" (diagnosis §9 — the +BSA /
Expand Down