Causal validation and falsification protocols for factor investing
Most factor research can't tell signal from self-deception. causal-quant is the validation layer that can. Declare a causal graph, run a falsification battery, and quantify how much of a reported edge could survive search and selection — implementing the published de Prado (2023/2026) and Bailey (2014/2017) protocols, not hand-waving. It pins down the three ways a backtest lies — luck, confounding, and selection across everything you tried — and returns HAC-correct, skeptical verdicts instead of flattering p-values. Built to falsify your own strategies before you risk capital.
A Python library that codifies the methodological framework from:
López de Prado, M. (2023). Causal Factor Investing: Can Factor Investing Become Scientific? Cambridge University Press (Elements in Quantitative Finance).
and the search-adjusted false-discovery-rate machinery from:
López de Prado, M., and F. Fabozzi (2026). The False Discovery Rate in Finance: Identification Failure and Search-Adjusted Estimation. ADIA Lab Research Paper Series, No. 24.
For the single-strategy selection problem (where the field-level FDR does not apply), it also implements the deflated-Sharpe and backtest-overfitting tools from:
Bailey, D. H., and M. López de Prado (2014). The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting, and Non-Normality. Journal of Portfolio Management 40 (5).
Bailey, D. H., J. Borwein, M. López de Prado, and Q. J. Zhu (2017). The Probability of Backtest Overfitting. Journal of Computational Finance 20 (4).
The library helps researchers and practitioners move factor strategies out of the "phenomenological" stage (associational claims, p-hacking, specification errors) and toward falsifiable, causally grounded research.
It is not another backtesting or factor-zoo library — it is the validation layer that sits after you have a candidate edge, to test whether that edge is real, identified, and robust to search.
The fastest way to put this library to work on your strategy is to let Claude drive it:
- Clone the repo and run
claudeinside it. - Say "evaluate my strategy" — then either describe it (the claim, a return series + a benchmark, and honestly how many specs you tried) or point Claude at a local returns file.
Claude auto-loads CLAUDE.md and follows RUNBOOK.md — the guided
A → B → C → evidence-ceiling protocol — running the real functions (evaluate_strategy,
jensens_alpha, CausalGraph, the FDR suite, run_falsification_battery) and returning an
honest verdict: is there an edge (HAC, not OLS), is it identified, did you data-mine, and what
it would take to move past observational (rank-4) evidence. It plays the skeptical reviewer —
the goal is to falsify, not to flatter.
Virtually all published factor research makes associational claims while using tools (OLS, p-values, long-short portfolios) that implicitly treat the factor as a cause. The literature produces spurious findings through three distinct failure modes, which this library treats as first-class concepts:
- Type-A spuriosity — mistaking noise for signal within a single study (p-hacking, backtest overfitting, multiple testing).
- Type-B spuriosity — mistaking a true association for causation through misspecification (missing confounders, controlling colliders, collapsing mediators, specification search).
- Type-C spuriosity — structural identification failure across an entire field: when reported statistics are the maximum of
Klatent within-study trials, the field-level FDR is not identifiable from the cross-section alone and is generally far higher than the optimistic single-trial estimate (López de Prado & Fabozzi 2026, Theorem 1).
Terminology note. Type-A and Type-B are de Prado's (2023, §6.4.1–6.4.2). "Type-C" is this library's label — a natural extension of his A/B taxonomy to the field-level identification failure that López de Prado & Fabozzi (2026) prove in Theorem 1. The result is theirs; only the "Type-C" naming is ours.
The one-line mnemonic:
- A — is it even real? (luck)
- B — is it real because of my strategy? (confounding)
- C — is it real after accounting for everything I tried? (selection)
In plain English, with the fix for each:
- Type-A — noise mistaken for signal. Your "edge" is just luck / overfitting; the pattern isn't real. Fix: HAC-significance, out-of-sample testing, more data.
- Type-B — real correlation, wrong cause. The pattern is real, but a confounder is driving it, not your strategy. Fix: a causal graph — control the confounder.
- Type-C — real for one strategy, but you searched many. Any single backtest looks fine, but you tried 500 and reported the best, so the field's false-discovery rate is far higher than it appears. Fix: search-adjusted FDR.
Without an explicit causal graph, falsification tests, and a search-adjusted view of the field, most factor claims are scientifically weak. This library provides the practical toolkit for all three.
| Version | 0.4.1 (the event-study circular rotation) |
| Tests | 303 passing (pytest) |
| Examples | 19 runnable (examples/01–19) — canonical structures, causal graphs, the evidence hierarchy, the falsification battery, real Ken-French / SPX / rates cases, the search-adjusted FDR, the evaluate-your-strategy / declare-your-graph walkthroughs, the factor-spanning + market-timing control, the owned-strategy overfitting suite (Deflated Sharpe + PBO), the mediator gate (don't neutralize the mechanism your edge runs on), the walk-forward + path-luck spine, the survival / drawdown overlay, and the event-study circular-rotation for single time-collocated events |
| Entry points | evaluate_strategy (returns → A/B/C scorecard) · walk_forward_evaluation (point-in-time consistency + the survival read) · run_falsification_battery (a claim + a graph) |
| Dependencies | numpy, statsmodels, scipy |
| Python | ≥ 3.9 |
Note: this repository is in active development. It is not yet published to PyPI. Install from source (below).
This package is not yet on PyPI. To work with it locally, clone and install in editable mode:
git clone <repository-url>
cd causal-quant
pip install -e ".[test]"Run the test suite to confirm your environment:
pytestfrom causalquant.graphs import CausalGraph
from causalquant.hierarchy import create_claim, EvidenceRank
from causalquant.protocol import run_falsification_battery
# 1. Declare your hypothesized causal mechanism (the key step the book demands)
g = CausalGraph.from_edges([
("MOM", "HML"),
("HML", "OI"),
("OI", "PC"),
("MOM", "PC"),
])
# 2. Run the full falsification battery
report = run_falsification_battery(
treatment="HML",
outcome="PC",
hypothesized_graph=g,
regression_results={
"beta": 0.38,
"pvalue": 0.0008,
"model_spec": "forward_returns ~ HML + controls",
"estimator": "OLS",
"uses_p_values": True,
},
evidence_claims=[
create_claim("Natural experiment on data release timing",
EvidenceRank.NATURAL_EXPERIMENT,
supports_causal_claim=True),
],
)
print(report.overall_risk_level)
print(report.summary)
print(report.prioritized_recommendations[:3])To bring field-level Type-C risk into the same report, pass the cross-section of reported statistics from the strategy's field (e.g. the Chen–Zimmermann predictor zoo):
report = run_falsification_battery(
treatment="HML",
outcome="forward_returns",
hypothesized_graph=g,
field_sharpe_ratios=zoo_sharpes, # np.ndarray of reported in-sample SRs
field_T_values=zoo_T, # per-predictor sample sizes (optional)
field_rho_values=zoo_rho, # per-predictor lag-1 autocorrelations (optional)
)
print(report.field_fdr_calibration) # search-adjusted FDR, fitted KSkip assembling the battery by hand — go straight to the returns-first entry points:
from causalquant import evaluate_strategy, walk_forward_evaluation
# one-shot A / B / C scorecard on a return series vs its benchmark
ev = evaluate_strategy(strat_returns, benchmark_returns, periods_per_year=12)
print(ev.scorecard())
# is the edge real, or a single-window artifact? re-run point-in-time at each fresh entry.
# for a survival overlay, judge it on the tail: statistic="max_drawdown", null_kind="timing"
wf = walk_forward_evaluation(strat_returns, benchmark_returns, in_market_mask)
print(wf.scorecard())See examples/ for fuller walkthroughs.
| Module | Purpose | Maps to |
|---|---|---|
graphs.py |
Declare & query small causal graphs; identify confounders / colliders / mediators; back-door criterion via d-separation | 2023, §4.3, 6.2, 6.4.2 |
simulation.py |
Exact Monte Carlo replications of the three canonical structures (fork, immorality, chain) | 2023, Ch. 7 |
bias.py |
Closed-form bias expressions for under- and over-controlling | 2023, App. A.1–A.2 |
diagnostics.py |
Causal-content scoring, confounder-bias ranges, specification-risk flagging | 2023, §6.1 & 6.4 |
hierarchy.py |
EvidenceRank (six-tier hierarchy) + reporter for assessing a body of evidence |
2023, §6.5 / Fig. 12 |
fdr.py |
Search-adjusted FDR: single-trial vs familywise error rates, max-of-mixtures MLE calibration, Fermi lower bound, non-identifiability gate | 2026, Thm 1, Eqs. 6–37 |
overfitting.py |
Owned-strategy selection tools: Deflated / Probabilistic Sharpe Ratio, Probability of Backtest Overfitting (PBO via CSCV), effective number of trials | Bailey & LdP 2014; Bailey, Borwein, LdP & Zhu 2017 |
metrics.py |
Jensen's α with HAC (Newey–West) errors, Sharpe, drawdown, regime / era splits — the Type-A return engine | 2023, §6.4.1 |
matched_null.py |
Matched-null placebo: does the in/out timing beat random placements (timing) or sign-flips (sign) of the same path — judged on the moment that matches the claim (Sharpe / α / drawdown) |
docs/matched_null_placebo_design.md |
strategy.py |
evaluate_strategy — full A/B/C scorecard on a return series; evaluate_parameter_sweep & Sobol sampled_parameter_sweep overfitting diagnostics |
core |
walkforward.py |
walk_forward_evaluation — re-run the edge battery point-in-time at each fresh entry signal; consistency + per-signal Type-B; the survival (drawdown) read |
de Prado walk-forward / OOS |
protocol.py |
run_falsification_battery() — ties graph + specification + evidence + field-FDR into one report |
core thesis |
types.py |
SpuriosityType (A / B / C), CausalRole, structured diagnosis containers |
taxonomy |
| Example | What it shows |
|---|---|
01_canonical_structures.py |
Reproduces the fork / immorality / chain structures (Ch. 7) and the bias each induces |
02_causal_graphs.py |
Declaring and querying a causal graph; confounder / collider / mediator identification |
03_hierarchy_of_evidence.py |
Scoring a body of evidence on the six-tier hierarchy (Fig. 12) |
04_falsification_protocol.py |
The flagship run_falsification_battery() end-to-end |
05_ken_french_real_data.py |
Full battery on real Ken French factor data |
06_fdr_calibration_chen_zimmermann.py |
Search-adjusted FDR calibration on the factor-zoo cross-section |
07_type_c_integrated_battery.py |
Type-C field-level identification risk folded into the validation report |
08_realworld_spx_moving_average.py |
The full A/B/C battery on the real SPX 200-day moving-average rule |
09_realworld_rates_and_returns.py |
Rates vs returns: Granger direction, HAC, and why it stays observational |
10_evaluate_your_strategy.py |
evaluate_strategy end-to-end on a return series — the deterministic A/B/C scorecard |
11_graph_roles_side_by_side.py |
Confounder vs collider vs mediator, side by side: same arrows, opposite actions |
12_declare_your_strategy_graph.py |
CausalGraph.from_plain_answers — build the graph from jargon-free yes/no answers |
13_news_surprise_rank2_or_rank4.py |
When an exogenous shock earns rank 2 vs when it stays rank 4 |
14_factor_spanning_and_timing.py |
Control a tradable-factor confounder; separate genuine selection from market timing (Treynor–Mazuy) |
15_owned_strategy_overfitting.py |
When the FDR is non-identifiable on your own grid: the reliability gate, the Deflated Sharpe Ratio, and PBO via CSCV |
16_mediator_gate.py |
The mediator gate: don't neutralize (control / vol-target) the mechanism your edge runs on |
17_walk_forward_and_path_luck.py |
The walk-forward spine: point-in-time consistency, the matched-null placebo (path luck), and the graph-inferred Type-B suspect |
18_survival_overlay_drawdown.py |
A survival overlay (carry / crash-dodger): return-α under-credits it, so judge it on the tail — the drawdown walk-forward certifies a consistent survival estimator |
19_event_study_santa_rally.py |
Falsifying a single, time-collocated event where CSCV/PBO don't apply: circular-rotate the event-window mask (path luck) + the confounder graph (real but not causal) — the Santa Claus rally on SPX |
- Severity scores and verdict thresholds in
protocol.pyare deliberately conservative defaults, not estimated constants. They are intended as a transparent starting point for triage; users running the battery at scale should calibrate them to their own field and document the choice. graphs.pyimplements a proper back-door criterion via d-separation (undirected-path enumeration with the correct collider rule and descendant exclusion) for small, explicitly declared graphs. It is not a general do-calculus engine and does not attempt automatic adjustment-set discovery on large graphs (see roadmap).- All randomized components (simulations, calibration) are seeded for reproducibility.
This library is an independent implementation of the methodology described in López de Prado (2023), López de Prado & Fabozzi (2026), and — for the overfitting module — Bailey & López de Prado (2014) and Bailey, Borwein, López de Prado & Zhu (2017). All code, docstrings, examples, and tests in this repository are original work. No verbatim text, figures, or code from the source works is reproduced.
- Code (this repository): released under the MIT License — see
LICENSE. Free to use, modify, and redistribute, including commercially, with attribution to this repository. - Methodology (the 2023 Element): published Open Access by Cambridge University Press under CC-BY-NC-4.0. That license governs the book; it does not propagate to independent code implementations of the methods it describes (under US/UK copyright law, ideas, methods, and procedures are not subject to copyright — only their specific expression is).
The methodological framework and all conceptual contributions belong to Marcos López de Prado (and, for the search-adjusted FDR, to López de Prado & Fabozzi; for the Deflated Sharpe Ratio and the Probability of Backtest Overfitting, to Bailey, López de Prado, and their respective co-authors Borwein and Zhu). Where this library uses terminology coined in those works — "Type-A / Type-B spuriosity," "causal commitment," the six-tier hierarchy of evidence, the three canonical causal structures, search-adjusted FDR — it is with explicit attribution. Conversely, "Type-C spuriosity" is this library's own term — it appears in neither source work; it labels the field-level identification failure that López de Prado & Fabozzi establish as Theorem 1.
- A survivorship-free, point-in-time cross-sectional harness for the asset-selection axis — the one axis a single-asset backtest structurally cannot close
- Expanded simulation coverage (more complex and time-series structures)
- Reporting / visualization of validation reports
- Publish to PyPI
Delivered in v0.4.1: the event-study circular rotation — examples/19_event_study_santa_rally.py
points the matched-null path-luck rotation (kind="timing", scheme="preserve") at a single,
time-collocated event window (where CSCV/PBO structurally don't apply), pairs it with the
confounder-identification layer, and ships an event_mask_from_dates adapt-your-own-event template;
a matched_null_test docstring note makes the in_market generality explicit. 303 tests, 19 examples.
Delivered in v0.4.0: the survival read & scalable Type-C — the walk-forward gains a
statistic="max_drawdown" survival mode (judge a survival overlay on the tail, not return-α),
sampled_parameter_sweep + effective_num_trials make the owned-grid Type-C scale without asking
"how many combos did you try," and the Ken French momentum factor ships as a bundled fixture for
Type-B spanning.
Delivered in v0.3.0: the walk-forward spine — walk_forward_evaluation re-runs the edge
battery point-in-time at each fresh entry signal (consistency + per-signal Type-B), and
matched_null_test adds the path-luck placebo.
Delivered in v0.2.0: the mediator gate — diagnostics.classify_adjustment and
CausalGraph.from_plain_answers turn a declared (or plain-language-elicited) graph into
licensed/blocked adjustment guidance, so you never neutralize the mechanism your edge runs on.
Contributions that improve falsifiability tooling for factor research are welcome.
If you use this library in academic work, please cite this repository alongside the original sources:
@book{lopezdeprado2023causal,
title = {Causal Factor Investing: Can Factor Investing Become Scientific?},
author = {López de Prado, Marcos M.},
year = {2023},
series = {Elements in Quantitative Finance},
publisher = {Cambridge University Press}
}
@techreport{lopezdeprado2026fdr,
title = {The False Discovery Rate in Finance: Identification Failure and Search-Adjusted Estimation},
author = {López de Prado, Marcos M. and Fabozzi, Frank J.},
year = {2026},
number = {24},
institution = {ADIA Lab Research Paper Series}
}
@article{bailey2014deflated,
title = {The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting, and Non-Normality},
author = {Bailey, David H. and López de Prado, Marcos M.},
journal = {Journal of Portfolio Management},
volume = {40},
number = {5},
year = {2014}
}
@article{bailey2017pbo,
title = {The Probability of Backtest Overfitting},
author = {Bailey, David H. and Borwein, Jonathan M. and López de Prado, Marcos M. and Zhu, Qiji J.},
journal = {Journal of Computational Finance},
volume = {20},
number = {4},
year = {2017}
}"...the factor investing community must first wake up from its associational slumber." — López de Prado, Causal Factor Investing (2023)