Skip to content

Latest commit

 

History

151 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

causal-quant

Causal validation and falsification protocols for factor investing

Most factor research can't tell signal from self-deception. causal-quant is the validation layer that can. Declare a causal graph, run a falsification battery, and quantify how much of a reported edge could survive search and selection — implementing the published de Prado (2023/2026) and Bailey (2014/2017) protocols, not hand-waving. It pins down the three ways a backtest lies — luck, confounding, and selection across everything you tried — and returns HAC-correct, skeptical verdicts instead of flattering p-values. Built to falsify your own strategies before you risk capital.

A Python library that codifies the methodological framework from:

López de Prado, M. (2023). Causal Factor Investing: Can Factor Investing Become Scientific? Cambridge University Press (Elements in Quantitative Finance).

and the search-adjusted false-discovery-rate machinery from:

López de Prado, M., and F. Fabozzi (2026). The False Discovery Rate in Finance: Identification Failure and Search-Adjusted Estimation. ADIA Lab Research Paper Series, No. 24.

For the single-strategy selection problem (where the field-level FDR does not apply), it also implements the deflated-Sharpe and backtest-overfitting tools from:

Bailey, D. H., and M. López de Prado (2014). The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting, and Non-Normality. Journal of Portfolio Management 40 (5).

Bailey, D. H., J. Borwein, M. López de Prado, and Q. J. Zhu (2017). The Probability of Backtest Overfitting. Journal of Computational Finance 20 (4).

The library helps researchers and practitioners move factor strategies out of the "phenomenological" stage (associational claims, p-hacking, specification errors) and toward falsifiable, causally grounded research.

It is not another backtesting or factor-zoo library — it is the validation layer that sits after you have a candidate edge, to test whether that edge is real, identified, and robust to search.


Using this with Claude Code

The fastest way to put this library to work on your strategy is to let Claude drive it:

  1. Clone the repo and run claude inside it.
  2. Say "evaluate my strategy" — then either describe it (the claim, a return series + a benchmark, and honestly how many specs you tried) or point Claude at a local returns file.

Claude auto-loads CLAUDE.md and follows RUNBOOK.md — the guided A → B → C → evidence-ceiling protocol — running the real functions (evaluate_strategy, jensens_alpha, CausalGraph, the FDR suite, run_falsification_battery) and returning an honest verdict: is there an edge (HAC, not OLS), is it identified, did you data-mine, and what it would take to move past observational (rank-4) evidence. It plays the skeptical reviewer — the goal is to falsify, not to flatter.


The Problem

Virtually all published factor research makes associational claims while using tools (OLS, p-values, long-short portfolios) that implicitly treat the factor as a cause. The literature produces spurious findings through three distinct failure modes, which this library treats as first-class concepts:

  • Type-A spuriosity — mistaking noise for signal within a single study (p-hacking, backtest overfitting, multiple testing).
  • Type-B spuriosity — mistaking a true association for causation through misspecification (missing confounders, controlling colliders, collapsing mediators, specification search).
  • Type-C spuriosity — structural identification failure across an entire field: when reported statistics are the maximum of K latent within-study trials, the field-level FDR is not identifiable from the cross-section alone and is generally far higher than the optimistic single-trial estimate (López de Prado & Fabozzi 2026, Theorem 1).

Terminology note. Type-A and Type-B are de Prado's (2023, §6.4.1–6.4.2). "Type-C" is this library's label — a natural extension of his A/B taxonomy to the field-level identification failure that López de Prado & Fabozzi (2026) prove in Theorem 1. The result is theirs; only the "Type-C" naming is ours.

The one-line mnemonic:

  • A — is it even real? (luck)
  • B — is it real because of my strategy? (confounding)
  • C — is it real after accounting for everything I tried? (selection)

In plain English, with the fix for each:

  • Type-A — noise mistaken for signal. Your "edge" is just luck / overfitting; the pattern isn't real. Fix: HAC-significance, out-of-sample testing, more data.
  • Type-B — real correlation, wrong cause. The pattern is real, but a confounder is driving it, not your strategy. Fix: a causal graph — control the confounder.
  • Type-C — real for one strategy, but you searched many. Any single backtest looks fine, but you tried 500 and reported the best, so the field's false-discovery rate is far higher than it appears. Fix: search-adjusted FDR.

Without an explicit causal graph, falsification tests, and a search-adjusted view of the field, most factor claims are scientifically weak. This library provides the practical toolkit for all three.


Status

Version 0.4.1 (the event-study circular rotation)
Tests 303 passing (pytest)
Examples 19 runnable (examples/01–19) — canonical structures, causal graphs, the evidence hierarchy, the falsification battery, real Ken-French / SPX / rates cases, the search-adjusted FDR, the evaluate-your-strategy / declare-your-graph walkthroughs, the factor-spanning + market-timing control, the owned-strategy overfitting suite (Deflated Sharpe + PBO), the mediator gate (don't neutralize the mechanism your edge runs on), the walk-forward + path-luck spine, the survival / drawdown overlay, and the event-study circular-rotation for single time-collocated events
Entry points evaluate_strategy (returns → A/B/C scorecard) · walk_forward_evaluation (point-in-time consistency + the survival read) · run_falsification_battery (a claim + a graph)
Dependencies numpy, statsmodels, scipy
Python ≥ 3.9

Note: this repository is in active development. It is not yet published to PyPI. Install from source (below).


Installation

This package is not yet on PyPI. To work with it locally, clone and install in editable mode:

git clone <repository-url>
cd causal-quant
pip install -e ".[test]"

Run the test suite to confirm your environment:

pytest

Quickstart

from causalquant.graphs import CausalGraph
from causalquant.hierarchy import create_claim, EvidenceRank
from causalquant.protocol import run_falsification_battery

# 1. Declare your hypothesized causal mechanism (the key step the book demands)
g = CausalGraph.from_edges([
    ("MOM", "HML"),
    ("HML", "OI"),
    ("OI", "PC"),
    ("MOM", "PC"),
])

# 2. Run the full falsification battery
report = run_falsification_battery(
    treatment="HML",
    outcome="PC",
    hypothesized_graph=g,
    regression_results={
        "beta": 0.38,
        "pvalue": 0.0008,
        "model_spec": "forward_returns ~ HML + controls",
        "estimator": "OLS",
        "uses_p_values": True,
    },
    evidence_claims=[
        create_claim("Natural experiment on data release timing",
                     EvidenceRank.NATURAL_EXPERIMENT,
                     supports_causal_claim=True),
    ],
)

print(report.overall_risk_level)
print(report.summary)
print(report.prioritized_recommendations[:3])

To bring field-level Type-C risk into the same report, pass the cross-section of reported statistics from the strategy's field (e.g. the Chen–Zimmermann predictor zoo):

report = run_falsification_battery(
    treatment="HML",
    outcome="forward_returns",
    hypothesized_graph=g,
    field_sharpe_ratios=zoo_sharpes,   # np.ndarray of reported in-sample SRs
    field_T_values=zoo_T,              # per-predictor sample sizes (optional)
    field_rho_values=zoo_rho,          # per-predictor lag-1 autocorrelations (optional)
)
print(report.field_fdr_calibration)    # search-adjusted FDR, fitted K

Already have a return series?

Skip assembling the battery by hand — go straight to the returns-first entry points:

from causalquant import evaluate_strategy, walk_forward_evaluation

# one-shot A / B / C scorecard on a return series vs its benchmark
ev = evaluate_strategy(strat_returns, benchmark_returns, periods_per_year=12)
print(ev.scorecard())

# is the edge real, or a single-window artifact? re-run point-in-time at each fresh entry.
# for a survival overlay, judge it on the tail:  statistic="max_drawdown", null_kind="timing"
wf = walk_forward_evaluation(strat_returns, benchmark_returns, in_market_mask)
print(wf.scorecard())

See examples/ for fuller walkthroughs.


Module Overview

Module Purpose Maps to
graphs.py Declare & query small causal graphs; identify confounders / colliders / mediators; back-door criterion via d-separation 2023, §4.3, 6.2, 6.4.2
simulation.py Exact Monte Carlo replications of the three canonical structures (fork, immorality, chain) 2023, Ch. 7
bias.py Closed-form bias expressions for under- and over-controlling 2023, App. A.1–A.2
diagnostics.py Causal-content scoring, confounder-bias ranges, specification-risk flagging 2023, §6.1 & 6.4
hierarchy.py EvidenceRank (six-tier hierarchy) + reporter for assessing a body of evidence 2023, §6.5 / Fig. 12
fdr.py Search-adjusted FDR: single-trial vs familywise error rates, max-of-mixtures MLE calibration, Fermi lower bound, non-identifiability gate 2026, Thm 1, Eqs. 6–37
overfitting.py Owned-strategy selection tools: Deflated / Probabilistic Sharpe Ratio, Probability of Backtest Overfitting (PBO via CSCV), effective number of trials Bailey & LdP 2014; Bailey, Borwein, LdP & Zhu 2017
metrics.py Jensen's α with HAC (Newey–West) errors, Sharpe, drawdown, regime / era splits — the Type-A return engine 2023, §6.4.1
matched_null.py Matched-null placebo: does the in/out timing beat random placements (timing) or sign-flips (sign) of the same path — judged on the moment that matches the claim (Sharpe / α / drawdown) docs/matched_null_placebo_design.md
strategy.py evaluate_strategy — full A/B/C scorecard on a return series; evaluate_parameter_sweep & Sobol sampled_parameter_sweep overfitting diagnostics core
walkforward.py walk_forward_evaluation — re-run the edge battery point-in-time at each fresh entry signal; consistency + per-signal Type-B; the survival (drawdown) read de Prado walk-forward / OOS
protocol.py run_falsification_battery() — ties graph + specification + evidence + field-FDR into one report core thesis
types.py SpuriosityType (A / B / C), CausalRole, structured diagnosis containers taxonomy

Examples

Example What it shows
01_canonical_structures.py Reproduces the fork / immorality / chain structures (Ch. 7) and the bias each induces
02_causal_graphs.py Declaring and querying a causal graph; confounder / collider / mediator identification
03_hierarchy_of_evidence.py Scoring a body of evidence on the six-tier hierarchy (Fig. 12)
04_falsification_protocol.py The flagship run_falsification_battery() end-to-end
05_ken_french_real_data.py Full battery on real Ken French factor data
06_fdr_calibration_chen_zimmermann.py Search-adjusted FDR calibration on the factor-zoo cross-section
07_type_c_integrated_battery.py Type-C field-level identification risk folded into the validation report
08_realworld_spx_moving_average.py The full A/B/C battery on the real SPX 200-day moving-average rule
09_realworld_rates_and_returns.py Rates vs returns: Granger direction, HAC, and why it stays observational
10_evaluate_your_strategy.py evaluate_strategy end-to-end on a return series — the deterministic A/B/C scorecard
11_graph_roles_side_by_side.py Confounder vs collider vs mediator, side by side: same arrows, opposite actions
12_declare_your_strategy_graph.py CausalGraph.from_plain_answers — build the graph from jargon-free yes/no answers
13_news_surprise_rank2_or_rank4.py When an exogenous shock earns rank 2 vs when it stays rank 4
14_factor_spanning_and_timing.py Control a tradable-factor confounder; separate genuine selection from market timing (Treynor–Mazuy)
15_owned_strategy_overfitting.py When the FDR is non-identifiable on your own grid: the reliability gate, the Deflated Sharpe Ratio, and PBO via CSCV
16_mediator_gate.py The mediator gate: don't neutralize (control / vol-target) the mechanism your edge runs on
17_walk_forward_and_path_luck.py The walk-forward spine: point-in-time consistency, the matched-null placebo (path luck), and the graph-inferred Type-B suspect
18_survival_overlay_drawdown.py A survival overlay (carry / crash-dodger): return-α under-credits it, so judge it on the tail — the drawdown walk-forward certifies a consistent survival estimator
19_event_study_santa_rally.py Falsifying a single, time-collocated event where CSCV/PBO don't apply: circular-rotate the event-window mask (path luck) + the confounder graph (real but not causal) — the Santa Claus rally on SPX

Design Notes

  • Severity scores and verdict thresholds in protocol.py are deliberately conservative defaults, not estimated constants. They are intended as a transparent starting point for triage; users running the battery at scale should calibrate them to their own field and document the choice.
  • graphs.py implements a proper back-door criterion via d-separation (undirected-path enumeration with the correct collider rule and descendant exclusion) for small, explicitly declared graphs. It is not a general do-calculus engine and does not attempt automatic adjustment-set discovery on large graphs (see roadmap).
  • All randomized components (simulations, calibration) are seeded for reproducibility.

License & Attribution

This library is an independent implementation of the methodology described in López de Prado (2023), López de Prado & Fabozzi (2026), and — for the overfitting module — Bailey & López de Prado (2014) and Bailey, Borwein, López de Prado & Zhu (2017). All code, docstrings, examples, and tests in this repository are original work. No verbatim text, figures, or code from the source works is reproduced.

  • Code (this repository): released under the MIT License — see LICENSE. Free to use, modify, and redistribute, including commercially, with attribution to this repository.
  • Methodology (the 2023 Element): published Open Access by Cambridge University Press under CC-BY-NC-4.0. That license governs the book; it does not propagate to independent code implementations of the methods it describes (under US/UK copyright law, ideas, methods, and procedures are not subject to copyright — only their specific expression is).

The methodological framework and all conceptual contributions belong to Marcos López de Prado (and, for the search-adjusted FDR, to López de Prado & Fabozzi; for the Deflated Sharpe Ratio and the Probability of Backtest Overfitting, to Bailey, López de Prado, and their respective co-authors Borwein and Zhu). Where this library uses terminology coined in those works — "Type-A / Type-B spuriosity," "causal commitment," the six-tier hierarchy of evidence, the three canonical causal structures, search-adjusted FDR — it is with explicit attribution. Conversely, "Type-C spuriosity" is this library's own term — it appears in neither source work; it labels the field-level identification failure that López de Prado & Fabozzi establish as Theorem 1.


Roadmap (toward v0.5)

  • A survivorship-free, point-in-time cross-sectional harness for the asset-selection axis — the one axis a single-asset backtest structurally cannot close
  • Expanded simulation coverage (more complex and time-series structures)
  • Reporting / visualization of validation reports
  • Publish to PyPI

Delivered in v0.4.1: the event-study circular rotation — examples/19_event_study_santa_rally.py points the matched-null path-luck rotation (kind="timing", scheme="preserve") at a single, time-collocated event window (where CSCV/PBO structurally don't apply), pairs it with the confounder-identification layer, and ships an event_mask_from_dates adapt-your-own-event template; a matched_null_test docstring note makes the in_market generality explicit. 303 tests, 19 examples.

Delivered in v0.4.0: the survival read & scalable Type-C — the walk-forward gains a statistic="max_drawdown" survival mode (judge a survival overlay on the tail, not return-α), sampled_parameter_sweep + effective_num_trials make the owned-grid Type-C scale without asking "how many combos did you try," and the Ken French momentum factor ships as a bundled fixture for Type-B spanning.

Delivered in v0.3.0: the walk-forward spine — walk_forward_evaluation re-runs the edge battery point-in-time at each fresh entry signal (consistency + per-signal Type-B), and matched_null_test adds the path-luck placebo.

Delivered in v0.2.0: the mediator gate — diagnostics.classify_adjustment and CausalGraph.from_plain_answers turn a declared (or plain-language-elicited) graph into licensed/blocked adjustment guidance, so you never neutralize the mechanism your edge runs on.

Contributions that improve falsifiability tooling for factor research are welcome.


Citation

If you use this library in academic work, please cite this repository alongside the original sources:

@book{lopezdeprado2023causal,
  title     = {Causal Factor Investing: Can Factor Investing Become Scientific?},
  author    = {López de Prado, Marcos M.},
  year      = {2023},
  series    = {Elements in Quantitative Finance},
  publisher = {Cambridge University Press}
}

@techreport{lopezdeprado2026fdr,
  title       = {The False Discovery Rate in Finance: Identification Failure and Search-Adjusted Estimation},
  author      = {López de Prado, Marcos M. and Fabozzi, Frank J.},
  year        = {2026},
  number      = {24},
  institution = {ADIA Lab Research Paper Series}
}

@article{bailey2014deflated,
  title     = {The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting, and Non-Normality},
  author    = {Bailey, David H. and López de Prado, Marcos M.},
  journal   = {Journal of Portfolio Management},
  volume    = {40},
  number    = {5},
  year      = {2014}
}

@article{bailey2017pbo,
  title     = {The Probability of Backtest Overfitting},
  author    = {Bailey, David H. and Borwein, Jonathan M. and López de Prado, Marcos M. and Zhu, Qiji J.},
  journal   = {Journal of Computational Finance},
  volume    = {20},
  number    = {4},
  year      = {2017}
}

"...the factor investing community must first wake up from its associational slumber." — López de Prado, Causal Factor Investing (2023)

About

Independent, MIT-licensed Python implementation of the López de Prado causal factor-investing and falsification framework. Run a falsification battery, measure how much of a backtest edge survives search and selection, and falsify your edge before you risk capital.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages