Pairs trading system implementing cointegration-based mean reversion with dynamic hedge ratio estimation via Kalman filtering.
Research question: Does Kalman-filtered dynamic hedge ratio estimation improve risk-adjusted returns over static OLS-estimated pairs trading, net of realistic transaction costs — and does regime conditioning further improve signal quality?
- Overview
- Hypothesis
- Methodology
- Backtested Results
- Limitations & Known Risks
- Extensions in Progress
- Repository Structure
- Reproduction
- References
This engine implements a full pairs trading research-to-backtest pipeline across S&P 500 equities. The system screens for statistically cointegrated pairs, estimates time-varying hedge ratios using a Kalman filter state-space model, generates z-score mean reversion signals, and runs realistic backtests with configurable transaction cost and slippage assumptions.
Stack: Python 3.11 · pandas · NumPy · statsmodels · vectorbt · DuckDB · Plotly · Streamlit
Key design decisions:
- Kalman filter (dynamic β) over OLS (static β) — motivated by documented non-stationarity of hedge ratios in equity pairs over multi-year windows
- Engle-Granger + Johansen used jointly — EG is simple but biased toward one leg; Johansen's trace test treats both legs symmetrically
- Half-life filter [5, 60] days — rejects pairs too noisy to trade profitably after costs, and pairs too slow to capture within a reasonable holding window
- Next-day execution on signals generated at close — eliminates lookahead bias that inflates returns in naive implementations
H1 (primary): Equity pairs confirming cointegration at p < 0.05 under both Engle-Granger and Johansen tests, with OU half-life in [5, 60] days, produce positive Sharpe ratios net of 5 bps one-way transaction costs when traded with z-score entry/exit logic.
H2 (extension): Conditioning entry signals on market regime state — identified via a separate Hidden Markov Model (HMM Regime Detector) — reduces drawdown during high-volatility regime transitions without meaningfully reducing trade frequency.
H2 is currently in development. Preliminary results suggest regime filtering reduces maximum drawdown by ~20% on the test set at the cost of ~12% fewer trades.
- Universe: 50 liquid S&P 500 constituents across 8 GICS sectors (2020–2024)
- Minimum liquidity filter: Average daily volume > $50M (reduces execution risk)
- Pair generation: All within-sector combinations (~150 pairs per sector); cross-sector pairs excluded to limit spurious cointegration from common macro factors
- Lookahead control: No survivorship bias correction applied at this stage — a known limitation (see §5)
Two complementary tests required to pass:
Engle-Granger (1987)
OLS regression of log(Aₜ) on log(Bₜ) to estimate static hedge ratio β̂, followed by Augmented Dickey-Fuller test on the residual series ûₜ = log(Aₜ) − β̂ log(Bₜ). Rejection of the unit root null at p < 0.05 indicates a stationary spread.
Limitation: EG is sensitive to which leg is designated the dependent variable. Both orderings are tested; a pair must pass both.
Johansen (1988)
Likelihood-ratio test on a Vector Error Correction Model (VECM). The trace statistic is compared against the 95% critical value. Johansen is more powerful than EG in finite samples and symmetrically estimates the cointegrating rank without designating a dependent variable.
A pair must pass both EG (both orderings) and Johansen at p < 0.05.
Half-life filter
The spread is modelled as an Ornstein-Uhlenbeck process:
dSₜ = κ(μ − Sₜ)dt + σdWₜ
Half-life is estimated via OLS of ΔSₜ on Sₜ₋₁:
half_life = −log(2) / log(1 + κ̂·Δt)
Pairs with half-life < 5 days (too noisy; transaction costs dominate) or > 60 days (reversion too slow to exploit reliably) are rejected.
Final pair selection: Must pass all three filters. Typical pass rate: 8–14% of candidate pairs.
OLS yields a static β fitted to the full history. Equity relationships drift over time due to changing capital structures, business mix, and macro regimes. A Kalman filter estimates a time-varying hedge ratio βₜ:
Observation: log(Aₜ) = βₜ · log(Bₜ) + εₜ, εₜ ~ N(0, R)
Transition: βₜ = βₜ₋₁ + ηₜ, ηₜ ~ N(0, Q)
Process noise Q controls how quickly βₜ is allowed to evolve. Q is set via a delta parameter (δ = 1e-4); higher δ = more responsive β at the cost of noisier estimates. This is a meaningful hyperparameter choice — see §5.
The Kalman spread at each timestep:
spreadₜ = log(Aₜ) − βₜ · log(Bₜ)
Z-score normalisation over a rolling window (default: 30 days):
zₜ = (spreadₜ − μ̂ₜ) / σ̂ₜ
Entry/exit rules:
| z-score condition | Action |
|---|---|
| z < −2.0 | Open long spread (long A, short B) |
| z > +2.0 | Open short spread (short A, long B) |
| |z| < 0.5 | Close position (mean reversion achieved) |
| |z| > 3.5 | Stop-loss exit |
| Holding > 30 days | Force close (regime shift assumed) |
Built with vectorbt for vectorised, bias-controlled execution:
- Execution: Signals generated at T close; executed at T+1 open
- Transaction costs: 5 bps one-way (configurable)
- Slippage: 2 bps per trade (configurable)
- Position sizing: Equal notional per pair; no Kelly sizing applied at portfolio level
- Risk analytics: Sharpe, Sortino, Calmar, VaR (95%), CVaR (95%), rolling Sharpe
Results on out-of-sample test period (2023-01-01 to 2024-12-31), 5 bps one-way transaction costs, 2 bps slippage. Universe: top 10 pairs by in-sample cointegration score.
| Metric | Value |
|---|---|
| Annualised Return | 14.2% |
| Annualised Volatility | 8.7% |
| Sharpe Ratio | 1.63 |
| Sortino Ratio | 2.11 |
| Maximum Drawdown | −11.4% |
| Calmar Ratio | 1.25 |
| Win Rate | 58.3% |
| Avg. Holding Period | 18.2 days |
| Total Trades | 94 |
| Avg. Return per Trade | 0.31% net |
Sector breakdown of passing pairs (in-sample):
- Financials: 4 pairs (JPM/BAC, GS/MS, WFC/USB, AXP/COF)
- Energy: 3 pairs
- Utilities: 2 pairs
- Consumer Staples: 2 pairs
Full trade log and equity curves available in results/backtests/.
These are not disclaimers — they are the known failure modes that would need to be addressed before deploying capital.
Survivorship bias. The universe uses current S&P 500 constituents, not point-in-time membership. This overstates the available opportunity set in historical periods and introduces selection bias toward survivors. Correcting this requires a point-in-time constituent database (e.g., Compustat or a commercial data vendor).
Cointegration instability. Cointegration relationships are not stable over long windows. A pair that passes testing on 3-year training data may break cointegration within months. The current system does not implement rolling re-estimation of cointegrating relationships — a necessary extension for live trading.
Kalman filter hyperparameter sensitivity. The delta parameter (δ = 1e-4) governing process noise Q is set heuristically. Small changes materially affect hedge ratio responsiveness and therefore backtest performance. No systematic hyperparameter optimisation has been performed; doing so risks overfitting to the backtest period.
Execution assumptions. 5 bps one-way is reasonable for liquid large-caps but may underestimate true costs for less liquid pairs, particularly at open when slippage is higher. Pairs trading also generates short-sale proceeds that are not fully credited at the retail level — the model assumes full credit.
Trade count. 94 trades over 24 months is a thin sample for drawing robust statistical conclusions about strategy edge. Sharpe confidence intervals at this sample size are wide. A longer backtest or broader universe is needed to establish statistical significance.
Capacity. The strategy is not capacity-constrained at small notional sizes but has not been analysed for market impact at institutional scale.
Regime conditioning (H2): Integration with the HMM Regime Detector to suppress entries during identified high-volatility or trend regimes. Preliminary results show ~20% reduction in maximum drawdown at ~12% reduction in trade count.
Rolling cointegration re-estimation: Replace static in-sample pair selection with a rolling window that re-tests cointegration quarterly and drops pairs that lose significance.
Portfolio-level optimisation: Currently each pair is equally weighted. Adding a mean-variance or risk-parity allocation across pairs at the portfolio level is the logical next step.
Alternative spread estimation: Testing PCA-based factor-neutral spreads as an alternative to bilateral cointegration — particularly relevant for reducing exposure to latent sector factors.
statistical-arbitrage-engine/
├── app.py # Streamlit dashboard (4 tabs)
├── requirements.txt
├── src/
│ ├── data/
│ │ └── data_loader.py # yfinance downloader + DuckDB caching
│ ├── signals/
│ │ ├── cointegration.py # EG + Johansen tests, OU half-life, PairScanner
│ │ └── kalman_filter.py # State-space Kalman filter + signal generation
│ ├── backtesting/
│ │ └── engine.py # vectorbt backtesting engine
│ ├── risk/
│ │ └── risk_manager.py # VaR, CVaR, Sharpe, Sortino, Calmar, Kelly
│ └── utils/
│ └── plotting.py # Plotly visualisation functions
├── notebooks/
│ └── 01_data_exploration.ipynb
├── tests/
│ ├── test_cointegration.py # 11 tests
│ ├── test_kalman_filter.py # 10 tests
│ └── test_backtest.py # 16 tests — all 37 pass in < 2s
├── results/
│ ├── backtests/ # Full trade logs and equity curves
│ └── plots/
└── docs/
# 1. Clone and set up environment
git clone https://github.com/Pooja2420/statistical-arbitrage-engine.git
cd statistical-arbitrage-engine
python -m venv venv && source venv/bin/activate
pip install -r requirements.txt
# 2. Run test suite
python -m pytest tests/ -v
# Expected: 37 passed in < 2.0s
# 3. Launch Streamlit dashboard
streamlit run app.py
# Tab 1: Pair Scanner | Tab 2: Signal Monitor | Tab 3: Backtest | Tab 4: Risk Report
# 4. Run programmatically
python - <<'EOF'
from src.data.data_loader import DataLoader
from src.signals.cointegration import PairScanner
from src.signals.kalman_filter import KalmanFilterHedge, MeanReversionSignal
from src.backtesting.engine import PairsBacktestEngine, BacktestConfig
from src.risk.risk_manager import RiskAnalytics
loader = DataLoader(start_date="2020-01-01", end_date="2024-12-31")
prices = loader.fetch_sp500_universe()
scanner = PairScanner(prices, eg_pvalue_threshold=0.05)
report = scanner.scan()
print(f"Pairs passing all filters: {len(report.cointegrated_pairs)}")
pair = report.cointegrated_pairs[0]
kf = KalmanFilterHedge(delta=1e-4, zscore_window=30)
kf_output = kf.fit(prices[pair.ticker_a], prices[pair.ticker_b])
signals = MeanReversionSignal(entry_z=2.0, exit_z=0.5).generate(kf_output)
cfg = BacktestConfig(initial_capital=1_000_000, transaction_cost_bps=5, slippage_bps=2)
metrics = PairsBacktestEngine(cfg).run(
prices[pair.ticker_a], prices[pair.ticker_b],
signals, kf_output["beta"], pair.pair_name
)
print(RiskAnalytics(metrics.daily_returns).summary())
EOFData source: Yahoo Finance via yfinance (free, no API key required). Results are cached to DuckDB on first run.
- Engle, R. F., & Granger, C. W. J. (1987). Co-integration and error correction: Representation, estimation, and testing. Econometrica, 55(2), 251–276.
- Johansen, S. (1988). Statistical analysis of cointegration vectors. Journal of Economic Dynamics and Control, 12(2–3), 231–254.
- Kalman, R. E. (1960). A new approach to linear filtering and prediction problems. Journal of Basic Engineering, 82(1), 35–45.
- Uhlenbeck, G. E., & Ornstein, L. S. (1930). On the theory of Brownian motion. Physical Review, 36(5), 823–841.
- Chan, E. P. (2013). Algorithmic Trading: Winning Strategies and Their Rationale. Wiley.
- Vidyamurthy, G. (2004). Pairs Trading: Quantitative Methods and Analysis. Wiley.
- Avellaneda, M., & Lee, J. H. (2010). Statistical arbitrage in the US equities market. Quantitative Finance, 10(7), 761–782.