Skip to content

Repository files navigation

RNN Classifiers for EUR/USD H1 Directional Breakouts

Deep Learning Project — Università degli Studi di Ferrara
A.Y. 2025/2026 — Corso di Deep Learning (Prof. Riccardo Zese)
Author: Marco Perozzi


Recurrent neural networks (LSTM, GRU) for short-term directional breakout prediction in the EUR/USD forex market. Uses the Triple-Barrier Method (TBM) with Purged K-Fold cross-validation under two risk/reward regimes. Three RNN architectures are trained per regime, ensembled over five folds, and evaluated on a chronological 20% holdout test set. Stateful variants (hidden state carried across consecutive batches) are also explored under both Purged K-Fold (V1) and single 70/10/20 split (V2) protocols. The best model overall (GRU_Mid_2L, asymmetric regime) achieves a weighted PR-AUC of 0.5362.

Project status: completed.

Methodology

Data & Preprocessing

  • Source: EUR/USD hourly OHLC from 2002 to 2025 (semicolon-separated CSV, ~200k samples)
  • Train/Test split: chronologically ordered — 80% CV pool, 20% holdout test (no random shuffle)
  • Feature engineering: 14 technical indicators over fast (120-period) and slow (480-period) sliding windows
  • Normalisation: sliding RobustScaler (median/IQR) applied independently inside each fold — never global

Triple-Barrier Labeling

Each trading window of lookback_window=24 hours is labelled via TBM with vertical_barrier=12 hours:

Regime pt_multiplier sl_multiplier Risk/Reward Label distribution (approx.)
Symmetric 3.2 3.2 1:1 ~42% Hold, ~29% Short, ~29% Long
Asymmetric 4.6 2.3 1:2 ~41% Hold, ~43.5% Short, ~15.5% Long

Switching regimes requires only changing these two values in config.yaml, then re-running the data pipeline and CV training.

Purged K-Fold Cross-Validation

  • 5 folds with 12-hour purge (based on vertical barrier) between train/val sets to remove leaked overlapping labels
  • Train samples shuffled due to stateless RNNs (de Prado)
  • 24-hour embargo after each purge, based on lookback window, to further prevent information leakage
  • Model selection via BestModelTracker (10% grace period, max val PR-AUC selection)
  • Inverse-frequency weighted CrossEntropyLoss with power smoothing: beta=1.0 (symmetric), beta=0.85 (asymmetric)

RNN Architectures

Model RNN type Hidden dim Layers RNN dropout FC dropout Weight decay
GRU_2L GRU 32 2 0.25 0.40 1e-2
LSTM_Mid_2L LSTM 48 2 0.40 0.50 7.5e-2
GRU_Mid_2L GRU 64 2 0.30 0.45 2e-2

All models: batch_first=True, stateless (hidden state zeroed per batch) by default, single linear classification head (3 classes). Ensemble inference averages logits across the 5 fold models.

Stateful variants (modules/stateful_pipeline.py): same architectures but with hidden state carried between consecutive batches within each epoch (truncated BPTT). Two protocols:

  • V1 — Purged K-Fold CV with StatefulForexEnsemble (identical structure to the stateless pipeline, only shuffle=False)
  • V2 — Single 70/10/20 chronological split with no ensemble, no threshold optimisation

Threshold Optimization

Decision thresholds for Short and Long classes (Hold is the default at 0) are optimised on OOF predictions via 2D grid search:

  1. Academic Global Optimum: maximises macro F1 — (F1_Short + F1_Hold + F1_Long) / 3
  2. Financial Optimum: maximises expectancy — reward × TP − risk × FP

Results

Symmetric Regime (PT=3.2, SL=3.2)

Model PR-AUC ↑ AUROC ↑ Accuracy F1 Precision Recall
GRU_2L 0.5102 0.6874 49.34% 48.72% 48.70% 49.34%
LSTM_Mid_2L 0.4994 0.6754 48.34% 47.87% 47.53% 48.34%
GRU_Mid_2L 0.5104 0.6869 49.47% 48.96% 48.93% 49.47%

Asymmetric Regime (PT=4.6, SL=2.3)

Model PR-AUC ↑ AUROC ↑ Accuracy F1 Precision Recall
GRU_2L 0.5351 0.6749 50.17% 50.30% 50.88% 50.17%
LSTM_Mid_2L 0.5257 0.6659 50.18% 49.98% 50.35% 50.18%
GRU_Mid_2L 0.5362 0.6754 50.92% 50.93% 51.24% 50.92%

The asymmetric regime yields higher PR-AUC across all models — the 1:2 risk/reward setup makes the classification task more discriminable. GRU_Mid_2L (64 hidden units, GRU cells) consistently outperforms GRU_2L and LSTM_Mid_2L in both regimes.

Stateful Results

Stateful models perform near-identically to their stateless counterparts (ΔPR-AUC < 0.005), confirming that the H1 market is predominantly Markovian — carrying hidden state across 24-hour windows provides no measurable benefit.

Regime Best Model Protocol PR-AUC Accuracy
Symmetric GRU_Mid_2L V1 (KFold) 0.5107 49.96%
Symmetric GRU_Mid_2L V2 (70/10/20) 0.5082 49.55%
Asymmetric GRU_Mid_2L V1 (KFold) 0.5367 51.07%
Asymmetric GRU_Mid_2L V2 (70/10/20) 0.5377 52.98%

Project Structure

.
├── modules/                      # Python backend (namespace package)
│   ├── config.py                 # YAML config loader
│   ├── data_processing.py        # Data loading, features, TBM, PurgedKFold
│   ├── models.py                 # ForexRNN (LSTM/GRU), ForexEnsemble
│   ├── training.py               # Train loop, Purged CV pipeline, test eval
│   ├── stateful_pipeline.py      # Stateful RNN experiments (V1 KFold + V2 70/10/20)
│   ├── evaluation.py             # Plots, threshold optim, KDE analysis
│   └── utils.py                  # Logging, seed, save/load vaults
├── docs/
│   └── report.md                 # Full technical report
├── config.yaml                   # All hyperparameters
├── 01_data_preparation.ipynb     # Cyclical features (hour/dow sin/cos), pips
├── 02_training_pipeline.ipynb    # Core pipeline: features, TBM, CV, test, figures
├── data/                         # Source CSV (gitignored)
│   └── EURUSD_download_notebook.ipynb  # M1→H1 download & conversion
├── results/                      # *.pt experiment vaults (gitignored)
├── figures/                      # Generated comparison plots (gitignored)
└── README.md

Setup & Usage

git clone <repo>
cd deep_learning_projectwork
uv sync

Data Preparation

Run the notebooks in order:

  1. data/EURUSD_download_notebook.ipynb — downloads M1 tick data, converts to H1 candles
  2. 01_data_preparation.ipynb — cyclical features (hour/dow sin/cos) and pips
  3. 02_training_pipeline.ipynb — feature engineering, EWMA volatility, TBM labels, Purged CV (k=5, 3 architectures), ensemble evaluation, comparison figures

Switching Trading Regime

Edit config.yaml:

pt_multiplier: 4.6   # Take Profit in ATR multiples
sl_multiplier: 2.3   # Stop Loss in ATR multiples

Then re-run 02_training_pipeline.ipynb. If the prepared H1 data isn't already available, also re-run 01_data_preparation.ipynb.

Key Gotchas

  • Logging is initialised once via setup_global_logging("pipeline.log") — called in the notebooks
  • torch.load(..., weights_only=False) required to load vaults (contain NumPy arrays inside dicts)
  • Data pipeline reads config.get("normalization_window_fast", 120) — consistent with config.yaml

Requirements

  • Python 3.10+
  • CUDA-capable GPU recommended
  • LaTeX distribution with lualatex and latexmk (for presentation)
  • See pyproject.toml for full dependency list (managed via uv)

References

  • Triple-Barrier Method & Purged K-Fold CV: M. Prado, Advances in Financial Machine Learning (2018)
  • Dataset: EUR/USD historical data (H1, 2002–2025), publicly available from HistData

About

EUR/USD H1 directional breakout prediction using RNNs (LSTM/GRU) with Triple-Barrier Method and Purged K-Fold Cross-Validation in PyTorch.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages