Deep Learning Project — Università degli Studi di Ferrara
A.Y. 2025/2026 — Corso di Deep Learning (Prof. Riccardo Zese)
Author: Marco Perozzi
Recurrent neural networks (LSTM, GRU) for short-term directional breakout prediction in the EUR/USD forex market. Uses the Triple-Barrier Method (TBM) with Purged K-Fold cross-validation under two risk/reward regimes. Three RNN architectures are trained per regime, ensembled over five folds, and evaluated on a chronological 20% holdout test set. Stateful variants (hidden state carried across consecutive batches) are also explored under both Purged K-Fold (V1) and single 70/10/20 split (V2) protocols. The best model overall (GRU_Mid_2L, asymmetric regime) achieves a weighted PR-AUC of 0.5362.
Project status: completed.
- Source: EUR/USD hourly OHLC from 2002 to 2025 (semicolon-separated CSV, ~200k samples)
- Train/Test split: chronologically ordered — 80% CV pool, 20% holdout test (no random shuffle)
- Feature engineering: 14 technical indicators over fast (120-period) and slow (480-period) sliding windows
- Normalisation: sliding RobustScaler (median/IQR) applied independently inside each fold — never global
Each trading window of lookback_window=24 hours is labelled via TBM with vertical_barrier=12 hours:
| Regime | pt_multiplier |
sl_multiplier |
Risk/Reward | Label distribution (approx.) |
|---|---|---|---|---|
| Symmetric | 3.2 | 3.2 | 1:1 | ~42% Hold, ~29% Short, ~29% Long |
| Asymmetric | 4.6 | 2.3 | 1:2 | ~41% Hold, ~43.5% Short, ~15.5% Long |
Switching regimes requires only changing these two values in config.yaml, then re-running the data pipeline and CV training.
- 5 folds with 12-hour purge (based on vertical barrier) between train/val sets to remove leaked overlapping labels
- Train samples shuffled due to stateless RNNs (de Prado)
- 24-hour embargo after each purge, based on lookback window, to further prevent information leakage
- Model selection via
BestModelTracker(10% grace period, max val PR-AUC selection) - Inverse-frequency weighted CrossEntropyLoss with power smoothing:
beta=1.0(symmetric),beta=0.85(asymmetric)
| Model | RNN type | Hidden dim | Layers | RNN dropout | FC dropout | Weight decay |
|---|---|---|---|---|---|---|
| GRU_2L | GRU | 32 | 2 | 0.25 | 0.40 | 1e-2 |
| LSTM_Mid_2L | LSTM | 48 | 2 | 0.40 | 0.50 | 7.5e-2 |
| GRU_Mid_2L | GRU | 64 | 2 | 0.30 | 0.45 | 2e-2 |
All models: batch_first=True, stateless (hidden state zeroed per batch) by default, single linear classification head (3 classes). Ensemble inference averages logits across the 5 fold models.
Stateful variants (modules/stateful_pipeline.py): same architectures but with hidden state carried between consecutive batches within each epoch (truncated BPTT). Two protocols:
- V1 — Purged K-Fold CV with
StatefulForexEnsemble(identical structure to the stateless pipeline, onlyshuffle=False) - V2 — Single 70/10/20 chronological split with no ensemble, no threshold optimisation
Decision thresholds for Short and Long classes (Hold is the default at 0) are optimised on OOF predictions via 2D grid search:
- Academic Global Optimum: maximises macro F1 —
(F1_Short + F1_Hold + F1_Long) / 3 - Financial Optimum: maximises expectancy —
reward × TP − risk × FP
| Model | PR-AUC ↑ | AUROC ↑ | Accuracy | F1 | Precision | Recall |
|---|---|---|---|---|---|---|
| GRU_2L | 0.5102 | 0.6874 | 49.34% | 48.72% | 48.70% | 49.34% |
| LSTM_Mid_2L | 0.4994 | 0.6754 | 48.34% | 47.87% | 47.53% | 48.34% |
| GRU_Mid_2L | 0.5104 | 0.6869 | 49.47% | 48.96% | 48.93% | 49.47% |
| Model | PR-AUC ↑ | AUROC ↑ | Accuracy | F1 | Precision | Recall |
|---|---|---|---|---|---|---|
| GRU_2L | 0.5351 | 0.6749 | 50.17% | 50.30% | 50.88% | 50.17% |
| LSTM_Mid_2L | 0.5257 | 0.6659 | 50.18% | 49.98% | 50.35% | 50.18% |
| GRU_Mid_2L | 0.5362 | 0.6754 | 50.92% | 50.93% | 51.24% | 50.92% |
The asymmetric regime yields higher PR-AUC across all models — the 1:2 risk/reward setup makes the classification task more discriminable. GRU_Mid_2L (64 hidden units, GRU cells) consistently outperforms GRU_2L and LSTM_Mid_2L in both regimes.
Stateful models perform near-identically to their stateless counterparts (ΔPR-AUC < 0.005), confirming that the H1 market is predominantly Markovian — carrying hidden state across 24-hour windows provides no measurable benefit.
| Regime | Best Model | Protocol | PR-AUC | Accuracy |
|---|---|---|---|---|
| Symmetric | GRU_Mid_2L | V1 (KFold) | 0.5107 | 49.96% |
| Symmetric | GRU_Mid_2L | V2 (70/10/20) | 0.5082 | 49.55% |
| Asymmetric | GRU_Mid_2L | V1 (KFold) | 0.5367 | 51.07% |
| Asymmetric | GRU_Mid_2L | V2 (70/10/20) | 0.5377 | 52.98% |
.
├── modules/ # Python backend (namespace package)
│ ├── config.py # YAML config loader
│ ├── data_processing.py # Data loading, features, TBM, PurgedKFold
│ ├── models.py # ForexRNN (LSTM/GRU), ForexEnsemble
│ ├── training.py # Train loop, Purged CV pipeline, test eval
│ ├── stateful_pipeline.py # Stateful RNN experiments (V1 KFold + V2 70/10/20)
│ ├── evaluation.py # Plots, threshold optim, KDE analysis
│ └── utils.py # Logging, seed, save/load vaults
├── docs/
│ └── report.md # Full technical report
├── config.yaml # All hyperparameters
├── 01_data_preparation.ipynb # Cyclical features (hour/dow sin/cos), pips
├── 02_training_pipeline.ipynb # Core pipeline: features, TBM, CV, test, figures
├── data/ # Source CSV (gitignored)
│ └── EURUSD_download_notebook.ipynb # M1→H1 download & conversion
├── results/ # *.pt experiment vaults (gitignored)
├── figures/ # Generated comparison plots (gitignored)
└── README.md
git clone <repo>
cd deep_learning_projectwork
uv syncRun the notebooks in order:
data/EURUSD_download_notebook.ipynb— downloads M1 tick data, converts to H1 candles01_data_preparation.ipynb— cyclical features (hour/dow sin/cos) and pips02_training_pipeline.ipynb— feature engineering, EWMA volatility, TBM labels, Purged CV (k=5, 3 architectures), ensemble evaluation, comparison figures
Edit config.yaml:
pt_multiplier: 4.6 # Take Profit in ATR multiples
sl_multiplier: 2.3 # Stop Loss in ATR multiplesThen re-run 02_training_pipeline.ipynb. If the prepared H1 data isn't already available, also re-run 01_data_preparation.ipynb.
- Logging is initialised once via
setup_global_logging("pipeline.log")— called in the notebooks torch.load(..., weights_only=False)required to load vaults (contain NumPy arrays inside dicts)- Data pipeline reads
config.get("normalization_window_fast", 120)— consistent withconfig.yaml
- Python 3.10+
- CUDA-capable GPU recommended
- LaTeX distribution with
lualatexandlatexmk(for presentation) - See
pyproject.tomlfor full dependency list (managed viauv)
- Triple-Barrier Method & Purged K-Fold CV: M. Prado, Advances in Financial Machine Learning (2018)
- Dataset: EUR/USD historical data (H1, 2002–2025), publicly available from HistData