Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
103 changes: 100 additions & 3 deletions Readme.md
Original file line number Diff line number Diff line change
@@ -1,8 +1,103 @@
> 🔗 **Live demo:** https://iambatman07-fraud-detection-mlops.hf.space · [HF Space](https://huggingface.co/spaces/IamBatman07/Fraud-Detection-MLOps)

# Fraud-Detection-MLOps

> 🔗 **Live demo:** https://iambatman07-fraud-detection-mlops.hf.space · [HF Space](https://huggingface.co/spaces/IamBatman07/Fraud-Detection-MLOps)
A fraud-detection system for payment transactions, built as a full MLOps pipeline rather than a notebook. It trains an XGBoost classifier on 6.1M transactions from two sources (PaySim and Sparkov), serves it behind a FastAPI endpoint, and keeps it healthy in production: a drift monitor watches eight features, and when the data shifts the system retrains, evaluates the candidate in shadow mode, and promotes it only if it beats the model currently live. Every model version is registered in MLflow and can be rolled back in milliseconds.

The point of the project is the operational layer — the split discipline, the registry, the drift loop — not the classifier itself.

---

## Architecture

![Architecture — training, registry, and the drift to retrain loop](assets/architecture.png)

---

## Measured results

All numbers below are read from committed artifacts in `results/`. Evaluation is on `sparkov_test.csv` — 555,719 transactions from a later time window that never enters training.

### The split fixes the score

The original training code used a random split on time-series data, which leaked future transactions into training. Fixing it to a per-source chronological split moved the number:

| Split | OOT AUC | |
|---|---:|---|
| Random stratified (as originally written) | 0.9210 | inflated by leakage |
| **Per-source temporal (final 20% held out)** | **0.7949** | honest baseline |

Source: [`results/baseline_metrics.json`](results/baseline_metrics.json)

Fraud detection on payment transactions (PaySim + Sparkov), built as an MLOps pipeline: a temporal train/test split, MLflow registry promotion and rollback, KS+PSI drift detection, auto-retrain-and-promote, Dask feature engineering, and Postgres-backed telemetry on a DVC pipeline.
### Tuning closes the gap to AutoGluon

| Model | OOT AUC | Δ vs AutoGluon |
|---|---:|---:|
| AutoGluon (published FDB baseline) | 0.9520 | — |
| **XGBoost + Optuna + source-balanced weights** | **0.95196** | **−0.00004** |

Source: [`results/day05/day05_leaderboard.csv`](results/day05/day05_leaderboard.csv) · weights: paysim 0.60×, sparkov 2.95×

### Each discipline layer is load-bearing

| Layer | OOT AUC | Δ |
|---|---:|---:|
| L0 — random split, XGBoost defaults | 0.5467 | — |
| L1 — + temporal split | 0.6574 | +0.1108 |
| L2 — + source-balanced weights | 0.6962 | +0.0388 |
| L3 — + Optuna tuning | **0.9480** | **+0.2518** |

**Total L0 → L3: +0.4013 AUC.** Source: [`results/day06/ablation_modelling.csv`](results/day06/ablation_modelling.csv)

### Specialised model vs a frontier LLM

Same 200-row out-of-time sample, scored three ways:

| Strategy | AUC | AUPRC | Latency/query | Cost @ 1k qps/day |
|---|---:|---:|---:|---:|
| **XGBoost champion** | **0.9156** | **0.5260** | 60 µs | **$0.43** |
| Naive notebook XGBoost | 0.6264 | 0.4153 | 73 µs | $0.43 |
| Claude Opus 4.6 as judge | 0.6225 | 0.3512 | 1.82 s | $1,250,691 |

The LLM is 30,400× slower and 2.9M× more expensive at equal throughput, and ranks worse than the tuned model. The naive XGBoost effectively ties the LLM — the 0.29 AUC gap comes from the discipline layers above, not from the algorithm. Source: [`results/day06/frontier_comparison.csv`](results/day06/frontier_comparison.csv)

### Operational metrics

| Capability | Measured | Source |
|---|---|---|
| Registry rollback | **3.9 ms** median alias flip · 11.9 ms fully audited | [`registry_rollback_times.csv`](results/registry_rollback_times.csv) |
| Drift detection | **precision 1.0, recall 1.0, 0-day lag** on a 2σ shift injected at day 23 of a 30-day replay | [`drift_replay_summary.json`](results/drift_replay_summary.json) |
| Auto-retrain | **median 6.85 s** detect → train → shadow-eval → promote · 3/3 events promoted | [`drift_retrain_events.csv`](results/drift_retrain_events.csv) |
| Dask vs pandas | bit-exact to **5.5e-12**; pandas 1.15M rows/s vs Dask 0.26M rows/s on one host | [`throughput_speedup.csv`](results/throughput_speedup.csv) |

Dask is slower on a single host by design — it is the scaling primitive, and being bit-exact with pandas is what lets you swap engines without changing results.

---

## How it works

1. **Ingest** — PaySim and Sparkov are pulled through a DVC pipeline so every stage is reproducible from a clean checkout.
2. **Features** — engineered identically in pandas or Dask; a determinism test asserts the two agree to within 5.5e-12.
3. **Split** — each source is sorted chronologically and the final 20% held out. A regression test fails if a random split reappears.
4. **Train** — XGBoost with source-balanced sample weights, tuned by a 30-trial Optuna sweep. Every run is logged to MLflow.
5. **Register** — the winning run is registered and promoted by moving the `production` alias. Rollback is the same operation in reverse, and both write an audit tag.
6. **Serve** — FastAPI resolves the `production` alias at startup and scores in ~60 µs.
7. **Monitor** — KS and PSI run over eight features per day. Two consecutive drift days trigger a retrain.
8. **Retrain → shadow → promote** — the candidate is scored on held-out data and promoted only if its AUPRC is within tolerance of the live model. If it loses, the live model stays.

## Infrastructure

| Layer | Technology |
|---|---|
| Pipeline | DVC |
| Training | XGBoost · Optuna |
| Tracking / registry | MLflow (SQLite locally, Postgres in compose) |
| Feature engineering | pandas · Dask |
| Serving | FastAPI · Uvicorn |
| Dashboard | Streamlit |
| Store / cache | Postgres · Redis |
| Packaging | Docker Compose |
| CI | GitHub Actions |

---

Expand Down Expand Up @@ -34,6 +129,8 @@ pytest tests/ -q
pytest tests/ -q -m "not requires_data" # CI mode (skips full-data replay)
```

Regenerate the architecture diagram with `python assets/make_architecture.py`.

---

## Datasets
Expand All @@ -44,7 +141,7 @@ pytest tests/ -q -m "not requires_data" # CI mode (skips full-data repl
| `sparkov.csv` | 1.30M | 2019-01 → 2020-06 | training window |
| `sparkov_test.csv` | 555,719 | 2020-06 → 2020-12 | held-out out-of-time |

DVC-tracked; `dvc pull` to materialize. The split is chronological per source — the final 20% of each is held out.
DVC-tracked; `dvc pull` to materialize.

---

Expand Down
Binary file added assets/architecture.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
167 changes: 167 additions & 0 deletions assets/make_architecture.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,167 @@
"""Render assets/architecture.png.

Drawn with Pillow at 2x and downsampled, so the PNG stays crisp on HiDPI
screens. Dark card background with light text — readable on both the GitHub
light and dark themes.

Run: python assets/make_architecture.py
"""
from pathlib import Path

from PIL import Image, ImageDraw, ImageFont

S = 2 # supersample factor
W, H = 940 * S, 560 * S
OUT = Path(__file__).with_name("architecture.png")

BG = (13, 17, 23)
FG = (201, 209, 217)
MUTED = (139, 148, 158)
LINE = (110, 118, 129)
ACCENT = (188, 140, 255)
GREEN = (63, 185, 80)

FONTS = r"C:\Windows\Fonts"


def font(name, size):
return ImageFont.truetype(f"{FONTS}\\{name}", size * S)


f_title = font("seguisb.ttf", 15)
f_head = font("seguisb.ttf", 13)
f_small = font("segoeui.ttf", 11)
f_lbl = font("segoeuii.ttf", 10)

img = Image.new("RGB", (W, H), BG)
d = ImageDraw.Draw(img)


def box(x, y, w, h, colour=LINE, width=2):
d.rounded_rectangle([x * S, y * S, (x + w) * S, (y + h) * S],
radius=6 * S, outline=colour, width=int(width * S))


def text(x, y, s, f=f_small, fill=MUTED, anchor="mm"):
d.text((x * S, y * S), s, font=f, fill=fill, anchor=anchor)


def arrow(pts, colour=LINE, width=1.6, dash=False):
pts = [(x * S, y * S) for x, y in pts]
for i in range(len(pts) - 1):
if dash:
_dashed(pts[i], pts[i + 1], colour, width)
else:
d.line([pts[i], pts[i + 1]], fill=colour, width=int(width * S))
_head(pts[-2], pts[-1], colour)


def _dashed(p0, p1, colour, width, on=6, off=4):
(x0, y0), (x1, y1) = p0, p1
dx, dy = x1 - x0, y1 - y0
dist = max((dx * dx + dy * dy) ** 0.5, 1e-6)
ux, uy = dx / dist, dy / dist
pos = 0.0
while pos < dist:
seg = min(on * S, dist - pos)
d.line([(x0 + ux * pos, y0 + uy * pos),
(x0 + ux * (pos + seg), y0 + uy * (pos + seg))],
fill=colour, width=int(width * S))
pos += (on + off) * S


def _head(p0, p1, colour, size=7):
(x0, y0), (x1, y1) = p0, p1
dx, dy = x1 - x0, y1 - y0
dist = max((dx * dx + dy * dy) ** 0.5, 1e-6)
ux, uy = dx / dist, dy / dist
px, py = -uy, ux
s = size * S
d.polygon([(x1, y1),
(x1 - ux * s + px * s * 0.5, y1 - uy * s + py * s * 0.5),
(x1 - ux * s - px * s * 0.5, y1 - uy * s - py * s * 0.5)],
fill=colour)


text(470, 26, "Fraud-Detection-MLOps — training, registry, and the drift → retrain loop",
f_title, FG)

# ── data sources ────────────────────────────────────────────────────────────
box(24, 52, 170, 66)
text(109, 74, "PaySim", f_head, FG)
text(109, 92, "6.36M rows · simulated")
text(109, 108, "training")

box(212, 52, 170, 66)
text(297, 74, "Sparkov", f_head, FG)
text(297, 92, "1.30M · 2019-01→2020-06")
text(297, 108, "training window")

box(400, 52, 188, 66, GREEN)
text(494, 74, "sparkov_test", f_head, FG)
text(494, 92, "555,719 · 2020-06→2020-12")
text(494, 108, "held-out out-of-time")

# ── features ────────────────────────────────────────────────────────────────
arrow([(109, 118), (109, 150)])
arrow([(297, 118), (297, 150)])
box(24, 152, 358, 72)
text(203, 176, "Feature engineering — Dask / pandas", f_head, FG)
text(203, 195, "bit-exact across engines (max abs diff 5.5e-12)")
text(203, 212, "pandas 1.15M rows/s · Dask 0.26M rows/s on one host")

# ── temporal split ──────────────────────────────────────────────────────────
arrow([(382, 188), (410, 188)])
box(412, 152, 176, 72, ACCENT)
text(500, 176, "Temporal split", f_head, FG)
text(500, 195, "per source, chronological")
text(500, 212, "final 20% held out")

# ── train ───────────────────────────────────────────────────────────────────
arrow([(500, 224), (500, 248)])
box(330, 250, 340, 72)
text(500, 274, "Train — XGBoost + Optuna (30 trials)", f_head, FG)
text(500, 293, "source-balanced: paysim 0.60× · sparkov 2.95×")
text(500, 310, "OOT AUC 0.9520")

# ── registry ────────────────────────────────────────────────────────────────
arrow([(500, 322), (500, 350)])
box(330, 352, 340, 72, ACCENT)
text(500, 376, "MLflow Model Registry", f_head, FG)
text(500, 395, "promote / rollback via alias · audited")
text(500, 412, "alias flip 3.9 ms · audited rollback 11.9 ms")

# ── serving ─────────────────────────────────────────────────────────────────
arrow([(670, 388), (714, 388)])
box(718, 352, 198, 72)
text(817, 376, "FastAPI serving :8000", f_head, FG)
text(817, 395, "alias “production”")
text(817, 412, "60 µs/query · $0.43 per 1k qps/day")

# ── drift monitor ───────────────────────────────────────────────────────────
arrow([(817, 424), (817, 450)])
box(600, 454, 316, 72, ACCENT)
text(758, 477, "Drift monitor — KS + PSI", f_head, FG)
text(758, 496, "8 features · precision 1.0 · recall 1.0")
text(758, 512, "0-day detection lag on a 2σ shift")

# ── retrain loop ────────────────────────────────────────────────────────────
arrow([(600, 490), (484, 490)])
box(164, 454, 316, 72)
text(322, 477, "Auto-retrain → shadow eval → promote", f_head, FG)
text(322, 496, "promote only if shadow AUPRC ≥ prod − 0.010")
text(322, 512, "median 6.85 s end-to-end · 3/3 promoted")

arrow([(164, 490), (96, 490), (96, 388), (328, 388)], ACCENT, dash=True)
text(150, 440, "registers new version", f_lbl, MUTED, anchor="lm")

# held-out set is only ever scored against
arrow([(588, 85), (700, 85), (700, 286), (674, 286)], GREEN, dash=True)
text(710, 178, "scored against only,", f_lbl, GREEN, anchor="lm")
text(710, 194, "never trained on", f_lbl, GREEN, anchor="lm")

text(24, 544, "DVC orchestrates every stage · MLflow records every run",
f_lbl, MUTED, anchor="lm")

img.resize((W // S, H // S), Image.LANCZOS).save(OUT, "PNG", optimize=True)
print(f"wrote {OUT} ({OUT.stat().st_size // 1024} KB)")
67 changes: 67 additions & 0 deletions results/baseline_metrics.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,67 @@
{
"captured_on": "2026-05-18",
"day": 1,
"phase": "1 — audit + temporal-split fix + baseline",
"story": "Pre-fix benchmark AUC of 0.921 was inflated by two leakage paths: (1) train_test_split with stratify=y in src/train.py was a RANDOM split that mixed future txns into training; (2) benchmark_fdb fell back to data/raw/sparkov.csv which is the SAME file used for training. After fixing both — per-source temporal split in train.py + held-out sparkov_test.csv in benchmark_fdb — the honest sparkov AUC vs AutoGluon is 0.795, not 0.921.",
"pre_fix": {
"split_strategy": "random_stratified",
"benchmark_data_source": "data/raw/sparkov.csv (same as train set)",
"reported_sparkov_benchmark_auc": 0.9210076846056613,
"delta_vs_AutoGluon_baseline_0.952": -0.030992,
"honest_assessment": "INVALID — both random split and training-set test → leakage stacked"
},
"post_fix_combined_temporal_test": {
"split_strategy": "temporal_per_source (last 20% of each source by txn_timestamp)",
"test_rows": 1531859,
"test_fraud_rate": 0.00378,
"test_accuracy": 0.996564,
"test_auc": 0.998906,
"test_average_precision": 0.827993,
"test_precision": 0.5268,
"test_recall": 0.8974,
"test_f1": 0.6639,
"note": "Inflated by paysim's near-deterministic balance_change_orig signal (paysim is 83% of combined data)."
},
"post_fix_per_source_train_time_test": {
"source_paysim": {
"rows": 1272524,
"fraud_rate": 0.003343,
"auc": 0.999650,
"average_precision": 0.917024
},
"source_sparkov": {
"rows": 259335,
"fraud_rate": 0.005931,
"auc": 0.996608,
"average_precision": 0.809023,
"note": "Held-out window is the final 20% of sparkov.csv (~Mar–Jun 2020). Same data-collection regime as training so the model interpolates well."
}
},
"post_fix_benchmark_held_out_sparkov_test_file": {
"test_file": "data/raw/sparkov_test.csv",
"date_range": "2020-06-21 12:14:25 → 2020-12-31 23:59:34",
"rows": 555719,
"our_auc": 0.794896,
"baselines": {
"AutoGluon": {"baseline_auc": 0.952, "delta": -0.157104, "status": "BELOW"},
"H2O AutoML": {"baseline_auc": 0.947, "delta": -0.152104, "status": "BELOW"},
"AutoSklearn": {"baseline_auc": 0.931, "delta": -0.136104, "status": "BELOW"}
},
"headline": "0.795 — this is Sentinel's honest sparkov AUC for the rest of the sprint. The 0.921 claim retires today.",
"interpretation": "Held-out file covers Jun–Dec 2020, a time period never seen in any form during training. The 0.20 AP/0.20 AUC drop vs the train-time test slice quantifies real distribution shift — and that's the gap auto-retrain (Day 3) must close."
},
"mlflow_run": {
"experiment": "sentinel-day01-temporal-split",
"run_name": "day01_phase1_temporal_split",
"tracking_uri": "sqlite:///mlflow.db (local)",
"logged_params": ["split_strategy", "test_size", "random_state", "n_estimators", "max_depth", "learning_rate", "scale_pos_weight", "use_smote", "early_stopping_rounds", "train_rows", "test_rows", "train_fraud_rate", "test_fraud_rate"],
"logged_metrics": ["test_accuracy", "test_auc", "test_average_precision"],
"logged_artifacts": ["xgboost_model (mlflow.xgboost)", "models/fraud_model.pkl"]
},
"files_changed_today": [
"src/combine_datasets.py — added txn_timestamp column, replaced shuffle with chronological sort",
"src/preprocess.py — added txn_timestamp to required schema + passthrough",
"src/train.py — replaced train_test_split with temporal_split_per_source; wrapped run in mlflow.start_run()",
"src/benchmark_fdb.py — prefer held-out sparkov_test.csv over leaking sparkov.csv fallback"
]
}
6 changes: 6 additions & 0 deletions results/day05/day05_leaderboard.csv
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
rank,strategy,training_scope,sample_weight,threshold,in_dist_sparkov_auc,oot_sparkov_test_auc,oot_ap,oot_recall,oot_precision,oot_f1,delta_vs_autogluon,notes
1,AutoGluon (FDB published baseline),FDB standard,n/a,n/a,,0.952,,,,,0.0,reference baseline
2,Day-5 Optuna + source-balanced weights -threshold=0.5,"full (6.1M, paysim 0.60x / sparkov 2.95x)",source_balanced,0.5,0.9971775606973103,0.951958845551572,0.20830353279765584,0.2592074592074592,0.2744323790720632,0.2666027331575162,-0.0,AUC ties AutoGluon; recall jumps 6x
3,Day-5 Optuna + source-balanced weights -threshold=tau*,"full (6.1M, paysim 0.60x / sparkov 2.95x)",source_balanced,0.8939489722251892,0.9971775606973103,0.951958845551572,0.20830353279765584,0.14685314685314685,0.5526315789473685,0.23204419889502761,-0.0,"F1 trade: higher precision (0.55), lower recall (0.15)"
4,Day-5 Optuna best (30 trials) -full retrain,full (6.1M),uniform,0.5,0.9970274251784879,0.915380324555611,0.067267067319249,0.042890442890442894,0.34328358208955223,0.0762536261914629,-0.0366,tuning recovers ranking; recall still collapses at 0.5
5,"Day-1 honest baseline (random XGB defaults, temporal split fix)",full (6.1M),uniform,0.5,0.996608,0.794896,,,,,-0.1571,honest baseline -pre-tuning
5 changes: 5 additions & 0 deletions results/day06/ablation_modelling.csv
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
view,layer,description,oot_auc,oot_auprc,oot_recall_at_0_5,oot_precision_at_0_5,delta_auc_vs_prev_layer,train_seconds,n_train_rows
modelling,L0_naive_random_default,"Naive notebook: random split, XGB defaults, no balancing",0.5466625224585316,0.09136609725819243,0.08158508158508158,0.5271084337349398,,19.57696139998734,480000
modelling,L1_temporal_default,+temporal split (Day-1 honest baseline fix),0.6574214679548385,0.16115083339577455,0.068997668997669,0.6166666666666667,0.11075894549630683,16.252756399917416,479999
modelling,L2_temporal_source_balanced_default,+source-balanced sample weights,0.696180215592977,0.1608393025972209,0.1076923076923077,0.7751677852348994,0.03875874763813858,17.605792300077155,479999
modelling,L3_temporal_source_balanced_optuna,+Optuna tuning (Day-5 champion),0.9479773886870319,0.23198044768793025,0.15897435897435896,0.41585365853658535,0.2517971730940548,45.4373118999647,479999
Loading
Loading