Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
36 commits
Select commit Hold shift + click to select a range
5609109
Remove .claude/launch.json (generated artifact)
Mark007-R Aug 5, 2026
43c0370
Remove docs/DATA_SPLIT.md (generated artifact)
Mark007-R Aug 5, 2026
d0256c7
Remove docs/MLOPS_AUDIT.md (generated artifact)
Mark007-R Aug 5, 2026
11c48f7
Remove results/baseline_metrics.json (generated artifact)
Mark007-R Aug 5, 2026
57ee4e5
Remove results/day04_api_smoke.json (generated artifact)
Mark007-R Aug 5, 2026
71cd901
Remove results/day04_shadow_smoke.json (generated artifact)
Mark007-R Aug 5, 2026
02421af
Remove results/day05/day05_leaderboard.csv (generated artifact)
Mark007-R Aug 5, 2026
b1791ed
Remove results/day05/failure_modes.csv (generated artifact)
Mark007-R Aug 5, 2026
663488b
Remove results/day05/failure_modes_summary.json (generated artifact)
Mark007-R Aug 5, 2026
90c9557
Remove results/day05/optuna_best_params.json (generated artifact)
Mark007-R Aug 5, 2026
069d808
Remove results/day05/optuna_trials.csv (generated artifact)
Mark007-R Aug 5, 2026
2dd49cb
Remove results/day05/targeted_fix_eval.json (generated artifact)
Mark007-R Aug 5, 2026
69aebbb
Remove results/day05/tuned_eval.json (generated artifact)
Mark007-R Aug 5, 2026
b7a089e
Remove results/day05/tuning_partition_info.json (generated artifact)
Mark007-R Aug 5, 2026
c4009c4
Remove results/day06/ablation.csv (generated artifact)
Mark007-R Aug 5, 2026
2f7508c
Remove results/day06/ablation_mlops_capability.csv (generated artifact)
Mark007-R Aug 5, 2026
39abb2e
Remove results/day06/ablation_modelling.csv (generated artifact)
Mark007-R Aug 5, 2026
5ad6b6a
Remove results/day06/ablation_summary.json (generated artifact)
Mark007-R Aug 5, 2026
f69d88e
Remove results/day06/frontier_comparison.csv (generated artifact)
Mark007-R Aug 5, 2026
c19b998
Remove results/day06/frontier_comparison.json (generated artifact)
Mark007-R Aug 5, 2026
fda2216
Remove results/day06/llm_fraud_negative_result.csv (generated artifact)
Mark007-R Aug 5, 2026
eb24e3e
Remove results/day06/llm_predictions.csv (generated artifact)
Mark007-R Aug 5, 2026
087d521
Remove results/day06/llm_summary.json (generated artifact)
Mark007-R Aug 5, 2026
f372728
Remove results/drift_metrics.json (generated artifact)
Mark007-R Aug 5, 2026
e7cafd8
Remove results/drift_reference.json (generated artifact)
Mark007-R Aug 5, 2026
357eef1
Remove results/drift_replay_per_day.csv (generated artifact)
Mark007-R Aug 5, 2026
3f72da7
Remove results/drift_replay_summary.json (generated artifact)
Mark007-R Aug 5, 2026
2a364ae
Remove results/drift_retrain_events.csv (generated artifact)
Mark007-R Aug 5, 2026
ce03bd3
Remove results/drift_retrain_metrics.json (generated artifact)
Mark007-R Aug 5, 2026
51b9988
Remove results/per_source_test_metrics.json (generated artifact)
Mark007-R Aug 5, 2026
3cc8ba4
Remove results/phase2_leaderboard.csv (generated artifact)
Mark007-R Aug 5, 2026
c69a4b4
Remove results/registry_metrics.json (generated artifact)
Mark007-R Aug 5, 2026
4e2a3c7
Remove results/registry_rollback_times.csv (generated artifact)
Mark007-R Aug 5, 2026
e18df23
Remove results/throughput_metrics.json (generated artifact)
Mark007-R Aug 5, 2026
861172c
Remove results/throughput_speedup.csv (generated artifact)
Mark007-R Aug 5, 2026
9598d0a
Update README for removed docs/ and results/ trees
Mark007-R Aug 5, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 0 additions & 11 deletions .claude/launch.json

This file was deleted.

31 changes: 14 additions & 17 deletions Readme.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,38 +10,38 @@ Fraud-Detection-MLOps detects fraud in payment transactions (PaySim + Sparkov),

## The headline: the leakage fix is the story

The original `train.py` used `train_test_split(..., stratify=y)` — a **random** split on time-series fraud data, which leaks future transaction patterns into training. The Day-1 audit ([docs/MLOPS_AUDIT.md](docs/MLOPS_AUDIT.md)) found a second leak too: the benchmark fell back to the *training* file instead of the held-out test file.
The original `train.py` used `train_test_split(..., stratify=y)` — a **random** split on time-series fraud data, which leaks future transaction patterns into training. The Day-1 audit found a second leak too: the benchmark fell back to the *training* file instead of the held-out test file.

| Metric (Sparkov OOT, vs AutoGluon 0.952) | Before fix | After fix |
|---|---:|---:|
| **Honest AUC** | **0.9210** (leaked) | **0.7949** (honest baseline) |
| Apparent gap to AutoGluon | −0.031 (mirage) | −0.157 (real) |

The fix: `temporal_split_per_source()` sorts each source chronologically and takes the final 20% as test ([docs/DATA_SPLIT.md](docs/DATA_SPLIT.md)). **The ~0.126 AUC drop was the data leakage. The 0.7949 is the number that was always real** — and the rest of the sprint earns it back honestly.
The fix: `temporal_split_per_source()` sorts each source chronologically and takes the final 20% as test. **The ~0.126 AUC drop was the data leakage. The 0.7949 is the number that was always real** — and the rest of the sprint earns it back honestly.

---

## Final champion + the gap closed honestly

A 30-trial Optuna sweep + source-balanced sample weights closed the entire gap to AutoGluon — **no new features, no ensembling, no SHAP** (that's the joint Fraud-Detection project's territory; Fraud-Detection-MLOps is deliberately MLOps-only).

| Model | OOT AUC (sparkov_test.csv) | Δ vs AutoGluon 0.952 | Source |
|---|---:|---:|---|
| Day-1 honest baseline (temporal split) | 0.7949 | −0.157 | `results/baseline_metrics.json` |
| **Day-5 champion (Optuna + source-balanced)** | **0.9520** | **−0.00004** (tied) | `results/day05/day05_leaderboard.csv` |
| Model | OOT AUC (sparkov_test.csv) | Δ vs AutoGluon 0.952 |
|---|---:|---:|
| Day-1 honest baseline (temporal split) | 0.7949 | −0.157 |
| **Day-5 champion (Optuna + source-balanced)** | **0.9520** | **−0.00004** (tied) |

Champion: XGBoost, per-source temporal split + source-balanced weights (paysim 0.60×, sparkov 2.95×) + Optuna tuning. Stored at `models/fraud_model_tuned_fixed.pkl` (MLflow run `day05_targeted_fix_v1`).

---

## The four MLOps capabilities (what a notebook doesn't have)

| Capability | Headline metric | Source |
|---|---|---|
| **Dask distributed features** | Pandas 1.15M rows/s vs Dask 0.26M rows/s on one host, **bit-exact within 5.5e-12** | `results/throughput_speedup.csv` |
| **MLflow registry rollback** | **3.9 ms** median alias flip; 11.9 ms full audited rollback | `results/registry_rollback_times.csv` |
| **KS+PSI drift detection** | **0-day** detection lag on a synthetic 2σ shift; **precision 1.0, recall 1.0** on the 7-day window | `results/drift_replay_summary.json` |
| **Auto-retrain + shadow-promote** | drift → train → shadow-eval → promote in **median 6.85 s**; 3/3 events auto-promoted | `results/drift_retrain_events.csv` |
| Capability | Headline metric |
|---|---|
| **Dask distributed features** | Pandas 1.15M rows/s vs Dask 0.26M rows/s on one host, **bit-exact within 5.5e-12** |
| **MLflow registry rollback** | **3.9 ms** median alias flip; 11.9 ms full audited rollback |
| **KS+PSI drift detection** | **0-day** detection lag on a synthetic 2σ shift; **precision 1.0, recall 1.0** on the 7-day window |
| **Auto-retrain + shadow-promote** | drift → train → shadow-eval → promote in **median 6.85 s**; 3/3 events auto-promoted |

Dask "loses" the single-host throughput race but is the *scaling primitive* — and it's bit-exact with Pandas, which is the property that lets you swap engines without changing results.

Expand All @@ -57,7 +57,7 @@ Day 6 ran Claude Opus 4.6 as an LLM fraud judge on the same 200-row OOT sample (
| Naive notebook XGBoost | 0.626 | 0.415 | 73 µs | $0.43 |
| Claude Opus 4.6 LLM-judged | 0.622 | 0.351 | 1.82 s | **$1,250,691** |

Two findings: (1) the LLM is **30,400× slower** and **2.9M× more expensive**, and ranks 0.29 AUC worse — it caught only textbook patterns (large online txns at night) and missed the "card skimmed at POS → grocery abuse" behavioral class. (2) **The naive notebook XGBoost (0.626) ties the LLM (0.622)** — the 0.29-AUC jump to the champion comes from the *discipline layers*, not from "XGBoost vs LLM". The discipline is the model. Source: `results/day06/frontier_comparison.csv`. The LLM judge ran in simulate mode (no API key on host); `--mode api` is the same code path with a real key.
Two findings: (1) the LLM is **30,400× slower** and **2.9M× more expensive**, and ranks 0.29 AUC worse — it caught only textbook patterns (large online txns at night) and missed the "card skimmed at POS → grocery abuse" behavioral class. (2) **The naive notebook XGBoost (0.626) ties the LLM (0.622)** — the 0.29-AUC jump to the champion comes from the *discipline layers*, not from "XGBoost vs LLM". The discipline is the model. The LLM judge ran in simulate mode (no API key on host); `--mode api` is the same code path with a real key.

### MLOps ablation (L0 → L3, full OOT)
| Layer | OOT AUC | Δ vs prev |
Expand All @@ -67,7 +67,7 @@ Two findings: (1) the LLM is **30,400× slower** and **2.9M× more expensive**,
| L2 + source-balanced weights | 0.696 | +0.039 |
| L3 + Optuna (champion) | **0.948** | **+0.252** |

Total L0→L3 gain: **+0.401 OOT AUC**. Every layer is load-bearing; Optuna is the biggest single contributor. Source: `results/day06/ablation_modelling.csv`.
Total L0→L3 gain: **+0.401 OOT AUC**. Every layer is load-bearing; Optuna is the biggest single contributor.

---

Expand Down Expand Up @@ -159,9 +159,6 @@ Fraud-Detection-MLOps/
│ └── frontier/{llm_judge,compare_models,ablation}.py
├── tests/ # 31-test suite
├── models/fraud_model_tuned_fixed.pkl # champion
├── results/ # baseline, throughput, rollback, drift, day05/, day06/
├── reports/ # day01..day07 phase reports
├── docs/ # MLOPS_AUDIT.md, DATA_SPLIT.md
└── data/raw/{paysim,sparkov,sparkov_test}.csv
```

Expand Down
76 changes: 0 additions & 76 deletions docs/DATA_SPLIT.md

This file was deleted.

116 changes: 0 additions & 116 deletions docs/MLOPS_AUDIT.md

This file was deleted.

Loading
Loading