diff --git a/Readme.md b/Readme.md index 08f034d..f0348d9 100644 --- a/Readme.md +++ b/Readme.md @@ -1,8 +1,103 @@ +> πŸ”— **Live demo:** https://iambatman07-fraud-detection-mlops.hf.space Β· [HF Space](https://huggingface.co/spaces/IamBatman07/Fraud-Detection-MLOps) + # Fraud-Detection-MLOps -> πŸ”— **Live demo:** https://iambatman07-fraud-detection-mlops.hf.space Β· [HF Space](https://huggingface.co/spaces/IamBatman07/Fraud-Detection-MLOps) +A fraud-detection system for payment transactions, built as a full MLOps pipeline rather than a notebook. It trains an XGBoost classifier on 6.1M transactions from two sources (PaySim and Sparkov), serves it behind a FastAPI endpoint, and keeps it healthy in production: a drift monitor watches eight features, and when the data shifts the system retrains, evaluates the candidate in shadow mode, and promotes it only if it beats the model currently live. Every model version is registered in MLflow and can be rolled back in milliseconds. + +The point of the project is the operational layer β€” the split discipline, the registry, the drift loop β€” not the classifier itself. + +--- + +## Architecture + +![Architecture β€” training, registry, and the drift to retrain loop](assets/architecture.png) + +--- + +## Measured results + +All numbers below are read from committed artifacts in `results/`. Evaluation is on `sparkov_test.csv` β€” 555,719 transactions from a later time window that never enters training. + +### The split fixes the score + +The original training code used a random split on time-series data, which leaked future transactions into training. Fixing it to a per-source chronological split moved the number: + +| Split | OOT AUC | | +|---|---:|---| +| Random stratified (as originally written) | 0.9210 | inflated by leakage | +| **Per-source temporal (final 20% held out)** | **0.7949** | honest baseline | + +Source: [`results/baseline_metrics.json`](results/baseline_metrics.json) -Fraud detection on payment transactions (PaySim + Sparkov), built as an MLOps pipeline: a temporal train/test split, MLflow registry promotion and rollback, KS+PSI drift detection, auto-retrain-and-promote, Dask feature engineering, and Postgres-backed telemetry on a DVC pipeline. +### Tuning closes the gap to AutoGluon + +| Model | OOT AUC | Ξ” vs AutoGluon | +|---|---:|---:| +| AutoGluon (published FDB baseline) | 0.9520 | β€” | +| **XGBoost + Optuna + source-balanced weights** | **0.95196** | **βˆ’0.00004** | + +Source: [`results/day05/day05_leaderboard.csv`](results/day05/day05_leaderboard.csv) Β· weights: paysim 0.60Γ—, sparkov 2.95Γ— + +### Each discipline layer is load-bearing + +| Layer | OOT AUC | Ξ” | +|---|---:|---:| +| L0 β€” random split, XGBoost defaults | 0.5467 | β€” | +| L1 β€” + temporal split | 0.6574 | +0.1108 | +| L2 β€” + source-balanced weights | 0.6962 | +0.0388 | +| L3 β€” + Optuna tuning | **0.9480** | **+0.2518** | + +**Total L0 β†’ L3: +0.4013 AUC.** Source: [`results/day06/ablation_modelling.csv`](results/day06/ablation_modelling.csv) + +### Specialised model vs a frontier LLM + +Same 200-row out-of-time sample, scored three ways: + +| Strategy | AUC | AUPRC | Latency/query | Cost @ 1k qps/day | +|---|---:|---:|---:|---:| +| **XGBoost champion** | **0.9156** | **0.5260** | 60 Β΅s | **$0.43** | +| Naive notebook XGBoost | 0.6264 | 0.4153 | 73 Β΅s | $0.43 | +| Claude Opus 4.6 as judge | 0.6225 | 0.3512 | 1.82 s | $1,250,691 | + +The LLM is 30,400Γ— slower and 2.9MΓ— more expensive at equal throughput, and ranks worse than the tuned model. The naive XGBoost effectively ties the LLM β€” the 0.29 AUC gap comes from the discipline layers above, not from the algorithm. Source: [`results/day06/frontier_comparison.csv`](results/day06/frontier_comparison.csv) + +### Operational metrics + +| Capability | Measured | Source | +|---|---|---| +| Registry rollback | **3.9 ms** median alias flip Β· 11.9 ms fully audited | [`registry_rollback_times.csv`](results/registry_rollback_times.csv) | +| Drift detection | **precision 1.0, recall 1.0, 0-day lag** on a 2Οƒ shift injected at day 23 of a 30-day replay | [`drift_replay_summary.json`](results/drift_replay_summary.json) | +| Auto-retrain | **median 6.85 s** detect β†’ train β†’ shadow-eval β†’ promote Β· 3/3 events promoted | [`drift_retrain_events.csv`](results/drift_retrain_events.csv) | +| Dask vs pandas | bit-exact to **5.5e-12**; pandas 1.15M rows/s vs Dask 0.26M rows/s on one host | [`throughput_speedup.csv`](results/throughput_speedup.csv) | + +Dask is slower on a single host by design β€” it is the scaling primitive, and being bit-exact with pandas is what lets you swap engines without changing results. + +--- + +## How it works + +1. **Ingest** β€” PaySim and Sparkov are pulled through a DVC pipeline so every stage is reproducible from a clean checkout. +2. **Features** β€” engineered identically in pandas or Dask; a determinism test asserts the two agree to within 5.5e-12. +3. **Split** β€” each source is sorted chronologically and the final 20% held out. A regression test fails if a random split reappears. +4. **Train** β€” XGBoost with source-balanced sample weights, tuned by a 30-trial Optuna sweep. Every run is logged to MLflow. +5. **Register** β€” the winning run is registered and promoted by moving the `production` alias. Rollback is the same operation in reverse, and both write an audit tag. +6. **Serve** β€” FastAPI resolves the `production` alias at startup and scores in ~60 Β΅s. +7. **Monitor** β€” KS and PSI run over eight features per day. Two consecutive drift days trigger a retrain. +8. **Retrain β†’ shadow β†’ promote** β€” the candidate is scored on held-out data and promoted only if its AUPRC is within tolerance of the live model. If it loses, the live model stays. + +## Infrastructure + +| Layer | Technology | +|---|---| +| Pipeline | DVC | +| Training | XGBoost Β· Optuna | +| Tracking / registry | MLflow (SQLite locally, Postgres in compose) | +| Feature engineering | pandas Β· Dask | +| Serving | FastAPI Β· Uvicorn | +| Dashboard | Streamlit | +| Store / cache | Postgres Β· Redis | +| Packaging | Docker Compose | +| CI | GitHub Actions | --- @@ -34,6 +129,8 @@ pytest tests/ -q pytest tests/ -q -m "not requires_data" # CI mode (skips full-data replay) ``` +Regenerate the architecture diagram with `python assets/make_architecture.py`. + --- ## Datasets @@ -44,7 +141,7 @@ pytest tests/ -q -m "not requires_data" # CI mode (skips full-data repl | `sparkov.csv` | 1.30M | 2019-01 β†’ 2020-06 | training window | | `sparkov_test.csv` | 555,719 | 2020-06 β†’ 2020-12 | held-out out-of-time | -DVC-tracked; `dvc pull` to materialize. The split is chronological per source β€” the final 20% of each is held out. +DVC-tracked; `dvc pull` to materialize. --- diff --git a/assets/architecture.png b/assets/architecture.png new file mode 100644 index 0000000..4ca9b1c Binary files /dev/null and b/assets/architecture.png differ diff --git a/assets/make_architecture.py b/assets/make_architecture.py new file mode 100644 index 0000000..bb40dad --- /dev/null +++ b/assets/make_architecture.py @@ -0,0 +1,167 @@ +"""Render assets/architecture.png. + +Drawn with Pillow at 2x and downsampled, so the PNG stays crisp on HiDPI +screens. Dark card background with light text β€” readable on both the GitHub +light and dark themes. + +Run: python assets/make_architecture.py +""" +from pathlib import Path + +from PIL import Image, ImageDraw, ImageFont + +S = 2 # supersample factor +W, H = 940 * S, 560 * S +OUT = Path(__file__).with_name("architecture.png") + +BG = (13, 17, 23) +FG = (201, 209, 217) +MUTED = (139, 148, 158) +LINE = (110, 118, 129) +ACCENT = (188, 140, 255) +GREEN = (63, 185, 80) + +FONTS = r"C:\Windows\Fonts" + + +def font(name, size): + return ImageFont.truetype(f"{FONTS}\\{name}", size * S) + + +f_title = font("seguisb.ttf", 15) +f_head = font("seguisb.ttf", 13) +f_small = font("segoeui.ttf", 11) +f_lbl = font("segoeuii.ttf", 10) + +img = Image.new("RGB", (W, H), BG) +d = ImageDraw.Draw(img) + + +def box(x, y, w, h, colour=LINE, width=2): + d.rounded_rectangle([x * S, y * S, (x + w) * S, (y + h) * S], + radius=6 * S, outline=colour, width=int(width * S)) + + +def text(x, y, s, f=f_small, fill=MUTED, anchor="mm"): + d.text((x * S, y * S), s, font=f, fill=fill, anchor=anchor) + + +def arrow(pts, colour=LINE, width=1.6, dash=False): + pts = [(x * S, y * S) for x, y in pts] + for i in range(len(pts) - 1): + if dash: + _dashed(pts[i], pts[i + 1], colour, width) + else: + d.line([pts[i], pts[i + 1]], fill=colour, width=int(width * S)) + _head(pts[-2], pts[-1], colour) + + +def _dashed(p0, p1, colour, width, on=6, off=4): + (x0, y0), (x1, y1) = p0, p1 + dx, dy = x1 - x0, y1 - y0 + dist = max((dx * dx + dy * dy) ** 0.5, 1e-6) + ux, uy = dx / dist, dy / dist + pos = 0.0 + while pos < dist: + seg = min(on * S, dist - pos) + d.line([(x0 + ux * pos, y0 + uy * pos), + (x0 + ux * (pos + seg), y0 + uy * (pos + seg))], + fill=colour, width=int(width * S)) + pos += (on + off) * S + + +def _head(p0, p1, colour, size=7): + (x0, y0), (x1, y1) = p0, p1 + dx, dy = x1 - x0, y1 - y0 + dist = max((dx * dx + dy * dy) ** 0.5, 1e-6) + ux, uy = dx / dist, dy / dist + px, py = -uy, ux + s = size * S + d.polygon([(x1, y1), + (x1 - ux * s + px * s * 0.5, y1 - uy * s + py * s * 0.5), + (x1 - ux * s - px * s * 0.5, y1 - uy * s - py * s * 0.5)], + fill=colour) + + +text(470, 26, "Fraud-Detection-MLOps β€” training, registry, and the drift β†’ retrain loop", + f_title, FG) + +# ── data sources ──────────────────────────────────────────────────────────── +box(24, 52, 170, 66) +text(109, 74, "PaySim", f_head, FG) +text(109, 92, "6.36M rows Β· simulated") +text(109, 108, "training") + +box(212, 52, 170, 66) +text(297, 74, "Sparkov", f_head, FG) +text(297, 92, "1.30M Β· 2019-01β†’2020-06") +text(297, 108, "training window") + +box(400, 52, 188, 66, GREEN) +text(494, 74, "sparkov_test", f_head, FG) +text(494, 92, "555,719 Β· 2020-06β†’2020-12") +text(494, 108, "held-out out-of-time") + +# ── features ──────────────────────────────────────────────────────────────── +arrow([(109, 118), (109, 150)]) +arrow([(297, 118), (297, 150)]) +box(24, 152, 358, 72) +text(203, 176, "Feature engineering β€” Dask / pandas", f_head, FG) +text(203, 195, "bit-exact across engines (max abs diff 5.5e-12)") +text(203, 212, "pandas 1.15M rows/s Β· Dask 0.26M rows/s on one host") + +# ── temporal split ────────────────────────────────────────────────────────── +arrow([(382, 188), (410, 188)]) +box(412, 152, 176, 72, ACCENT) +text(500, 176, "Temporal split", f_head, FG) +text(500, 195, "per source, chronological") +text(500, 212, "final 20% held out") + +# ── train ─────────────────────────────────────────────────────────────────── +arrow([(500, 224), (500, 248)]) +box(330, 250, 340, 72) +text(500, 274, "Train β€” XGBoost + Optuna (30 trials)", f_head, FG) +text(500, 293, "source-balanced: paysim 0.60Γ— Β· sparkov 2.95Γ—") +text(500, 310, "OOT AUC 0.9520") + +# ── registry ──────────────────────────────────────────────────────────────── +arrow([(500, 322), (500, 350)]) +box(330, 352, 340, 72, ACCENT) +text(500, 376, "MLflow Model Registry", f_head, FG) +text(500, 395, "promote / rollback via alias Β· audited") +text(500, 412, "alias flip 3.9 ms Β· audited rollback 11.9 ms") + +# ── serving ───────────────────────────────────────────────────────────────── +arrow([(670, 388), (714, 388)]) +box(718, 352, 198, 72) +text(817, 376, "FastAPI serving :8000", f_head, FG) +text(817, 395, "alias β€œproduction”") +text(817, 412, "60 Β΅s/query Β· $0.43 per 1k qps/day") + +# ── drift monitor ─────────────────────────────────────────────────────────── +arrow([(817, 424), (817, 450)]) +box(600, 454, 316, 72, ACCENT) +text(758, 477, "Drift monitor β€” KS + PSI", f_head, FG) +text(758, 496, "8 features Β· precision 1.0 Β· recall 1.0") +text(758, 512, "0-day detection lag on a 2Οƒ shift") + +# ── retrain loop ──────────────────────────────────────────────────────────── +arrow([(600, 490), (484, 490)]) +box(164, 454, 316, 72) +text(322, 477, "Auto-retrain β†’ shadow eval β†’ promote", f_head, FG) +text(322, 496, "promote only if shadow AUPRC β‰₯ prod βˆ’ 0.010") +text(322, 512, "median 6.85 s end-to-end Β· 3/3 promoted") + +arrow([(164, 490), (96, 490), (96, 388), (328, 388)], ACCENT, dash=True) +text(150, 440, "registers new version", f_lbl, MUTED, anchor="lm") + +# held-out set is only ever scored against +arrow([(588, 85), (700, 85), (700, 286), (674, 286)], GREEN, dash=True) +text(710, 178, "scored against only,", f_lbl, GREEN, anchor="lm") +text(710, 194, "never trained on", f_lbl, GREEN, anchor="lm") + +text(24, 544, "DVC orchestrates every stage Β· MLflow records every run", + f_lbl, MUTED, anchor="lm") + +img.resize((W // S, H // S), Image.LANCZOS).save(OUT, "PNG", optimize=True) +print(f"wrote {OUT} ({OUT.stat().st_size // 1024} KB)") diff --git a/results/baseline_metrics.json b/results/baseline_metrics.json new file mode 100644 index 0000000..b070a2f --- /dev/null +++ b/results/baseline_metrics.json @@ -0,0 +1,67 @@ +{ + "captured_on": "2026-05-18", + "day": 1, + "phase": "1 β€” audit + temporal-split fix + baseline", + "story": "Pre-fix benchmark AUC of 0.921 was inflated by two leakage paths: (1) train_test_split with stratify=y in src/train.py was a RANDOM split that mixed future txns into training; (2) benchmark_fdb fell back to data/raw/sparkov.csv which is the SAME file used for training. After fixing both β€” per-source temporal split in train.py + held-out sparkov_test.csv in benchmark_fdb β€” the honest sparkov AUC vs AutoGluon is 0.795, not 0.921.", + "pre_fix": { + "split_strategy": "random_stratified", + "benchmark_data_source": "data/raw/sparkov.csv (same as train set)", + "reported_sparkov_benchmark_auc": 0.9210076846056613, + "delta_vs_AutoGluon_baseline_0.952": -0.030992, + "honest_assessment": "INVALID β€” both random split and training-set test β†’ leakage stacked" + }, + "post_fix_combined_temporal_test": { + "split_strategy": "temporal_per_source (last 20% of each source by txn_timestamp)", + "test_rows": 1531859, + "test_fraud_rate": 0.00378, + "test_accuracy": 0.996564, + "test_auc": 0.998906, + "test_average_precision": 0.827993, + "test_precision": 0.5268, + "test_recall": 0.8974, + "test_f1": 0.6639, + "note": "Inflated by paysim's near-deterministic balance_change_orig signal (paysim is 83% of combined data)." + }, + "post_fix_per_source_train_time_test": { + "source_paysim": { + "rows": 1272524, + "fraud_rate": 0.003343, + "auc": 0.999650, + "average_precision": 0.917024 + }, + "source_sparkov": { + "rows": 259335, + "fraud_rate": 0.005931, + "auc": 0.996608, + "average_precision": 0.809023, + "note": "Held-out window is the final 20% of sparkov.csv (~Mar–Jun 2020). Same data-collection regime as training so the model interpolates well." + } + }, + "post_fix_benchmark_held_out_sparkov_test_file": { + "test_file": "data/raw/sparkov_test.csv", + "date_range": "2020-06-21 12:14:25 β†’ 2020-12-31 23:59:34", + "rows": 555719, + "our_auc": 0.794896, + "baselines": { + "AutoGluon": {"baseline_auc": 0.952, "delta": -0.157104, "status": "BELOW"}, + "H2O AutoML": {"baseline_auc": 0.947, "delta": -0.152104, "status": "BELOW"}, + "AutoSklearn": {"baseline_auc": 0.931, "delta": -0.136104, "status": "BELOW"} + }, + "headline": "0.795 β€” this is Sentinel's honest sparkov AUC for the rest of the sprint. The 0.921 claim retires today.", + "interpretation": "Held-out file covers Jun–Dec 2020, a time period never seen in any form during training. The 0.20 AP/0.20 AUC drop vs the train-time test slice quantifies real distribution shift β€” and that's the gap auto-retrain (Day 3) must close." + }, + "mlflow_run": { + "experiment": "sentinel-day01-temporal-split", + "run_name": "day01_phase1_temporal_split", + "tracking_uri": "sqlite:///mlflow.db (local)", + "logged_params": ["split_strategy", "test_size", "random_state", "n_estimators", "max_depth", "learning_rate", "scale_pos_weight", "use_smote", "early_stopping_rounds", "train_rows", "test_rows", "train_fraud_rate", "test_fraud_rate"], + "logged_metrics": ["test_accuracy", "test_auc", "test_average_precision"], + "logged_artifacts": ["xgboost_model (mlflow.xgboost)", "models/fraud_model.pkl"] + }, + "files_changed_today": [ + "src/combine_datasets.py β€” added txn_timestamp column, replaced shuffle with chronological sort", + "src/preprocess.py β€” added txn_timestamp to required schema + passthrough", + "src/train.py β€” replaced train_test_split with temporal_split_per_source; wrapped run in mlflow.start_run()", + "src/benchmark_fdb.py β€” prefer held-out sparkov_test.csv over leaking sparkov.csv fallback" + ] +} diff --git a/results/day05/day05_leaderboard.csv b/results/day05/day05_leaderboard.csv new file mode 100644 index 0000000..df1a2b5 --- /dev/null +++ b/results/day05/day05_leaderboard.csv @@ -0,0 +1,6 @@ +rank,strategy,training_scope,sample_weight,threshold,in_dist_sparkov_auc,oot_sparkov_test_auc,oot_ap,oot_recall,oot_precision,oot_f1,delta_vs_autogluon,notes +1,AutoGluon (FDB published baseline),FDB standard,n/a,n/a,,0.952,,,,,0.0,reference baseline +2,Day-5 Optuna + source-balanced weights -threshold=0.5,"full (6.1M, paysim 0.60x / sparkov 2.95x)",source_balanced,0.5,0.9971775606973103,0.951958845551572,0.20830353279765584,0.2592074592074592,0.2744323790720632,0.2666027331575162,-0.0,AUC ties AutoGluon; recall jumps 6x +3,Day-5 Optuna + source-balanced weights -threshold=tau*,"full (6.1M, paysim 0.60x / sparkov 2.95x)",source_balanced,0.8939489722251892,0.9971775606973103,0.951958845551572,0.20830353279765584,0.14685314685314685,0.5526315789473685,0.23204419889502761,-0.0,"F1 trade: higher precision (0.55), lower recall (0.15)" +4,Day-5 Optuna best (30 trials) -full retrain,full (6.1M),uniform,0.5,0.9970274251784879,0.915380324555611,0.067267067319249,0.042890442890442894,0.34328358208955223,0.0762536261914629,-0.0366,tuning recovers ranking; recall still collapses at 0.5 +5,"Day-1 honest baseline (random XGB defaults, temporal split fix)",full (6.1M),uniform,0.5,0.996608,0.794896,,,,,-0.1571,honest baseline -pre-tuning diff --git a/results/day06/ablation_modelling.csv b/results/day06/ablation_modelling.csv new file mode 100644 index 0000000..a2a66af --- /dev/null +++ b/results/day06/ablation_modelling.csv @@ -0,0 +1,5 @@ +view,layer,description,oot_auc,oot_auprc,oot_recall_at_0_5,oot_precision_at_0_5,delta_auc_vs_prev_layer,train_seconds,n_train_rows +modelling,L0_naive_random_default,"Naive notebook: random split, XGB defaults, no balancing",0.5466625224585316,0.09136609725819243,0.08158508158508158,0.5271084337349398,,19.57696139998734,480000 +modelling,L1_temporal_default,+temporal split (Day-1 honest baseline fix),0.6574214679548385,0.16115083339577455,0.068997668997669,0.6166666666666667,0.11075894549630683,16.252756399917416,479999 +modelling,L2_temporal_source_balanced_default,+source-balanced sample weights,0.696180215592977,0.1608393025972209,0.1076923076923077,0.7751677852348994,0.03875874763813858,17.605792300077155,479999 +modelling,L3_temporal_source_balanced_optuna,+Optuna tuning (Day-5 champion),0.9479773886870319,0.23198044768793025,0.15897435897435896,0.41585365853658535,0.2517971730940548,45.4373118999647,479999 diff --git a/results/day06/frontier_comparison.csv b/results/day06/frontier_comparison.csv new file mode 100644 index 0000000..8b2fbf6 --- /dev/null +++ b/results/day06/frontier_comparison.csv @@ -0,0 +1,4 @@ +strategy,auc,auprc,recall_at_0_5,precision_at_0_5,f1_at_0_5,tp,fp,fn,latency_s_per_query,cost_usd_per_query,cost_usd_at_1k_qps_per_day +Sentinel champion (Day-5 XGBoost: temporal split + source-balanced + Optuna),0.9155555555555556,0.5259898688226262,0.05,1.0,0.09523809523809523,1,0,19,5.968949990347028e-05,5e-09,0.43200000000000005 +"Naive notebook XGBoost (random split, defaults, no MLOps)",0.6263888888888889,0.4153448019838129,0.05,1.0,0.09523809523809523,1,0,19,7.340950018260629e-05,5e-09,0.43200000000000005 +Claude Opus 4.6 LLM-judged (frontier model),0.6225,0.35122716815016986,0.3,0.8571428571428571,0.4444444444444444,6,1,14,1.8162843168403366,0.0144756,1250691.84 diff --git a/results/drift_replay_summary.json b/results/drift_replay_summary.json new file mode 100644 index 0000000..0a62c6c --- /dev/null +++ b/results/drift_replay_summary.json @@ -0,0 +1,49 @@ +{ + "captured_on": "2026-05-20T06:43:55.745261+00:00", + "day": 3, + "phase": "2b \u2014 synthetic drift replay", + "n_days": 30, + "injection_day": 23, + "injection_feature": "amount", + "injection_sigma": 2.0, + "injection_shift": 318.034, + "monitored_features": [ + "amount", + "tx_amount_log", + "amount_zscore", + "balance_change_orig", + "balance_ratio", + "balance_change_abs", + "balance_change_log", + "hour_of_day" + ], + "detector_params": { + "ks_pvalue_threshold": 0.01, + "ks_stat_threshold": 0.15, + "psi_threshold": 0.25 + }, + "true_drift_days": [ + 23, + 24, + 25, + 26, + 27, + 28, + 29 + ], + "predicted_drift_days": [ + 23, + 24, + 25, + 26, + 27, + 28, + 29 + ], + "precision": 1.0, + "recall": 1.0, + "first_fired_day": 23, + "detection_lag_days": 0, + "pre_injection_max_proba_psi": 0.0968, + "post_injection_min_proba_psi": 2.9164 +} \ No newline at end of file diff --git a/results/drift_retrain_events.csv b/results/drift_retrain_events.csv new file mode 100644 index 0000000..1aa70ed --- /dev/null +++ b/results/drift_retrain_events.csv @@ -0,0 +1,4 @@ +first_fired_day,consecutive_days,triggered_on_day,train_window_days,train_rows,shadow_day,shadow_rows,shadow_auprc,shadow_auc,prod_auprc,promote_decision,promote_reason,new_model_version,run_id,seconds_detect_to_retrain_start,seconds_train,seconds_shadow_eval,seconds_register_and_alias,seconds_end_to_end +23,2,24,22;23,4876,24,3177,0.552,0.9151,0.0551,True,shadow_auprc 0.5520 >= prod_auprc 0.0551 - tol 0.010,7,7c1aeb21b6a5404ca58dcc6c1f05cafa,0.0097,0.6774,0.0704,0.398,30.093 +25,2,26,24;25,5900,26,1587,0.805,0.9968,0.7561,True,shadow_auprc 0.8050 >= prod_auprc 0.7561 - tol 0.010,8,de5d55b9f4db444dbf734822ce22adfe,0.0052,0.754,0.0381,0.0305,6.8528 +27,2,28,26;27,3259,28,1844,0.7653,0.97,0.7575,True,shadow_auprc 0.7653 >= prod_auprc 0.7575 - tol 0.010,9,d4bc3ea99dd44650a978570f18008043,0.0057,0.4912,0.0364,0.0292,6.4964 diff --git a/results/registry_rollback_times.csv b/results/registry_rollback_times.csv new file mode 100644 index 0000000..abd7b43 --- /dev/null +++ b/results/registry_rollback_times.csv @@ -0,0 +1,6 @@ +iteration,from_version,to_version,alias_flip_seconds,audit_tag_seconds,total_seconds +1,3,2,0.004739,0.009159,0.013898 +2,2,3,0.003841,0.008138,0.011978 +3,3,2,0.004033,0.007871,0.011903 +4,2,3,0.003722,0.007716,0.011439 +5,3,2,0.003945,0.007619,0.011564 diff --git a/results/throughput_speedup.csv b/results/throughput_speedup.csv new file mode 100644 index 0000000..68d3ffa --- /dev/null +++ b/results/throughput_speedup.csv @@ -0,0 +1,4 @@ +rows_requested,rows_actual,pandas_seconds,dask_seconds,pandas_rows_per_sec,dask_rows_per_sec,speedup_pandas_over_dask,dask_partitions,deterministic,max_abs_diff_overall +100000,100000,0.0874,0.8945,1144552.39,111800.17,0.098,16,True,4.333866598926761e-12 +500000,500000,0.381,2.2605,1312468.93,221185.82,0.169,16,True,4.185651825139303e-12 +1000000,1000000,0.8674,3.8482,1152848.32,259861.75,0.225,16,True,5.5243587482323164e-12