Train a model on your laptop the same way you will later train it on Amazon SageMaker — without an AWS account and without a cloud bill.
This repo is a local-first MLOps workbench. You start one Docker stack, open a console in the browser, click Train, and watch a real sklearn churn model land in MLflow with its data and artifacts sitting in a local fake-S3.
A typical SageMaker workflow is: data in S3 → training job in a container → metrics in a tracker → a registered model you can promote.
Teams usually skip that on a laptop and write a notebook instead. Then they rewrite everything for SageMaker. FLO-ML keeps one contract:
| Step | On your laptop (now) | On AWS (later) |
|---|---|---|
| Object storage | Floci at http://localhost:4566 — S3-shaped, dummy keys test/test |
Real Amazon S3 |
| Experiment tracking | MLflow at http://localhost:5000 |
SageMaker MLflow App |
| Training compute | Your Python, or the flo-ml/sklearn-train Docker image |
SageMaker Training (ml.m5.xlarge, GPU, …) |
| Control surface | FLO-ML console at http://localhost:8088 |
Same scripts, different env file |
Floci is not SageMaker. It does not implement CreateTrainingJob. It stands in for S3, IAM, and logs. The training code is src/train.py — a SageMaker-shaped entry point that reads SM_CHANNEL_* and writes to SM_MODEL_DIR.
The sample job is a small churn classifier (HistGradientBoosting) on an anonymized fixture. It is meant to prove the pipeline, not to be a production model.
Interactive diagram (open in a browser):
open architecture.htmlThat file is the workbench map: Local MLOps Stack in the middle, User (browser / CLI) on the left.
| Node | Kind | What it is |
|---|---|---|
| User | External | You, via browser or CLI |
FLO-ML Console :8088 |
Frontend | Control surface — train, inspect runs, promote Staging |
Training Engine src/train.py |
Backend | SageMaker-shaped entry point (laptop Python or the training image) |
MLflow Server :5000 |
Backend | Experiment tracking + model registry |
Floci S3 :4566 |
Cloud (local) | AWS-shaped object store for datasets, artifacts, models |
Scoring Client :8090 |
Backend | Inference API — loads the registered model, does not train |
Traffic in the diagram:
flowchart LR
user[User<br/>browser / CLI]
subgraph stack [Local MLOps Stack]
console[FLO-ML Console :8088]
trainer[Training Engine<br/>src/train.py]
mlflow[MLflow Server :5000]
floci[Floci S3 :4566]
scorer[Scoring Client :8090]
end
user -->|Manage Pipeline| console
user -->|Get Prediction| scorer
console -->|Trigger Train| trainer
trainer -->|S3 I/O datasets and models| floci
trainer -->|Log metrics / register| mlflow
scorer -->|Load registered model| mlflow
| From → to | Label |
|---|---|
| User → Console | Manage pipeline |
| User → Scoring client | Get prediction |
| Console → Training engine | Trigger train |
| Training engine → Floci | S3 I/O (datasets / models) |
| Training engine → MLflow | Log metrics / register |
| Scoring client → MLflow | Load registered model |
Three jobs the stack is split into (same cards as in architecture.html):
- Control & train — console triggers runs; trainer uses SageMaker-shaped entry points.
- Storage & tracking — Floci emulates S3 for artifacts; MLflow tracks lineage.
- Inference — scoring client is decoupled from training; registered models are promoted to Staging.
Two training modes — same script, same data, same MLflow experiment:
- Fast (laptop Python) — runs
src/train.pydirectly. Seconds. Use this while you iterate. - Container (SageMaker-shaped) —
docker run flo-ml/sklearn-train:local. This is the image you would send to AWS.
You do not need both every time. Fast is the default in the console.
Local runs can be promoted to Staging. They cannot be marked Production. That is deliberate: a laptop experiment must never become a live endpoint.
The tables in data/fixtures/ are synthetic. They are generated by sklearn.datasets.make_classification (seed 42) and then rescaled so the columns look like a small internet/telco account file. They are not real customers. Metrics here prove the pipeline, not a business lift.
| Column | What it means | Why it is in the model |
|---|---|---|
tenure_months |
Months the account has been open (1–72) | New accounts leave more often |
monthly_charges |
Recurring bill | High bill + short tenure is a churn pattern |
total_charges |
Rough lifetime spend | Derived from tenure × monthly |
contract_month_to_month |
1 = no term commitment | Classic churn flag |
has_fiber |
1 = high-speed product | Product mix |
support_tickets |
Recent support contacts (0–8) | Friction |
late_payments |
Late payments (0–6) | Payment stress |
addon_count |
Add-on products (0–5) | Attach / stickiness |
churn |
1 = left, 0 = stayed | Label (~28% positives) |
500 train / 150 valid rows, split with stratification so both classes appear in each file. Bootstrap copies them to s3://datasets/churn/train/ and s3://datasets/churn/valid/.
src/train.py is a SageMaker-shaped entry point: it reads the train and validation channels, writes model.joblib to SM_MODEL_DIR, and logs the run to MLflow.
Algorithm: scikit-learn HistGradientBoostingClassifier — histogram gradient boosting, the sklearn analogue of LightGBM. It is a strong CPU default for small tabular data and does not need a GPU.
Pipeline (in order):
- Median-impute missing numerics
- Standard-scale every column
- Boosted trees with
max_depth=3,learning_rate=0.08,max_iter=80,l2_regularization=0.1,random_state=42 - Predict P(churn); label = 1 if probability ≥ 0.5
Same script whether you click Fast (laptop Python) or Container (the image AWS would run).
Before the model is treated as a successful run:
| Check | What happens |
|---|---|
| Schema | Train and valid must share the same feature columns |
| Target | churn must be 0/1; both classes must appear in each split |
| Empty features | A column that is entirely missing fails the job |
| Holdout ROC-AUC ≥ 0.70 | Primary gate (FLO_ML_MIN_ROC_AUC). Fail → non-zero exit, nothing to promote |
| Average precision | Logged (better than accuracy under 28% churn) |
| Accuracy at 0.5 | Logged as a secondary number; do not optimize it alone |
| Confusion matrix | PNG artifact in MLflow |
| Lineage | Data S3 URI + etag, git SHA, image digest, model signature + input example |
Unit tests in tests/test_train_unit.py train on the fixture and assert ROC-AUC > 0.70, and they assert the gate rejects a weak score.
- Docker Desktop (or Engine + Compose v2)
- No AWS account
- Optional: Python 3.11+ only if you want the CLI instead of the console
cd flo-ml
cp .env.example .env # already filled with dummy Floci keys
docker compose up --build -dWait until these three are green (first build takes a few minutes):
curl -fsS http://localhost:4566 >/dev/null && echo "Floci ok"
curl -fsS http://localhost:5000/health && echo
curl -fsS http://localhost:8088/api/status | python3 -m json.tool | head| Open this | What it is |
|---|---|
architecture.html |
Interactive workbench diagram (open the file in a browser) |
| http://localhost:8088 | FLO-ML console — train, track, promote |
| http://localhost:8090 | Scoring client — consumes the registered model |
| http://localhost:5000 | MLflow — run history, metrics, model versions |
localhost:4566 |
Floci S3 API (not a webpage) |
- Open http://localhost:8088.
- Confirm the pills at the top: Floci S3 up, MLflow up, buckets
datasets,mlflow-artifacts,models. - Leave Fast — laptop Python selected (or pick Container if you want the SageMaker image).
- Click Train and log to MLflow. Watch the log until you see a ROC-AUC.
- A new row appears under Recent MLflow runs. Click MLflow to inspect it.
- Click Promote latest → Staging (Production is rejected on purpose).
- Click Score. You should get a churn probability for the sample customer. Change the numbers and score again.
That is the entire inner loop: data → train → track → register → score.
Training writes a registered model churn in MLflow. A separate scoring client loads that model and is what a CRM or billing app would call. It does not train.
Open http://localhost:8090, or:
# Python client / CLI (from the repo root)
python -m client describe
python -m client score --json client/examples/likely_churn.json
python -m client score --json client/examples/likely_stay.json
python -m client score --csv data/fixtures/churn_valid.csv
# HTTP (same client, as a service)
curl -sS http://localhost:8090/v1/model | python3 -m json.tool
curl -sS -X POST http://localhost:8090/v1/score \
-H 'Content-Type: application/json' \
-d @client/examples/likely_churn.jsonThe JSON body for HTTP is {"features": { ... }}. The CLI --json file can be the feature object itself.
Each score returns churn_probability, a 0/1 label at 0.5, and a retention action:
| P(churn) | action |
Meaning |
|---|---|---|
| ≥ 0.70 | offer_retention |
High risk — outreach / save offer |
| ≥ 0.50 | monitor |
Elevated — watch the next cycle |
| < 0.50 | no_action |
Likely to stay |
POST /v1/reload after you train a new version so the client drops its cached model.
Code: client/churn.py (ChurnClient.from_registry / from_joblib), CLI python -m client, HTTP client/service.py.
docker compose downBuckets survive a restart (floci-data volume). Add -v only if you want a clean slate.
Only needed if you prefer a shell to the console.
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements-ui.txt
# Fast
python jobs/estimator.py --env local --instance process
# Same training image SageMaker will use (builds it on first run)
python jobs/estimator.py --env local --instance local
# Promote
python jobs/cli.py promote --name churn --stage StagingMakefile shortcuts: make up, make train-process, make train, make ui, make score, make down.
| Path | Role |
|---|---|
ui/ |
Browser console + API |
client/ |
Scoring client: Python API, CLI, HTTP service on :8090 |
src/train.py |
SageMaker-shaped training entry point |
jobs/estimator.py |
CLI runner (process / docker / later AWS) |
configs/local.yaml |
Local environment contract |
architecture.html |
Interactive architecture diagram |
docker-compose.yml |
Floci + MLflow + console + scoring client |
data/fixtures/ |
Tiny anonymized churn CSVs (not production data) |
infra/aws/ |
IAM role shape for the real account (P1) |
Switching from laptop to AWS is config, not a rewrite.
| Variable | Local | AWS |
|---|---|---|
AWS_ENDPOINT_URL |
http://localhost:4566 |
unset |
AWS_ACCESS_KEY_ID / secret |
test / test |
role / SSO |
MLFLOW_TRACKING_URI |
http://localhost:5000 |
SageMaker MLflow App ARN |
S3_DATA_URI |
s3://datasets/churn/train/ |
your account bucket |
SAGEMAKER_INSTANCE |
process or local |
ml.m5.xlarge / ml.g5.xlarge |
ENV_NAME |
local |
dev / staging / prod |
Is Floci a SageMaker emulator? No. It is local AWS for S3/IAM/ECR/logs. Training is Docker (or laptop Python).
Why two train buttons?
Fast = iterate. Container = prove the image that AWS will run. Same src/train.py.
Will this spend money?
Not while ENV_NAME=local. Dummy keys only.
GPU?
CPU is the supported default. local_gpu needs NVIDIA Container Toolkit; on macOS it is best-effort.
Can I mark a local model Production?
No. Change ENV_NAME off local after you point at real AWS.
MLflow client/server versions?
Pin mlflow==3.7.0 on both. A 3.x client against a 2.x server fails.
Cloud training (--env aws --instance ml.m5.xlarge) is the next phase. The IAM template is infra/aws/sagemaker-execution-role.yaml.