Skip to content

Repository files navigation

FLO-ML

Train a model on your laptop the same way you will later train it on Amazon SageMaker — without an AWS account and without a cloud bill.

This repo is a local-first MLOps workbench. You start one Docker stack, open a console in the browser, click Train, and watch a real sklearn churn model land in MLflow with its data and artifacts sitting in a local fake-S3.


What this is doing

A typical SageMaker workflow is: data in S3 → training job in a container → metrics in a tracker → a registered model you can promote.

Teams usually skip that on a laptop and write a notebook instead. Then they rewrite everything for SageMaker. FLO-ML keeps one contract:

Step On your laptop (now) On AWS (later)
Object storage Floci at http://localhost:4566 — S3-shaped, dummy keys test/test Real Amazon S3
Experiment tracking MLflow at http://localhost:5000 SageMaker MLflow App
Training compute Your Python, or the flo-ml/sklearn-train Docker image SageMaker Training (ml.m5.xlarge, GPU, …)
Control surface FLO-ML console at http://localhost:8088 Same scripts, different env file

Floci is not SageMaker. It does not implement CreateTrainingJob. It stands in for S3, IAM, and logs. The training code is src/train.py — a SageMaker-shaped entry point that reads SM_CHANNEL_* and writes to SM_MODEL_DIR.

The sample job is a small churn classifier (HistGradientBoosting) on an anonymized fixture. It is meant to prove the pipeline, not to be a production model.


Architecture

Interactive diagram (open in a browser):

open architecture.html

That file is the workbench map: Local MLOps Stack in the middle, User (browser / CLI) on the left.

Node Kind What it is
User External You, via browser or CLI
FLO-ML Console :8088 Frontend Control surface — train, inspect runs, promote Staging
Training Engine src/train.py Backend SageMaker-shaped entry point (laptop Python or the training image)
MLflow Server :5000 Backend Experiment tracking + model registry
Floci S3 :4566 Cloud (local) AWS-shaped object store for datasets, artifacts, models
Scoring Client :8090 Backend Inference API — loads the registered model, does not train

Traffic in the diagram:

flowchart LR
  user[User<br/>browser / CLI]
  subgraph stack [Local MLOps Stack]
    console[FLO-ML Console :8088]
    trainer[Training Engine<br/>src/train.py]
    mlflow[MLflow Server :5000]
    floci[Floci S3 :4566]
    scorer[Scoring Client :8090]
  end
  user -->|Manage Pipeline| console
  user -->|Get Prediction| scorer
  console -->|Trigger Train| trainer
  trainer -->|S3 I/O datasets and models| floci
  trainer -->|Log metrics / register| mlflow
  scorer -->|Load registered model| mlflow
Loading
From → to Label
User → Console Manage pipeline
User → Scoring client Get prediction
Console → Training engine Trigger train
Training engine → Floci S3 I/O (datasets / models)
Training engine → MLflow Log metrics / register
Scoring client → MLflow Load registered model

Three jobs the stack is split into (same cards as in architecture.html):

  • Control & train — console triggers runs; trainer uses SageMaker-shaped entry points.
  • Storage & tracking — Floci emulates S3 for artifacts; MLflow tracks lineage.
  • Inference — scoring client is decoupled from training; registered models are promoted to Staging.

Two training modes — same script, same data, same MLflow experiment:

  • Fast (laptop Python) — runs src/train.py directly. Seconds. Use this while you iterate.
  • Container (SageMaker-shaped) — docker run flo-ml/sklearn-train:local. This is the image you would send to AWS.

You do not need both every time. Fast is the default in the console.

Local runs can be promoted to Staging. They cannot be marked Production. That is deliberate: a laptop experiment must never become a live endpoint.


The sample problem: telco-style churn

The tables in data/fixtures/ are synthetic. They are generated by sklearn.datasets.make_classification (seed 42) and then rescaled so the columns look like a small internet/telco account file. They are not real customers. Metrics here prove the pipeline, not a business lift.

Column What it means Why it is in the model
tenure_months Months the account has been open (1–72) New accounts leave more often
monthly_charges Recurring bill High bill + short tenure is a churn pattern
total_charges Rough lifetime spend Derived from tenure × monthly
contract_month_to_month 1 = no term commitment Classic churn flag
has_fiber 1 = high-speed product Product mix
support_tickets Recent support contacts (0–8) Friction
late_payments Late payments (0–6) Payment stress
addon_count Add-on products (0–5) Attach / stickiness
churn 1 = left, 0 = stayed Label (~28% positives)

500 train / 150 valid rows, split with stratification so both classes appear in each file. Bootstrap copies them to s3://datasets/churn/train/ and s3://datasets/churn/valid/.


How the model is trained

src/train.py is a SageMaker-shaped entry point: it reads the train and validation channels, writes model.joblib to SM_MODEL_DIR, and logs the run to MLflow.

Algorithm: scikit-learn HistGradientBoostingClassifier — histogram gradient boosting, the sklearn analogue of LightGBM. It is a strong CPU default for small tabular data and does not need a GPU.

Pipeline (in order):

  1. Median-impute missing numerics
  2. Standard-scale every column
  3. Boosted trees with max_depth=3, learning_rate=0.08, max_iter=80, l2_regularization=0.1, random_state=42
  4. Predict P(churn); label = 1 if probability ≥ 0.5

Same script whether you click Fast (laptop Python) or Container (the image AWS would run).


Quality checks

Before the model is treated as a successful run:

Check What happens
Schema Train and valid must share the same feature columns
Target churn must be 0/1; both classes must appear in each split
Empty features A column that is entirely missing fails the job
Holdout ROC-AUC ≥ 0.70 Primary gate (FLO_ML_MIN_ROC_AUC). Fail → non-zero exit, nothing to promote
Average precision Logged (better than accuracy under 28% churn)
Accuracy at 0.5 Logged as a secondary number; do not optimize it alone
Confusion matrix PNG artifact in MLflow
Lineage Data S3 URI + etag, git SHA, image digest, model signature + input example

Unit tests in tests/test_train_unit.py train on the fixture and assert ROC-AUC > 0.70, and they assert the gate rejects a weak score.


What you need

  • Docker Desktop (or Engine + Compose v2)
  • No AWS account
  • Optional: Python 3.11+ only if you want the CLI instead of the console

Run it (the whole flow)

1. Start the stack

cd flo-ml
cp .env.example .env          # already filled with dummy Floci keys
docker compose up --build -d

Wait until these three are green (first build takes a few minutes):

curl -fsS http://localhost:4566 >/dev/null && echo "Floci ok"
curl -fsS http://localhost:5000/health && echo
curl -fsS http://localhost:8088/api/status | python3 -m json.tool | head
Open this What it is
architecture.html Interactive workbench diagram (open the file in a browser)
http://localhost:8088 FLO-ML console — train, track, promote
http://localhost:8090 Scoring client — consumes the registered model
http://localhost:5000 MLflow — run history, metrics, model versions
localhost:4566 Floci S3 API (not a webpage)

2. Use the console

  1. Open http://localhost:8088.
  2. Confirm the pills at the top: Floci S3 up, MLflow up, buckets datasets, mlflow-artifacts, models.
  3. Leave Fast — laptop Python selected (or pick Container if you want the SageMaker image).
  4. Click Train and log to MLflow. Watch the log until you see a ROC-AUC.
  5. A new row appears under Recent MLflow runs. Click MLflow to inspect it.
  6. Click Promote latest → Staging (Production is rejected on purpose).
  7. Click Score. You should get a churn probability for the sample customer. Change the numbers and score again.

That is the entire inner loop: data → train → track → register → score.

3. Consume the model (client)

Training writes a registered model churn in MLflow. A separate scoring client loads that model and is what a CRM or billing app would call. It does not train.

Open http://localhost:8090, or:

# Python client / CLI (from the repo root)
python -m client describe
python -m client score --json client/examples/likely_churn.json
python -m client score --json client/examples/likely_stay.json
python -m client score --csv data/fixtures/churn_valid.csv

# HTTP (same client, as a service)
curl -sS http://localhost:8090/v1/model | python3 -m json.tool
curl -sS -X POST http://localhost:8090/v1/score \
  -H 'Content-Type: application/json' \
  -d @client/examples/likely_churn.json

The JSON body for HTTP is {"features": { ... }}. The CLI --json file can be the feature object itself.

Each score returns churn_probability, a 0/1 label at 0.5, and a retention action:

P(churn) action Meaning
≥ 0.70 offer_retention High risk — outreach / save offer
≥ 0.50 monitor Elevated — watch the next cycle
< 0.50 no_action Likely to stay

POST /v1/reload after you train a new version so the client drops its cached model.

Code: client/churn.py (ChurnClient.from_registry / from_joblib), CLI python -m client, HTTP client/service.py.

4. Stop

docker compose down

Buckets survive a restart (floci-data volume). Add -v only if you want a clean slate.


Optional: same flow from the terminal

Only needed if you prefer a shell to the console.

python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements-ui.txt

# Fast
python jobs/estimator.py --env local --instance process

# Same training image SageMaker will use (builds it on first run)
python jobs/estimator.py --env local --instance local

# Promote
python jobs/cli.py promote --name churn --stage Staging

Makefile shortcuts: make up, make train-process, make train, make ui, make score, make down.


Files that matter

Path Role
ui/ Browser console + API
client/ Scoring client: Python API, CLI, HTTP service on :8090
src/train.py SageMaker-shaped training entry point
jobs/estimator.py CLI runner (process / docker / later AWS)
configs/local.yaml Local environment contract
architecture.html Interactive architecture diagram
docker-compose.yml Floci + MLflow + console + scoring client
data/fixtures/ Tiny anonymized churn CSVs (not production data)
infra/aws/ IAM role shape for the real account (P1)

Environment contract

Switching from laptop to AWS is config, not a rewrite.

Variable Local AWS
AWS_ENDPOINT_URL http://localhost:4566 unset
AWS_ACCESS_KEY_ID / secret test / test role / SSO
MLFLOW_TRACKING_URI http://localhost:5000 SageMaker MLflow App ARN
S3_DATA_URI s3://datasets/churn/train/ your account bucket
SAGEMAKER_INSTANCE process or local ml.m5.xlarge / ml.g5.xlarge
ENV_NAME local dev / staging / prod

FAQ

Is Floci a SageMaker emulator? No. It is local AWS for S3/IAM/ECR/logs. Training is Docker (or laptop Python).

Why two train buttons? Fast = iterate. Container = prove the image that AWS will run. Same src/train.py.

Will this spend money? Not while ENV_NAME=local. Dummy keys only.

GPU? CPU is the supported default. local_gpu needs NVIDIA Container Toolkit; on macOS it is best-effort.

Can I mark a local model Production? No. Change ENV_NAME off local after you point at real AWS.

MLflow client/server versions? Pin mlflow==3.7.0 on both. A 3.x client against a 2.x server fails.

Cloud training (--env aws --instance ml.m5.xlarge) is the next phase. The IAM template is infra/aws/sagemaker-execution-role.yaml.

About

Local-first MLOps workbench that reproduces a SageMaker-style training pipeline locally — Floci (S3), MLflow tracking, and a portable scikit-learn churn classifier with Docker + console for training, evaluation, and model promotion.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages