Binary CNN classifier that flags defective cast pump impellers from photographs, trained and evaluated end-to-end with a cost-based deployment threshold.
- Best strategy (
last_block, fine-tuned ResNet18): test F₂ of 0.998 - 1 of 453 defective parts missed (0.22% escape rate), zero false alarms
- Estimated $0.09 per inspected part under the assumed cost model
Read the full project report for the complete story: data pipeline, model comparison, hyperparameter search, threshold selection, and held-out test results, with real numbers and plots rendered inline.
This repo does not include the dataset (data/ is gitignored). To set it up:
- Download casting product image data for quality inspection from Kaggle (requires a free account).
- Unzip into
data/, matching this layout:data/ train/ def_front/ ok_front/ test/ def_front/ ok_front/ evaluate.pyandselect_threshold.pyalso need a checkpoint underexperiments/checkpoints/, produced by training first.
Heads up, the source dataset contains 64 train/test duplicates. As shipped,
64 images are byte-identical between the train and test folders. This pipeline
finds them by content hash and drops them from the training pool automatically
(exclude_duplicates: true in configs/config.yaml), so nothing is needed on
your end here. Matching on filename is not sufficient: the dataset also contains
same-name/different-content pairs, so only hashing isolates the real duplicates.
Worth checking if you use this dataset elsewhere, since leaving them in inflates
test scores.
Docker:
docker build -t casting-defect-classifier .
docker run --shm-size=2g \
-v $(pwd)/data:/app/data \
-v $(pwd)/experiments:/app/experiments \
casting-defect-classifier \
python -m src.train --config configs/config.yamlLocal:
source venv/bin/activate
pip install -r requirements.txt
python -m src.train --config configs/config.yamlThen, to select a deployment threshold and score it once on test:
python -m src.select_threshold --config configs/config.yaml --checkpoint experiments/checkpoints/last_block.pt
python -m src.evaluate --config configs/config.yaml --checkpoint experiments/checkpoints/last_block.ptSee CLAUDE.md for the full command reference, including running via Docker in detail.
src/-- training pipeline, model factory, HPO, threshold selection, evaluationconfigs/config.yaml-- single source of truth for all hyperparameters, paths, and cost assumptionsnotebooks/project_report.ipynb-- the full write-updocs/decisions.md-- design decisions and their rationaletests/-- 40 pytest tests, run withpytest tests/ -v
Every non-obvious choice, why F-beta over a weighted loss, why the threshold is chosen on validation and never test, why the linear probe shares no code with the training pipeline, and more, is written up in docs/decisions.md.
Cost assumptions are estimated, not measured from a real production line; results are single-seed throughout; the test set is 715 images. See the notebook's Limitations and next steps section for the complete list.
Built to practice three things at once: an engineering-focused ML project
rather than a single notebook, a real comparison across transfer-learning
strategies, and containerizing a training pipeline with Docker, my first time
using it. All three show up above: docs/decisions.md records the pipeline's
design reasoning, the strategy comparison lives in the
report, and "Running via Docker" in
CLAUDE.md covers the last one.
- Decide the evaluation protocol before the first training run. Early
versions of this project selected checkpoints on the test split, because no
validation split existed yet. Introducing one meant re-running every
experiment. The pipeline now enforces the separation structurally:
select_threshold.pynever loads the test split at all, so the mistake is not expressible rather than merely discouraged. - Fix correctness issues when found, not later. The 64 duplicate train/test images were knowingly left in early as a negligible-impact compromise. Removing them later invalidated results that had already been computed.
- Write decisions down while making them.
docs/decisions.mdwas started on day three, and is the only reason the reasoning behind early choices was still recoverable months later. - Update docs in the same commit as the code. References to an abandoned Colab setup survived in the documentation for weeks after the code itself was deleted.
This project was developed with AI assistance (Claude, via Claude Code), used
as a mentor and reviewer rather than an autonomous author. Per a working
agreement kept in CLAUDE.md, the training pipeline (src/) was written by
me, Claude's role there was to guide, question, and review, not write code,
and any design or methodology decision was presented as a trade-off before I
decided on it. Some artifacts were written directly by Claude at my request.
