Skip to content

About

Binary CNN classifier for detecting defective cast pump impellers, comparing CNN-from-scratch vs. ResNet18 transfer learning. Trained and evaluated end-to-end with a cost-based deployment threshold.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

42 Commits

Folders and files

Repository files navigation

Casting Defect Classifier

Binary CNN classifier that flags defective cast pump impellers from photographs, trained and evaluated end-to-end with a cost-based deployment threshold.

Sample images: defective vs. ok castings

Result

  • Best strategy (last_block, fine-tuned ResNet18): test F₂ of 0.998
  • 1 of 453 defective parts missed (0.22% escape rate), zero false alarms
  • Estimated $0.09 per inspected part under the assumed cost model

Read the full project report for the complete story: data pipeline, model comparison, hyperparameter search, threshold selection, and held-out test results, with real numbers and plots rendered inline.

Getting the data

This repo does not include the dataset (data/ is gitignored). To set it up:

  1. Download casting product image data for quality inspection from Kaggle (requires a free account).
  2. Unzip into data/, matching this layout:
    data/
      train/
        def_front/
        ok_front/
      test/
        def_front/
        ok_front/
    
  3. evaluate.py and select_threshold.py also need a checkpoint under experiments/checkpoints/, produced by training first.

Heads up, the source dataset contains 64 train/test duplicates. As shipped, 64 images are byte-identical between the train and test folders. This pipeline finds them by content hash and drops them from the training pool automatically (exclude_duplicates: true in configs/config.yaml), so nothing is needed on your end here. Matching on filename is not sufficient: the dataset also contains same-name/different-content pairs, so only hashing isolates the real duplicates. Worth checking if you use this dataset elsewhere, since leaving them in inflates test scores.

Quickstart

Docker:

docker build -t casting-defect-classifier .
docker run --shm-size=2g \
  -v $(pwd)/data:/app/data \
  -v $(pwd)/experiments:/app/experiments \
  casting-defect-classifier \
  python -m src.train --config configs/config.yaml

Local:

source venv/bin/activate
pip install -r requirements.txt
python -m src.train --config configs/config.yaml

Then, to select a deployment threshold and score it once on test:

python -m src.select_threshold --config configs/config.yaml --checkpoint experiments/checkpoints/last_block.pt
python -m src.evaluate         --config configs/config.yaml --checkpoint experiments/checkpoints/last_block.pt

See CLAUDE.md for the full command reference, including running via Docker in detail.

What's in the repo

  • src/ -- training pipeline, model factory, HPO, threshold selection, evaluation
  • configs/config.yaml -- single source of truth for all hyperparameters, paths, and cost assumptions
  • notebooks/project_report.ipynb -- the full write-up
  • docs/decisions.md -- design decisions and their rationale
  • tests/ -- 40 pytest tests, run with pytest tests/ -v

Design decisions

Every non-obvious choice, why F-beta over a weighted loss, why the threshold is chosen on validation and never test, why the linear probe shares no code with the training pipeline, and more, is written up in docs/decisions.md.

Limitations

Cost assumptions are estimated, not measured from a real production line; results are single-seed throughout; the test set is 715 images. See the notebook's Limitations and next steps section for the complete list.

About this project

Built to practice three things at once: an engineering-focused ML project rather than a single notebook, a real comparison across transfer-learning strategies, and containerizing a training pipeline with Docker, my first time using it. All three show up above: docs/decisions.md records the pipeline's design reasoning, the strategy comparison lives in the report, and "Running via Docker" in CLAUDE.md covers the last one.

Lessons learned

  • Decide the evaluation protocol before the first training run. Early versions of this project selected checkpoints on the test split, because no validation split existed yet. Introducing one meant re-running every experiment. The pipeline now enforces the separation structurally: select_threshold.py never loads the test split at all, so the mistake is not expressible rather than merely discouraged.
  • Fix correctness issues when found, not later. The 64 duplicate train/test images were knowingly left in early as a negligible-impact compromise. Removing them later invalidated results that had already been computed.
  • Write decisions down while making them. docs/decisions.md was started on day three, and is the only reason the reasoning behind early choices was still recoverable months later.
  • Update docs in the same commit as the code. References to an abandoned Colab setup survived in the documentation for weeks after the code itself was deleted.

AI assistance

This project was developed with AI assistance (Claude, via Claude Code), used as a mentor and reviewer rather than an autonomous author. Per a working agreement kept in CLAUDE.md, the training pipeline (src/) was written by me, Claude's role there was to guide, question, and review, not write code, and any design or methodology decision was presented as a trade-off before I decided on it. Some artifacts were written directly by Claude at my request.

About

Binary CNN classifier for detecting defective cast pump impellers, comparing CNN-from-scratch vs. ResNet18 transfer learning. Trained and evaluated end-to-end with a cost-based deployment threshold.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages