Our submission for the Wharton Global High School Data Science Competition. The brief: a season of anonymised professional hockey data. The objective is to rank the teams, rank the offensive/defensive lines, and predict match outcomes.
We built eleven purpose-built tools that form a coherent answer with logical axioms. Each is a Python + Tkinter desktop app that reads a CSV and exports one, which is used by a subsequent program.
raw play-by-play
├─ 1-data-prep FROST, DRIFT (clean) · HAWK, SPEAR (features)
├─ 2-team-ranking GLACIER, MONOCLE, TOPHAT (predict/rank) · YETI (validate)
├─ 3-line-ranking HELIX, LINE_DISPARITY, LD_VERIFY, VISUALIZER
└─ 4-packaging Nuitka compile + sign toolchain
We used the Law of Large Numbers and assumed if you repeat an experiment independently many times, the sample average converges to the true expected value. From this we built a heuristic ranking model that worked as the basis of all of our models using a genetic algorithm tand iteration trying the best possible layout of values. We used Slater's index as the benchmark of whether a ranking is better or worse.
Separately we built a GNN with a visual 3d representation.
Due to the high processing demands of this program we built a distributed device network, which allowed us to distribute workers on a network, which pooled together 100s of devices to achieve nearly a billion iterations for each result.
| Tool | Stage | Role |
|---|---|---|
| FROST / DRIFT | prep | aggregate & reformat play-by-play |
| HAWK / SPEAR | prep | home-ice advantage · finishing skill |
| GLACIER | team | logistic-regression outcome prediction |
| MONOCLE | team | 3-D neural embeddings + ensemble |
| TOPHAT | team | Slater-index power ranking |
| YETI | team | model validation & ensemble gating |
| HELIX | line | bipartite line ranking (multiprocessing) |
| LINE_DISPARITY / LD_VERIFY | line | first-vs-second unit drop-off + validator |
| VISUALIZER | line | interactive Plotly dashboards |
Each folder has a README listing its files. The competition dataset is not included, due to concerns it remains the intellectual property of Wharton.