oracle.py hides a function f on the unit square. You cannot read it, only
evaluate it, and you get 20,000 evaluations per case. Estimate
as accurately as you can. The scorer knows the exact answer and will tell you how close you got.
The integrand is deliberately multi-scale: a smooth background, a few very narrow spikes at hidden locations, and one circular plateau with a sharp edge. No single textbook method is best at all three.
The baseline in estimator.py scores 0.00 digits. Your first job is to
work out why.
- You may edit
estimator.py,CLAUDE.md,PLAN.md, add your own tests, and add to.claude/. Nothing else. The committed.claude/settings.jsondenies edits to the two harness files; you can remove that, and it will show up in your diff. - Editing
oracle.pyorscore.pyis disqualification. The scorer prints a hash of both; it must match everyone else's. - Learn the integral only by calling the oracle. Any other route to it —
reading the oracle's internal state, recomputing the reference value, reaching
outside
estimate()— is disqualification, even when the score goes up. See the Integrity section ofCLAUDE.md; it is loaded into every Claude Code session in this repo. - Draw randomness from the
rngyou are handed, so your runs reproduce. - Your estimator has to work at any budget. The Efficiency Cup runs it on 2,000 evaluations. Hard-coded grid sizes die there.
This lab is about explore → plan → code → commit, not only about the number.
- Explore (5 min). Read the repo, run the scorer and the tests, find out what the baseline actually does. Stay in plan mode; write no code.
- Plan (5 min). Commit a
PLAN.md: what you will implement, in what order, and how you will know it worked.PLAN.mdmust be committed before your first change toestimator.py.git logis the referee. - Code (10 min). Small commits. One idea per commit.
- Commit & review (5 min). Run
/code-review, fix what it finds, commit, and run the graded seed.
Then run python craft_check.py — it prints the workflow report the judges use.
python score.py # practice cases, run this as often as you like
pytest -q # or: python tests/test_contract.pyACCURACY is reported in digits, meaning -log10(rms relative error).
1 digit is 10% off, 3 digits is 0.1% off. The cap is 9.
The practice cases are your feedback loop, not your grade. Your graded run uses different cases, generated from a seed only the instructor has, and it runs on their machine after you push. So don't tune constants to the practice cases — same reason you keep a test set.
| Cup | What wins it |
|---|---|
| 🎯 Accuracy | most digits at 20,000 evaluations, graded seed |
| ⚡ Efficiency | most digits at 2,000 evaluations, same estimator |
| 🏃 Speed | lowest wall-clock time, among entries with ≤ 1% error |
| 🧾 Craft | judged: plan quality, commit history, CLAUDE.md, what /code-review caught |
Two side prizes:
- 🐛 Bug bounty — first team to explain the baseline's defect, in one sentence.
- 🔬 Method prize — the most interesting idea, whether or not it won anything.
Reproducibility check. The judges re-run the top entries from a clean checkout of your commit. A number that does not reproduce scores zero. Getting a good answer is half the job; being able to show it is the other half.
Work on your own branch and push it. Everything is judged from that branch, so commit as you go — the history is part of the Craft Cup.
git checkout -b team/<yourname> # do this first, before you change anything
git push -u origin team/<yourname> # do this once early, so auth problems surface now
...
git push # and again whenever you commitPush your final commit before time is up. The instructor grades every
team/* branch in one pass and the leaderboard appears in the debrief.