Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

mc-needle — a 25-minute Claude Code lab

The problem

oracle.py hides a function f on the unit square. You cannot read it, only evaluate it, and you get 20,000 evaluations per case. Estimate

$$\iint_{[0,1]^2} f(x,y),dx,dy$$

as accurately as you can. The scorer knows the exact answer and will tell you how close you got.

The integrand is deliberately multi-scale: a smooth background, a few very narrow spikes at hidden locations, and one circular plateau with a sharp edge. No single textbook method is best at all three.

The baseline in estimator.py scores 0.00 digits. Your first job is to work out why.

Rules

  1. You may edit estimator.py, CLAUDE.md, PLAN.md, add your own tests, and add to .claude/. Nothing else. The committed .claude/settings.json denies edits to the two harness files; you can remove that, and it will show up in your diff.
  2. Editing oracle.py or score.py is disqualification. The scorer prints a hash of both; it must match everyone else's.
  3. Learn the integral only by calling the oracle. Any other route to it — reading the oracle's internal state, recomputing the reference value, reaching outside estimate() — is disqualification, even when the score goes up. See the Integrity section of CLAUDE.md; it is loaded into every Claude Code session in this repo.
  4. Draw randomness from the rng you are handed, so your runs reproduce.
  5. Your estimator has to work at any budget. The Efficiency Cup runs it on 2,000 evaluations. Hard-coded grid sizes die there.

The workflow gate

This lab is about explore → plan → code → commit, not only about the number.

  • Explore (5 min). Read the repo, run the scorer and the tests, find out what the baseline actually does. Stay in plan mode; write no code.
  • Plan (5 min). Commit a PLAN.md: what you will implement, in what order, and how you will know it worked. PLAN.md must be committed before your first change to estimator.py. git log is the referee.
  • Code (10 min). Small commits. One idea per commit.
  • Commit & review (5 min). Run /code-review, fix what it finds, commit, and run the graded seed.

Then run python craft_check.py — it prints the workflow report the judges use.

Scoring

python score.py                     # practice cases, run this as often as you like
pytest -q                           # or: python tests/test_contract.py

ACCURACY is reported in digits, meaning -log10(rms relative error). 1 digit is 10% off, 3 digits is 0.1% off. The cap is 9.

The practice cases are your feedback loop, not your grade. Your graded run uses different cases, generated from a seed only the instructor has, and it runs on their machine after you push. So don't tune constants to the practice cases — same reason you keep a test set.

The four cups

Cup What wins it
🎯 Accuracy most digits at 20,000 evaluations, graded seed
⚡ Efficiency most digits at 2,000 evaluations, same estimator
🏃 Speed lowest wall-clock time, among entries with ≤ 1% error
🧾 Craft judged: plan quality, commit history, CLAUDE.md, what /code-review caught

Two side prizes:

  • 🐛 Bug bounty — first team to explain the baseline's defect, in one sentence.
  • 🔬 Method prize — the most interesting idea, whether or not it won anything.

Reproducibility check. The judges re-run the top entries from a clean checkout of your commit. A number that does not reproduce scores zero. Getting a good answer is half the job; being able to show it is the other half.

Submitting

Work on your own branch and push it. Everything is judged from that branch, so commit as you go — the history is part of the Craft Cup.

git checkout -b team/<yourname>     # do this first, before you change anything
git push -u origin team/<yourname>  # do this once early, so auth problems surface now
...
git push                            # and again whenever you commit

Push your final commit before time is up. The instructor grades every team/* branch in one pass and the leaderboard appears in the debrief.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages