Skip to content
#

data-contamination

Here are 22 public repositories matching this topic...

Python .pyc decompiler (3.0–3.14) with a contamination-aware benchmark harness. Rule-only pass + one Codex call per module; evaluated on fuzz-synthetic (LLM-naïve) and *-obf (anonymised) corpora to put a number on the memorisation share. Three independent PyPI packages: pychd, pychd-pyfuzz, pychd-pyobf.

  • Updated May 27, 2026
  • Python

Leakage-free real-time evaluation of open-weights LLMs for US CPI inflation forecasting. Introduces the memorization premium (seen vs. unseen forecast-error gap) and a three-role decomposition (direct forecaster, FOMC-text extractor, combiner). Reproduces every number in the IJF manuscript's Table 3 from the committed checkpoint.

  • Updated Aug 9, 2026
  • Python

Deterministic evaluation harness for AP document-matching agents. Scores 3-way findings against a hand-audited, held-out golden dataset: per-category precision and recall, over-flagging measured on a zero-defect control, byte-reproducible scorecards, answer key structurally out of reach.

  • Updated Jul 28, 2026
  • Python

Promotion gate and verifier toolbox for AI systems: exposure audits, paired PASS/HOLD/BLOCK receipts, disposable public report cards, and an open MCP endpoint.

  • Updated Aug 8, 2026
  • Python

Improve this page

Add a description, image, and links to the data-contamination topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the data-contamination topic, visit your repo's landing page and select "manage topics."

Learn more