Code for "Prediction-Powered Ranking of Large Language Models", NeurIPS 2024.
-
Updated
Oct 28, 2024 - Jupyter Notebook
Code for "Prediction-Powered Ranking of Large Language Models", NeurIPS 2024.
Replication code and data for the journal article titled "The Mixed Subjects Design: Treating Large Language Models as Potentially Informative Observations."
Python and R package for semisupervised mean estimation and causal inference with AIPW, calibration, and practical uncertainty quantification.
Certifying retrieval-policy deployment with non-neutral AI judges: decision weights, control variates, and what judge accuracy does not tell you (TMLR submission artifact)
Reproducibility materials for prediction-powered inference in image-based maize height phenotyping
Statistics for LLM-judged evaluations: bias-corrected scores, real confidence intervals, and comparisons that hold up
Risk-controlled evaluation & release-gating for LLM/RAG/agent systems: LLM-as-judge with bias probes, Prediction-Powered Inference for label-efficient CIs, and a conformal CI release gate.
Finite-sample guarantees for Jev (TypeSafe's System One). Conformal risk control turns calibrated probabilities into certified routing thresholds; prediction-powered inference audits them. 2,412 decisions on CLINC150 for $0.23 — including the shift and prevalence cases where the guarantee breaks.
To associate your repository with the prediction-powered-inference topic, visit your repo's landing page and select "manage topics."