A published Llama fine-tune scores 100% on its historical split. The audit shows why: a depth-2 tree ties it, a two-field rule reproduces all 1,000 labels, and one unrelated note flips 84 of 84 negative predictions.
tabular-data supply-chain calibration llama reproducibility logistics case-study fine-tuning olist llm-evaluation temporal-validation label-leakage counterfactual-probing
-
Updated
Sep 22, 2026 - Python