EvalFloor: is your LLM eval improvement real? Trying k prompt variants and keeping the best scores points on noise alone — this computes how many.
-
Updated
Sep 21, 2026 - Python
EvalFloor: is your LLM eval improvement real? Trying k prompt variants and keeping the best scores points on noise alone — this computes how many.
Closed Q1 2026: Kalshi weather-contracts trading system. Calibration validation failed at Gate 1; full post-mortem included. Methodology extracted to ai-discipline.
Your A/B test winner is about half as good as it looked. The winner's curse, peeking and power measured on 32,363 real randomised experiments -- with a live demo that re-splits the traffic in your browser.
To associate your repository with the winners-curse topic, visit your repo's landing page and select "manage topics."