You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
When Better Means Less: Quantifying What Benchmarks Miss Between Model Generations. 2,310 controlled comparisons show GPT-5 series lost 6.7x creativity and gained 4.4x false refusals vs chatgpt-4o-latest — invisible to standard benchmarks.
Code, data, and results for "The Correlation Mirage: Benchmark Dependence Collapses for Top-Performing LLMs" (EMNLP 2026). Copula-based tail dependence analysis of LLM benchmark suites.
How much of a phishing benchmark's score is detection skill, and how much is an artifact of how the dataset was built? A zero-parameter regex scores TSS 0.99 on PhiUSIIL; the trained model scores 0.000 on live data.
Does a benchmark overstate operational skill when the distribution shift is ordinary and physical? Houston ground-level ozone exceedance forecasting — the control domain in a multi-domain study of benchmark-vs-operational ML performance.
Option-reversal control for paired binary vision–language benchmarks: estimates a model's answer-slot bias and perceptual accuracy from one extra inference pass per question, attributes each failure to slot or perception, and reports scores against the correct chance floors. Ongoing project; code only.
llmverify is a lightweight Python tool for externally auditing LLM APIs and verifying whether custom providers and resellers truly serve the frontier models they claim to run, using layered probes and benchmark-based statistics.
Audit of BixBench, FutureHouse's bioinformatics-agent benchmark: replicated runs, replicated grading, re-derived answer keys — and what a single pooled score hides.