A kernel-userland protocol enforcing information-theoretic bounds on AI adaptivity leakage, benchmark gaming, and capability spillover.
-
Updated
Mar 5, 2026 - Rust
A kernel-userland protocol enforcing information-theoretic bounds on AI adaptivity leakage, benchmark gaming, and capability spillover.
Your AI “improved.” Did it get smarter—or memorize the test? Find out before you trust the score.
Behavioral evaluation framework for sentience-, emotion-, and welfare-related AI claims, with anti-sandbagging analysis.
Horizon-Eval: evaluation-integrity framework for portable long-horizon agent benchmarks, with QA gates, trajectory auditing, replayable run bundles, and safety-gap analysis.
Reproducible audit of the computational claims in arXiv:2604.27092. Reproduces the released accuracies to within 5e-4 and finds that both benchmarks place repeated measurements of the same physical pair in training and test sets.
Scanner for evaluator trust boundaries — the places your eval pipeline trusts things it shouldn't. Born from published research; 100+ defect reports across 40+ orgs. Free, MIT, forever.
Evidence-bound AI systems: agent-benchmark audits, tool effects, checkpoint restart invariants, and fail-closed recovery.
To associate your repository with the benchmark-integrity topic, visit your repo's landing page and select "manage topics."