Skip to content
#

benchmark-evaluation

Here are 14 public repositories matching this topic...

Option-reversal control for paired binary vision–language benchmarks: estimates a model's answer-slot bias and perceptual accuracy from one extra inference pass per question, attributes each failure to slot or perception, and reports scores against the correct chance floors. Ongoing project; code only.

  • Updated Sep 19, 2026

Add this topic to your repo

To associate your repository with the benchmark-evaluation topic, visit your repo's landing page and select "manage topics."

Learn more