LLM Benchmarks Tracker is a sourced, schema-validated catalogue of how language models and agents are measured — for engineers, researchers and analysts who need to know which benchmark still separates frontier systems, and where a quoted score actually came from.
For every benchmark: what it tests, whether it still discriminates between frontier systems (status), how exposed its test set is (contamination_risk), a measured human baseline when one exists, and who reported the top score under which conditions.
31 model benchmarks · 23 agent benchmarks · 18 evaluators · 252 sourced results · data as of 2026-09-13
Why another list? Most benchmark pages copy vendor slide numbers with no provenance. Here every result row carries a source URL, a source kind (official leaderboard / paper / independent re-run / developer self-report / aggregator), the access date, and the evaluation conditions the source stated (tools, reasoning effort, scaffold, pass@k) — an unstated condition is left empty, never guessed. Numbers without a source do not get in.
| Resource | URL |
|---|---|
| Site | https://alloevil.github.io/llm-benchmarks-tracker/ · 中文 |
| API index | https://alloevil.github.io/llm-benchmarks-tracker/api/v1/index.json |
| All benchmarks + SOTA | https://alloevil.github.io/llm-benchmarks-tracker/api/v1/benchmarks.json |
| Per-benchmark ledger | https://alloevil.github.io/llm-benchmarks-tracker/api/v1/results/<id>.json |
| Schemas | schema/ (JSON Schema 2020-12) |
- What it is
- Model benchmarks
- Agent benchmarks
- Evaluators
- Timeline
- Data model
- Install
- Using the data
- When to use it
- When NOT to use it
- FAQ
- Contributing
Static prompt-and-response scoring of the model itself. Top score is the best row in the results ledger, not a model ranking; conditions differ between rows. Newest first.
| Benchmark | Released | Domains | Status | Top score | System | Source |
|---|---|---|---|---|---|---|
| BenchCAD | 2026-05 | multimodal, code, reasoning | active | 0.959 (vision2code-tools) | GPT-6 Astra | aggregator |
| AA-Omniscience | 2025-11 | knowledge, factuality | active | 44 | GPT-6 Astra (high) | independent |
| GDPval | 2025-09 | general-assistant, knowledge, instruction-following | active | 74.1% | GPT-5.2 Pro | self-reported |
| HealthBench | 2025-05 | knowledge, safety, instruction-following | active | 59.9% | o3 | paper |
| ARC-AGI-2 | 2025-03 | reasoning | saturating | 95% | GPT-6 Astra (Max) | official |
| ScreenSpot-Pro | 2025-01 | computer-use, multimodal | saturating | 92.7% | GPT-6 Astra | self-reported |
| Humanity's Last Exam | 2025-01 | knowledge, reasoning, science, math, multimodal | active | 65% | Claude Fable 5.1 (with tools) | self-reported |
| Aider Polyglot | 2024-12 | code, instruction-following | saturating | 88% | gpt-5 (high) | official |
| SimpleQA | 2024-11 | factuality, knowledge | active | 62.5% | gpt-4.5-preview-2025-02-27 | self-reported |
| FrontierMath | 2024-11 | math, reasoning, research | saturating | 93.7% (Tiers 1-3 (v2)) | gpt-6-astra (max) | independent |
| MMMU-Pro | 2024-09 | multimodal, knowledge, reasoning | active | 86.9% | Chance Vision 1.5 | self-reported |
| MMLU-Pro | 2024-06 | knowledge, reasoning | saturating | 91% | Gemini 3.1 Pro (High) | independent |
| RULER | 2024-04 | long-context | saturating | 95.1% | Jamba-1.5-large | official |
| LiveCodeBench | 2024-03 | code, reasoning | active | 93.5% | DeepSeek-V4-Pro (Think Max) | self-reported |
| BFCL | 2024-02 | tool-use | active | 77.47% | Claude-Opus-4-5-20251101 (FC) | official |
| MMMU | 2023-11 | multimodal, knowledge, reasoning | saturating | 85.4% | GPT-5.1 | self-reported |
| IFEval | 2023-11 | instruction-following | saturating | 95% | Qwen3.5-27B | self-reported |
| GPQA Diamond | 2023-11 | science, reasoning, knowledge | saturating | 96% | GPT-6 Astra | aggregator |
| LMArena Text | 2023-05 | human-preference, general-assistant | active | 1466 | gemini-2.5-pro | official |
| Saturated or retired — no longer used to compare frontier systems. Score shown is the last one reported, not a leaderboard top. | ||||||
| AIME 2025 | 2025-02 | math, reasoning | saturated | last reported 100% | GPT-5.2 (high) | independent |
| BIG-Bench Hard | 2022-10 | reasoning | saturated | last reported 87.5% | DeepSeek-V4-Pro-Base | self-reported |
| GSM8K | 2021-10 | math, reasoning | saturated | last reported 92.6% | DeepSeek-V4-Pro-Base | self-reported |
| TruthfulQA | 2021-09 | factuality, safety | saturated | last reported 58% | GPT-3-175B (helpful prompt) | paper |
| MBPP | 2021-08 | code | saturated | last reported 88.6% | Llama 3.1 405B Instruct | self-reported |
| HumanEval | 2021-07 | code | saturated | last reported 76.8% | DeepSeek-V4-Pro-Base | self-reported |
| MATH | 2021-03 | math, reasoning | saturated | last reported 98.1% | o3-high | self-reported |
| MMLU | 2020-09 | knowledge, reasoning | saturated | last reported 90.1% | DeepSeek-V4-Pro-Base | self-reported |
| ARC-AGI-1 | 2019-11 | reasoning | saturated | last reported 98.5% | GPT-6 Astra (XHigh) | official |
| WinoGrande | 2019-07 | commonsense, reasoning | saturated | last reported 81.5% | DeepSeek-V4-Pro-Base | self-reported |
| HellaSwag | 2019-05 | commonsense | saturated | last reported 88% | DeepSeek-V4-Pro-Base | self-reported |
| DROP | 2019-03 | reasoning, knowledge | saturated | last reported 88.7% | DeepSeek-V4-Pro-Base | self-reported |
Interactive environments where the system acts, uses tools, and is scored on task completion. Scaffold and budget move scores by tens of points; read conditions before comparing. Newest first.
| Benchmark | Released | Domains | Status | Top score | System | Source |
|---|---|---|---|---|---|---|
| OSWorld 2.0 | 2026-06 | computer-use, multimodal, tool-use | active | 72.6% (offline) | GPT-6 Astra | self-reported |
| AutomationBench | 2026-04 | tool-use, general-assistant, instruction-following | active | 50.3% (public) | Claude Opus 5 (max) | official |
| ARC-AGI-3 | 2026-03 | reasoning, tool-use | saturating | 99.9% | GPT-6 Astra (high) | official |
| Terminal-Bench | 2026-01 | software-engineering, code, tool-use, ml-engineering | active | 64.6% (science) | GPT-6 Astra | self-reported |
| DeepSearchQA | 2025-12 | web, research, factuality, tool-use | active | 95% | Claude Opus 5 | aggregator |
| Tool Decathlon (Toolathlon) | 2025-10 | tool-use, general-assistant, software-engineering | active | 78.4% (verified) | GLM 5.3 Flash (max) | official |
| SWE-Bench Pro | 2025-09 | software-engineering, code, tool-use | active | 61.5% (public) | Muse Spark 1.1 | official |
| tau2-bench | 2025-06 | tool-use, instruction-following, general-assistant | active | 87.9% (core) | Qwen3.5-397B-A17B | official |
| FieldWorkArena | 2025-05 | multimodal, general-assistant, safety, reasoning | active | 52% | GPT-5.2 (2025-12-11) | paper |
| PaperBench | 2025-04 | research, ml-engineering, code, tool-use | active | 43.4% (code-dev) | IterativeAgent o1-high | official |
| BrowseComp | 2025-04 | web, tool-use, factuality, reasoning | saturating | 92.2% | GPT-5.6 Sol | aggregator |
| SWE-bench Multilingual | 2025-03 | software-engineering, code, tool-use | active | 72.7% | Gemini 3 Flash | official |
| MLE-bench | 2024-10 | ml-engineering, code, tool-use | active | 64.44% | Famou-Agent 2.0 + Gemini-3-Pro-Preview | official |
| SWE-bench Verified | 2024-08 | software-engineering, code, tool-use | saturating | 96% | Claude Opus 5 | aggregator |
| OSWorld | 2024-04 | computer-use, multimodal, tool-use | active | 90.19% (verified) | Intelligence-Indeed Agent | official |
| VisualWebArena | 2024-01 | web, multimodal, computer-use | active | 54% | Gemini 2.5 Flash (SGV) | official |
| GAIA | 2023-11 | general-assistant, tool-use, web, reasoning, multimodal | saturating | 93.36% (test) | CustomGPT.ai Research Lab v44 | official |
| WebArena | 2023-07 | web, tool-use, computer-use | active | 74.3% | WebTactix + Deepseek v3.2 | official |
| Saturated or retired — no longer used to compare frontier systems. Score shown is the last one reported, not a leaderboard top. | ||||||
| Cybench | 2024-08 | safety, code, tool-use | saturated | last reported 96% | Claude Opus 4.7 | self-reported |
| tau-bench | 2024-06 | tool-use, instruction-following, general-assistant | retired | last reported 69.2% | TC (claude-3-5-sonnet-20241022), retail | official |
| AndroidWorld | 2024-05 | computer-use, multimodal, tool-use | saturated | last reported 85.3% | Qwen3.8 Max | aggregator |
| SWE-bench | 2023-10 | software-engineering, code, tool-use | saturated | last reported 52.62% | Sonar Foundation Agent + Claude 4.5 Opus | official |
| AgentBench | 2023-08 | tool-use, reasoning, web, code | retired | last reported 3.11 | claude-3 opus | paper |
Frameworks you run, leaderboards run by maintainers, organisations that independently re-run models, and aggregators that republish reported numbers. Independence matters: self-reported scores are routinely higher than independent re-runs of the same model.
| Evaluator | Kind | Maintainer | Methodology | Status |
|---|---|---|---|---|
| DeepEval | framework | Confident AI | self-run | active |
| EvalScope | framework | ModelScope (Alibaba) | self-run | active |
| HELM | framework | Stanford CRFM | self-run | archived |
| Inspect AI | framework | UK AI Security Institute | self-run | active |
| Lighteval | framework | Hugging Face | self-run | active |
| LM Evaluation Harness | framework | EleutherAI | self-run | active |
| OpenAI Evals | framework | OpenAI | self-run | archived |
| OpenCompass | framework | Shanghai AI Laboratory (OpenCompass team) | self-run | active |
| Holistic Agent Leaderboard (HAL) | leaderboard | Princeton University (SAgE team) | self-run | active |
| LiveBench | leaderboard | LiveBench team (Abacus.AI, NYU and collaborators) | self-run | active |
| LMArena | leaderboard | Arena (LMArena, spun out of LMSYS / UC Berkeley) | crowdsourced | active |
| Open LLM Leaderboard | leaderboard | Hugging Face | submission | archived |
| SWE-rebench | leaderboard | Nebius | self-run | active |
| Artificial Analysis | independent-evaluator | Artificial Analysis | self-run | active |
| Epoch AI Benchmarking Hub | independent-evaluator | Epoch AI | self-run | active |
| Scale SEAL Leaderboards | independent-evaluator | Scale AI (SEAL Research Lab) | self-run | active |
| Vals AI | independent-evaluator | Vals AI | self-run | active |
| BenchLM | aggregator | BenchLM.ai (independent, maintained by @glevd) | collected | active |
- 2019 — ARC-AGI-1, DROP, HellaSwag, WinoGrande
- 2020 — MMLU
- 2021 — GSM8K, HumanEval, LM Evaluation Harness (framework), MATH, MBPP, TruthfulQA
- 2022 — BIG-Bench Hard, HELM (framework)
- 2023 — AgentBench, DeepEval (framework), GAIA, GPQA Diamond, IFEval, LMArena (leaderboard), LMArena Text, MMMU, Open LLM Leaderboard (leaderboard), OpenAI Evals (framework), OpenCompass (framework), SWE-bench, WebArena
- 2024 — Aider Polyglot, AndroidWorld, Artificial Analysis (independent-evaluator), BFCL, Cybench, Epoch AI Benchmarking Hub (independent-evaluator), EvalScope (framework), FrontierMath, Inspect AI (framework), Lighteval (framework), LiveBench (leaderboard), LiveCodeBench, MLE-bench, MMLU-Pro, MMMU-Pro, OSWorld, RULER, Scale SEAL Leaderboards (independent-evaluator), SimpleQA, SWE-bench Verified, tau-bench, Vals AI (independent-evaluator), VisualWebArena
- 2025 — AA-Omniscience, AIME 2025, ARC-AGI-2, BenchLM (aggregator), BrowseComp, DeepSearchQA, FieldWorkArena, GDPval, HealthBench, Holistic Agent Leaderboard (HAL) (leaderboard), Humanity's Last Exam, PaperBench, ScreenSpot-Pro, SWE-bench Multilingual, SWE-Bench Pro, SWE-rebench (leaderboard), tau2-bench, Tool Decathlon (Toolathlon)
- 2026 — ARC-AGI-3, AutomationBench, BenchCAD, OSWorld 2.0, Terminal-Bench
data/
benchmarks/<id>.json one file per benchmark schema/benchmark.schema.json
results/<id>.json append-only score ledger schema/results.schema.json
evaluators/<id>.json framework / leaderboard / ... schema/evaluator.schema.json
Key fields:
| Field | Meaning |
|---|---|
layer |
model (static scoring) or agent (interactive environment) |
status |
active still separates frontier systems · saturating within ~5 pts of ceiling or human baseline · saturated no longer discriminative · retired maintainer stopped |
contamination_risk |
low live/rolling/private test set · medium public with mitigations · high public, static, widely scraped |
description_zh |
Simplified Chinese translation of description; names and metrics stay in Latin script |
human_baseline |
only when a measured number with a source exists |
results[].source.kind |
official-leaderboard · paper · independent-evaluation · developer-report · aggregator |
results[].conditions |
split, tools, reasoning_effort, scaffold, pass_k, shots, cost_usd_per_task, notes — omit what is unknown, never guess |
Results ledgers are append-only: a new score is a new row, never an edit. scripts/dataset.py::Dataset.sota() picks the best row according to metric.higher_is_better.
Reading the published data needs no install:
curl https://alloevil.github.io/llm-benchmarks-tracker/api/v1/benchmarks.jsonTo work on the catalogue (Python ≥ 3.11):
git clone https://github.com/alloevil/llm-benchmarks-tracker
cd llm-benchmarks-tracker
pip install -e ".[dev]"
python scripts/validate.pyimport json, urllib.request
api = "https://alloevil.github.io/llm-benchmarks-tracker/api/v1/"
benchmarks = json.load(urllib.request.urlopen(api + "benchmarks.json"))["benchmarks"]
active = [b for b in benchmarks if b["status"] == "active" and b["layer"] == "agent"]
for b in active:
s = b["sota"]
print(b["name"], s and f'{s["value"]} {s["system"]} ({s["source"]["kind"]})')Locally:
pip install -e ".[dev]"
python scripts/validate.py # schema + cross-file invariants
python scripts/build.py # regenerate README tables (en + zh), dist/ site and API
pytest # validator contract tests- You need to choose a benchmark for a capability and want to know whether it still separates frontier systems (
status) and how exposed its test set is (contamination_risk). - Someone quoted a score at you and you want the source URL, the source kind, the access date and the evaluation conditions behind it.
- You want benchmark and evaluator metadata as schema-validated JSON — free, no key, no rate limit — for a dashboard, a paper or an agent.
- You want to follow supersession chains (which benchmark replaced which) or cite a measured human baseline with its population and source.
- Not a model ranking. Rows within one benchmark differ in scaffold, reasoning effort, split and budget, so the top score is the best row in a ledger, not proof that one model beats another.
- Not a live leaderboard mirror. Structured sources sync twice a week and everything else lands by pull request, so a score published yesterday may not be here yet; read each row's
accesseddate. - Not an independent re-run. Developer self-reports and aggregator rows are included because they exist and are labelled as such; nothing here is re-verified.
- Not an evaluation harness. This repository cannot run a model. Use one of the catalogued frameworks (LM Evaluation Harness, Inspect AI, OpenCompass, …) for that.
- Not a cost or latency comparison. Cost appears only where a source stated
cost_usd_per_task.
Does this project run the benchmarks? No. Every score is a third-party result republished with its provenance: the URL it was published at, the kind of source that published it (official leaderboard, paper, independent evaluation, developer self-report, aggregator), the date it was read there, and the conditions the source stated. Numbers without a source do not get in.
How does a score get accepted? A result row must name the system, the developer, the value, the publication date and a source with URL, kind and access date; schema/results.schema.json makes all of that required, scripts/validate.py enforces cross-file invariants such as dangling references and impossible dates, and CI runs both on every pull request. Ledgers are append-only, so a corrected score is a new row and the earlier row stays visible.
What do the status values mean? active means the benchmark still separates frontier systems; saturating means the top score is within roughly five points of the ceiling or of the human baseline; saturated means it no longer discriminates; retired means the maintainer stopped running it. Saturated and retired benchmarks show the last reported score rather than a leaderboard top, because there is no meaningful current top.
Is there a machine-readable summary for LLMs? Yes. scripts/build.py generates llms.txt, llms-full.txt and claims.json from data/ on every deploy, so their counts, provenance breakdown and repro commands are always the ones the repository can back up.
Can I reuse the data? Yes, under the MIT License, including commercially. Benchmark names, papers and scores belong to their authors and every entry links back to them, so keep the attribution links when you republish.
Two tiers of provenance. Structured official/independent sources (ARC Prize leaderboard JSON, BFCL CSV, Epoch AI data export, Aider leaderboard YAML) carry system, score and conditions in machine-readable form, so sync.yml appends new top scores from them automatically twice a week (scripts/sync_ledgers.py; schema validation, append-only ordering and duplicate checks gate every write, and git history is the audit log). Everything else — vendor posts, aggregators, new benchmarks — is added by people. See CONTRIBUTING.md. Short version: edit or add a JSON file under data/, run python scripts/validate.py && python scripts/build.py, open a PR. CI rejects schema violations, dangling references, unsourced rows, and stale README tables.
@misc{llm-benchmarks-tracker,
title = {LLM Benchmarks Tracker},
author = {alloevil and contributors},
year = {2026},
url = {https://github.com/alloevil/llm-benchmarks-tracker}
}Code and data are released under the MIT License. Benchmark names, papers and scores belong to their respective authors; each entry links to them.