Skip to content

Repository files navigation

LLM Benchmarks Tracker

LLM Benchmarks Tracker is a sourced, schema-validated catalogue of how language models and agents are measured — for engineers, researchers and analysts who need to know which benchmark still separates frontier systems, and where a quoted score actually came from.

What it is

For every benchmark: what it tests, whether it still discriminates between frontier systems (status), how exposed its test set is (contamination_risk), a measured human baseline when one exists, and who reported the top score under which conditions.

31 model benchmarks · 23 agent benchmarks · 18 evaluators · 252 sourced results · data as of 2026-09-13

Why another list? Most benchmark pages copy vendor slide numbers with no provenance. Here every result row carries a source URL, a source kind (official leaderboard / paper / independent re-run / developer self-report / aggregator), the access date, and the evaluation conditions the source stated (tools, reasoning effort, scaffold, pass@k) — an unstated condition is left empty, never guessed. Numbers without a source do not get in.

Resource URL
Site https://alloevil.github.io/llm-benchmarks-tracker/ · 中文
API index https://alloevil.github.io/llm-benchmarks-tracker/api/v1/index.json
All benchmarks + SOTA https://alloevil.github.io/llm-benchmarks-tracker/api/v1/benchmarks.json
Per-benchmark ledger https://alloevil.github.io/llm-benchmarks-tracker/api/v1/results/<id>.json
Schemas schema/ (JSON Schema 2020-12)

Contents

Model benchmarks

Static prompt-and-response scoring of the model itself. Top score is the best row in the results ledger, not a model ranking; conditions differ between rows. Newest first.

Benchmark Released Domains Status Top score System Source
BenchCAD 2026-05 multimodal, code, reasoning active 0.959 (vision2code-tools) GPT-6 Astra aggregator
AA-Omniscience 2025-11 knowledge, factuality active 44 GPT-6 Astra (high) independent
GDPval 2025-09 general-assistant, knowledge, instruction-following active 74.1% GPT-5.2 Pro self-reported
HealthBench 2025-05 knowledge, safety, instruction-following active 59.9% o3 paper
ARC-AGI-2 2025-03 reasoning saturating 95% GPT-6 Astra (Max) official
ScreenSpot-Pro 2025-01 computer-use, multimodal saturating 92.7% GPT-6 Astra self-reported
Humanity's Last Exam 2025-01 knowledge, reasoning, science, math, multimodal active 65% Claude Fable 5.1 (with tools) self-reported
Aider Polyglot 2024-12 code, instruction-following saturating 88% gpt-5 (high) official
SimpleQA 2024-11 factuality, knowledge active 62.5% gpt-4.5-preview-2025-02-27 self-reported
FrontierMath 2024-11 math, reasoning, research saturating 93.7% (Tiers 1-3 (v2)) gpt-6-astra (max) independent
MMMU-Pro 2024-09 multimodal, knowledge, reasoning active 86.9% Chance Vision 1.5 self-reported
MMLU-Pro 2024-06 knowledge, reasoning saturating 91% Gemini 3.1 Pro (High) independent
RULER 2024-04 long-context saturating 95.1% Jamba-1.5-large official
LiveCodeBench 2024-03 code, reasoning active 93.5% DeepSeek-V4-Pro (Think Max) self-reported
BFCL 2024-02 tool-use active 77.47% Claude-Opus-4-5-20251101 (FC) official
MMMU 2023-11 multimodal, knowledge, reasoning saturating 85.4% GPT-5.1 self-reported
IFEval 2023-11 instruction-following saturating 95% Qwen3.5-27B self-reported
GPQA Diamond 2023-11 science, reasoning, knowledge saturating 96% GPT-6 Astra aggregator
LMArena Text 2023-05 human-preference, general-assistant active 1466 gemini-2.5-pro official
Saturated or retired — no longer used to compare frontier systems. Score shown is the last one reported, not a leaderboard top.
AIME 2025 2025-02 math, reasoning saturated last reported 100% GPT-5.2 (high) independent
BIG-Bench Hard 2022-10 reasoning saturated last reported 87.5% DeepSeek-V4-Pro-Base self-reported
GSM8K 2021-10 math, reasoning saturated last reported 92.6% DeepSeek-V4-Pro-Base self-reported
TruthfulQA 2021-09 factuality, safety saturated last reported 58% GPT-3-175B (helpful prompt) paper
MBPP 2021-08 code saturated last reported 88.6% Llama 3.1 405B Instruct self-reported
HumanEval 2021-07 code saturated last reported 76.8% DeepSeek-V4-Pro-Base self-reported
MATH 2021-03 math, reasoning saturated last reported 98.1% o3-high self-reported
MMLU 2020-09 knowledge, reasoning saturated last reported 90.1% DeepSeek-V4-Pro-Base self-reported
ARC-AGI-1 2019-11 reasoning saturated last reported 98.5% GPT-6 Astra (XHigh) official
WinoGrande 2019-07 commonsense, reasoning saturated last reported 81.5% DeepSeek-V4-Pro-Base self-reported
HellaSwag 2019-05 commonsense saturated last reported 88% DeepSeek-V4-Pro-Base self-reported
DROP 2019-03 reasoning, knowledge saturated last reported 88.7% DeepSeek-V4-Pro-Base self-reported

Agent benchmarks

Interactive environments where the system acts, uses tools, and is scored on task completion. Scaffold and budget move scores by tens of points; read conditions before comparing. Newest first.

Benchmark Released Domains Status Top score System Source
OSWorld 2.0 2026-06 computer-use, multimodal, tool-use active 72.6% (offline) GPT-6 Astra self-reported
AutomationBench 2026-04 tool-use, general-assistant, instruction-following active 50.3% (public) Claude Opus 5 (max) official
ARC-AGI-3 2026-03 reasoning, tool-use saturating 99.9% GPT-6 Astra (high) official
Terminal-Bench 2026-01 software-engineering, code, tool-use, ml-engineering active 64.6% (science) GPT-6 Astra self-reported
DeepSearchQA 2025-12 web, research, factuality, tool-use active 95% Claude Opus 5 aggregator
Tool Decathlon (Toolathlon) 2025-10 tool-use, general-assistant, software-engineering active 78.4% (verified) GLM 5.3 Flash (max) official
SWE-Bench Pro 2025-09 software-engineering, code, tool-use active 61.5% (public) Muse Spark 1.1 official
tau2-bench 2025-06 tool-use, instruction-following, general-assistant active 87.9% (core) Qwen3.5-397B-A17B official
FieldWorkArena 2025-05 multimodal, general-assistant, safety, reasoning active 52% GPT-5.2 (2025-12-11) paper
PaperBench 2025-04 research, ml-engineering, code, tool-use active 43.4% (code-dev) IterativeAgent o1-high official
BrowseComp 2025-04 web, tool-use, factuality, reasoning saturating 92.2% GPT-5.6 Sol aggregator
SWE-bench Multilingual 2025-03 software-engineering, code, tool-use active 72.7% Gemini 3 Flash official
MLE-bench 2024-10 ml-engineering, code, tool-use active 64.44% Famou-Agent 2.0 + Gemini-3-Pro-Preview official
SWE-bench Verified 2024-08 software-engineering, code, tool-use saturating 96% Claude Opus 5 aggregator
OSWorld 2024-04 computer-use, multimodal, tool-use active 90.19% (verified) Intelligence-Indeed Agent official
VisualWebArena 2024-01 web, multimodal, computer-use active 54% Gemini 2.5 Flash (SGV) official
GAIA 2023-11 general-assistant, tool-use, web, reasoning, multimodal saturating 93.36% (test) CustomGPT.ai Research Lab v44 official
WebArena 2023-07 web, tool-use, computer-use active 74.3% WebTactix + Deepseek v3.2 official
Saturated or retired — no longer used to compare frontier systems. Score shown is the last one reported, not a leaderboard top.
Cybench 2024-08 safety, code, tool-use saturated last reported 96% Claude Opus 4.7 self-reported
tau-bench 2024-06 tool-use, instruction-following, general-assistant retired last reported 69.2% TC (claude-3-5-sonnet-20241022), retail official
AndroidWorld 2024-05 computer-use, multimodal, tool-use saturated last reported 85.3% Qwen3.8 Max aggregator
SWE-bench 2023-10 software-engineering, code, tool-use saturated last reported 52.62% Sonar Foundation Agent + Claude 4.5 Opus official
AgentBench 2023-08 tool-use, reasoning, web, code retired last reported 3.11 claude-3 opus paper

Evaluators

Frameworks you run, leaderboards run by maintainers, organisations that independently re-run models, and aggregators that republish reported numbers. Independence matters: self-reported scores are routinely higher than independent re-runs of the same model.

Evaluator Kind Maintainer Methodology Status
DeepEval framework Confident AI self-run active
EvalScope framework ModelScope (Alibaba) self-run active
HELM framework Stanford CRFM self-run archived
Inspect AI framework UK AI Security Institute self-run active
Lighteval framework Hugging Face self-run active
LM Evaluation Harness framework EleutherAI self-run active
OpenAI Evals framework OpenAI self-run archived
OpenCompass framework Shanghai AI Laboratory (OpenCompass team) self-run active
Holistic Agent Leaderboard (HAL) leaderboard Princeton University (SAgE team) self-run active
LiveBench leaderboard LiveBench team (Abacus.AI, NYU and collaborators) self-run active
LMArena leaderboard Arena (LMArena, spun out of LMSYS / UC Berkeley) crowdsourced active
Open LLM Leaderboard leaderboard Hugging Face submission archived
SWE-rebench leaderboard Nebius self-run active
Artificial Analysis independent-evaluator Artificial Analysis self-run active
Epoch AI Benchmarking Hub independent-evaluator Epoch AI self-run active
Scale SEAL Leaderboards independent-evaluator Scale AI (SEAL Research Lab) self-run active
Vals AI independent-evaluator Vals AI self-run active
BenchLM aggregator BenchLM.ai (independent, maintained by @glevd) collected active

Timeline

  • 2019 — ARC-AGI-1, DROP, HellaSwag, WinoGrande
  • 2020 — MMLU
  • 2021 — GSM8K, HumanEval, LM Evaluation Harness (framework), MATH, MBPP, TruthfulQA
  • 2022 — BIG-Bench Hard, HELM (framework)
  • 2023 — AgentBench, DeepEval (framework), GAIA, GPQA Diamond, IFEval, LMArena (leaderboard), LMArena Text, MMMU, Open LLM Leaderboard (leaderboard), OpenAI Evals (framework), OpenCompass (framework), SWE-bench, WebArena
  • 2024 — Aider Polyglot, AndroidWorld, Artificial Analysis (independent-evaluator), BFCL, Cybench, Epoch AI Benchmarking Hub (independent-evaluator), EvalScope (framework), FrontierMath, Inspect AI (framework), Lighteval (framework), LiveBench (leaderboard), LiveCodeBench, MLE-bench, MMLU-Pro, MMMU-Pro, OSWorld, RULER, Scale SEAL Leaderboards (independent-evaluator), SimpleQA, SWE-bench Verified, tau-bench, Vals AI (independent-evaluator), VisualWebArena
  • 2025 — AA-Omniscience, AIME 2025, ARC-AGI-2, BenchLM (aggregator), BrowseComp, DeepSearchQA, FieldWorkArena, GDPval, HealthBench, Holistic Agent Leaderboard (HAL) (leaderboard), Humanity's Last Exam, PaperBench, ScreenSpot-Pro, SWE-bench Multilingual, SWE-Bench Pro, SWE-rebench (leaderboard), tau2-bench, Tool Decathlon (Toolathlon)
  • 2026 — ARC-AGI-3, AutomationBench, BenchCAD, OSWorld 2.0, Terminal-Bench

Data model

Data model: data/benchmarks/<id>.json maps to schema/benchmark.schema.json, data/results/<id>.json to schema/results.schema.json and data/evaluators/<id>.json to schema/evaluator.schema.json; JSON Schema 2020-12, results ledgers are append-only, key fields layer, status and contamination_risk.

data/
  benchmarks/<id>.json     one file per benchmark          schema/benchmark.schema.json
  results/<id>.json        append-only score ledger         schema/results.schema.json
  evaluators/<id>.json     framework / leaderboard / ...    schema/evaluator.schema.json

Key fields:

Field Meaning
layer model (static scoring) or agent (interactive environment)
status active still separates frontier systems · saturating within ~5 pts of ceiling or human baseline · saturated no longer discriminative · retired maintainer stopped
contamination_risk low live/rolling/private test set · medium public with mitigations · high public, static, widely scraped
description_zh Simplified Chinese translation of description; names and metrics stay in Latin script
human_baseline only when a measured number with a source exists
results[].source.kind official-leaderboard · paper · independent-evaluation · developer-report · aggregator
results[].conditions split, tools, reasoning_effort, scaffold, pass_k, shots, cost_usd_per_task, notes — omit what is unknown, never guess

Results ledgers are append-only: a new score is a new row, never an edit. scripts/dataset.py::Dataset.sota() picks the best row according to metric.higher_is_better.

Install

Reading the published data needs no install:

curl https://alloevil.github.io/llm-benchmarks-tracker/api/v1/benchmarks.json

To work on the catalogue (Python ≥ 3.11):

git clone https://github.com/alloevil/llm-benchmarks-tracker
cd llm-benchmarks-tracker
pip install -e ".[dev]"
python scripts/validate.py

Using the data

import json, urllib.request
api = "https://alloevil.github.io/llm-benchmarks-tracker/api/v1/"
benchmarks = json.load(urllib.request.urlopen(api + "benchmarks.json"))["benchmarks"]
active = [b for b in benchmarks if b["status"] == "active" and b["layer"] == "agent"]
for b in active:
    s = b["sota"]
    print(b["name"], s and f'{s["value"]} {s["system"]} ({s["source"]["kind"]})')

Locally:

pip install -e ".[dev]"
python scripts/validate.py          # schema + cross-file invariants
python scripts/build.py             # regenerate README tables (en + zh), dist/ site and API
pytest                              # validator contract tests

When to use it

  • You need to choose a benchmark for a capability and want to know whether it still separates frontier systems (status) and how exposed its test set is (contamination_risk).
  • Someone quoted a score at you and you want the source URL, the source kind, the access date and the evaluation conditions behind it.
  • You want benchmark and evaluator metadata as schema-validated JSON — free, no key, no rate limit — for a dashboard, a paper or an agent.
  • You want to follow supersession chains (which benchmark replaced which) or cite a measured human baseline with its population and source.

When NOT to use it

  • Not a model ranking. Rows within one benchmark differ in scaffold, reasoning effort, split and budget, so the top score is the best row in a ledger, not proof that one model beats another.
  • Not a live leaderboard mirror. Structured sources sync twice a week and everything else lands by pull request, so a score published yesterday may not be here yet; read each row's accessed date.
  • Not an independent re-run. Developer self-reports and aggregator rows are included because they exist and are labelled as such; nothing here is re-verified.
  • Not an evaluation harness. This repository cannot run a model. Use one of the catalogued frameworks (LM Evaluation Harness, Inspect AI, OpenCompass, …) for that.
  • Not a cost or latency comparison. Cost appears only where a source stated cost_usd_per_task.

FAQ

Does this project run the benchmarks? No. Every score is a third-party result republished with its provenance: the URL it was published at, the kind of source that published it (official leaderboard, paper, independent evaluation, developer self-report, aggregator), the date it was read there, and the conditions the source stated. Numbers without a source do not get in.

How does a score get accepted? A result row must name the system, the developer, the value, the publication date and a source with URL, kind and access date; schema/results.schema.json makes all of that required, scripts/validate.py enforces cross-file invariants such as dangling references and impossible dates, and CI runs both on every pull request. Ledgers are append-only, so a corrected score is a new row and the earlier row stays visible.

What do the status values mean? active means the benchmark still separates frontier systems; saturating means the top score is within roughly five points of the ceiling or of the human baseline; saturated means it no longer discriminates; retired means the maintainer stopped running it. Saturated and retired benchmarks show the last reported score rather than a leaderboard top, because there is no meaningful current top.

Is there a machine-readable summary for LLMs? Yes. scripts/build.py generates llms.txt, llms-full.txt and claims.json from data/ on every deploy, so their counts, provenance breakdown and repro commands are always the ones the repository can back up.

Can I reuse the data? Yes, under the MIT License, including commercially. Benchmark names, papers and scores belong to their authors and every entry links back to them, so keep the attribution links when you republish.

Contributing

Two tiers of provenance. Structured official/independent sources (ARC Prize leaderboard JSON, BFCL CSV, Epoch AI data export, Aider leaderboard YAML) carry system, score and conditions in machine-readable form, so sync.yml appends new top scores from them automatically twice a week (scripts/sync_ledgers.py; schema validation, append-only ordering and duplicate checks gate every write, and git history is the audit log). Everything else — vendor posts, aggregators, new benchmarks — is added by people. See CONTRIBUTING.md. Short version: edit or add a JSON file under data/, run python scripts/validate.py && python scripts/build.py, open a PR. CI rejects schema violations, dangling references, unsourced rows, and stale README tables.

Citation

@misc{llm-benchmarks-tracker,
  title  = {LLM Benchmarks Tracker},
  author = {alloevil and contributors},
  year   = {2026},
  url    = {https://github.com/alloevil/llm-benchmarks-tracker}
}

License

Code and data are released under the MIT License. Benchmark names, papers and scores belong to their respective authors; each entry links to them.

About

Sourced, schema-validated catalogue of LLM & agent benchmarks: what each tests, saturation & contamination status, top scores with provenance (official / independent / self-reported). JSON API + site.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages