[NeurIPS 2025] 🌐 WebThinker: Empowering Large Reasoning Models with Deep Research Capability
-
Updated
Dec 8, 2025 - Python
[NeurIPS 2025] 🌐 WebThinker: Empowering Large Reasoning Models with Deep Research Capability
Open LLM leaderboard featuring Xiaomi MiMo v2.5 & MiMo 100T head-to-head with GPT-5, Claude, Gemini, DeepSeek, Llama 4. ARC-AGI · SWE-Bench · MMLU-Pro · GPQA · HumanEval · BFCL.
Universal LLM evaluation framework — HumanEval, SWE-bench, GPQA, BigCodeBench, τ-Bench, Terminal-Bench via a single make command
Sourced, schema-validated catalogue of LLM & agent benchmarks: what each tests, saturation & contamination status, top scores with provenance (official / independent / self-reported). JSON API + site.
Analyzing and recovering reasoning degradation in LLMs under 4-bit quantization using QLoRA and GRPO.
To associate your repository with the gpqa topic, visit your repo's landing page and select "manage topics."