[COLM 2026] MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
-
Updated
Jul 9, 2026 - Python
[COLM 2026] MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
Reproducible and flexible LLM evaluations for scientific reasoning.
ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning
Foundational doctrine establishing the principles behind the Neurotransparency governance framework.
Make AI research agents accountable — give every conclusion a traceable argument graph. MCP server + Claude Code plugin for paper reproduction, hypothesis verification, and auditable scientific reasoning.
Code and lightweight audit artifacts for auditable repair of scientific reasoning graph extraction on a 350-row benchmark.
Explain why two scientific papers disagree — a contradiction explorer for bench scientists that compares study designs and grounds every claim in the source text.
BioReasoner: Training LLMs for grounded scientific reasoning. 0% hallucination rate on citations, 100% format adherence. Cross-domain polymathic insights via Scientific Tribunal evaluation.
LLM + PDDLStream pipeline for executable science problem solving on SciBench.
Scientific computing portfolio covering computational chemistry, biomedical modeling, literature mining, kinetic modeling, biomarker simulation, and AI-assisted scientific evaluation.
Open-ended benchmark and evaluation framework for biological reasoning in LLMs: 12 components, 400 tasks, deterministic dataset release.
A multimodal benchmark for evaluating biological reasoning, beginning with figure interpretation and evidence-calibrated scientific analysis.
Framework for building and evaluating explanatory accounts of evidence using Bayesian network inference. Includes interactive Shiny app.
Evidence-grounded Sci-Evo scientific evolution dataset built from MinerU-parsed open-access protein-engineering papers.
A scientific reasoning game where you win by choosing the experiment that proves you wrong. Built with Codex and GPT-5.6 for OpenAI Build Week.
Benchmark for whether LLMs flatten contested scientific mechanisms into false consensus
Benchmark for multimodal contradiction and evidence reconciliation in biological research
Turn messy research progress into a decision-ready advisor meeting — an evidence-grounded Agent Skill for graduate researchers.
A research program on whether adaptive systems can become appropriately different without losing their capacity for justified correction.
CASCADE in scientific shape — paper, experiments, licence. Research era, March 2026. Kept public.
To associate your repository with the scientific-reasoning topic, visit your repo's landing page and select "manage topics."