-
Notifications
You must be signed in to change notification settings - Fork 90
Pull requests: onejune2018/Awesome-LLM-Eval
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Add: AdverseMed-500 (Domain/Healthcare) and AI Scientist v2 Biomedical Audit (Agent-Capabilities)
#85
opened Aug 22, 2026 by
calnugget
Loading…
Add demoparity tool for demographic auditing of LLMs
#84
opened Aug 12, 2026 by
cindysteward
Loading…
Add StructEval (structured-output generation benchmark, TMLR 2025) to Benchmarks → General
#82
opened Aug 4, 2026 by
reacher-z
Loading…
Add MMESGBench (multimodal ESG benchmark) to Datasets/Benchmarks → Domain
#62
opened Jun 22, 2026 by
ChaoYue0307
Loading…
Add ESGenius (ESG & sustainability) to Datasets/Benchmarks → Domain
#61
opened Jun 22, 2026 by
ChaoYue0307
Loading…
Add Implicit Behavioral Alignment of Language Agents (EMNLP 2025)
#57
opened Jun 14, 2026 by
wangyz1999
Loading…
Add Prompt Evaluator — web-based prompt and workflow QA tool
#54
opened Jun 6, 2026 by
ariangibson
Loading…
Previous Next
ProTip!
Adding no:label will show everything without a label.