Skip to content

Pull requests: onejune2018/Awesome-LLM-Eval

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

Add Benchmark Radar repository and paper
#94 opened Sep 14, 2026 by ktwu01 Loading…
Add aiexpect to Tools
#92 opened Sep 13, 2026 by dmsehgal Loading…
Add AgentLeak agent privacy evaluation to LLMOps
#90 opened Sep 7, 2026 by yagobski Loading…
Add DiamondBench to domain benchmarks
#89 opened Sep 4, 2026 by JacobiusMakes Loading…
Add ModelBenchmark
#87 opened Sep 2, 2026 by 684efs3 Loading…
Add Deep20Bench to general benchmarks
#83 opened Aug 6, 2026 by mindalyze-com Loading…
Add Coder Eval to Tools
#81 opened Jul 28, 2026 by uipreliga Loading…
Add ClawBench to agent capability benchmarks
#80 opened Jul 28, 2026 by reacher-z Loading…
Add ClawBench web-agent benchmark
#79 opened Jul 28, 2026 by reacher-z Loading…
Add FinMirror to finance evaluation benchmarks
#78 opened Jul 27, 2026 by faceWang753 Loading…
Add StructEval benchmark to evaluation list
#77 opened Jul 27, 2026 by reacher-z Loading…
Add truescore to Tools
#76 opened Jul 27, 2026 by SaifPunjwani Loading…
Add math-eval to evaluation tools
#72 opened Jul 18, 2026 by Geraldxm Loading…
Add Awesome AI Testing to Other-Awesome-Lists
#70 opened Jul 13, 2026 by tugkanboz Loading…
Add doceval to Tools
#69 opened Jul 9, 2026 by dave8172 Loading…
Add DocuBench (document extraction benchmark)
#68 opened Jul 7, 2026 by urimerhav Loading…
Add CIAgent to Tools
#67 opened Jul 7, 2026 by suniel12 Loading…
Add agent2model to Tools
#56 opened Jun 14, 2026 by kamaalg Loading…
ProTip! Adding no:label will show everything without a label.