[ICLR 2026] A Framework for LLM-based Multi-Agent Reinforced Training and Inference
-
Updated
Aug 20, 2026 - Python
[ICLR 2026] A Framework for LLM-based Multi-Agent Reinforced Training and Inference
RL training environments with verifiable rewards for coding agents. Works with TRL, Unsloth, verl, OpenRLHF.
Exact-oracle conformance tests for turn-level credit assignment in agentic RL (GRPO, RLOO, GAE, GiGPO; verl, TRL, OpenRLHF).
A list of uv environments templates for LLM development.
Multi-turn Agent RL training in OpenRLHF, an AgentFlow reimplementation.
RLHF Annotation Studio — Web-based tool for collecting human preference data to train LLMs via Reinforcement Learning from Human Feedback (RLHF). Compare responses side-by-side, capture preferences, and export JSONL for reward model training.
🌐 Streamline LLM development with ready-to-use environment templates for efficient setup and deployment.
To associate your repository with the openrlhf topic, visit your repo's landing page and select "manage topics."