trl
Here are 181 public repositories matching this topic...
HuggingEnvs — RL Environments 101: building and scaling RL environments in the age of LLMs
-
Updated
Sep 18, 2026 - Python
Notus is a collection of fine-tuned LLMs using SFT, DPO, SFT+DPO, and/or any other RLHF techniques, while always keeping a data-first approach
-
Updated
Jan 15, 2024 - Python
Agentic RL 零基础中文教程:24 章从概念到 GRPO 实战,含 TRL 最小可跑示例 | Beginner-friendly Agentic RL tutorial with hands-on GRPO project
-
Updated
Jun 3, 2026 - Python
Distill teacher chains-of-thought into a LoRA adapter via a strict boxed-answer format contract + two-phase Train→Nudge (silver-medal NVIDIA Nemotron reasoning recipe, as a tested library).
-
Updated
Jul 27, 2026 - Python
An implementation of GRPO for Unsloth's VLMs training
-
Updated
Aug 7, 2025 - Python
Code repository dedicated to experimenting and research with tiny reasoning language model
-
Updated
Nov 24, 2025 - Python
Various training, inference and validation code and results related to Open LLM's that were pretrained (full or partially) on the Dutch language.
-
Updated
Apr 9, 2024 - Jupyter Notebook
simpleR1: A Simple Framework for Training R1-like Models
-
Updated
Aug 12, 2025 - Python
An open-source application that estimates an Open Source Software Technology Readiness Level (OSSTRL) 1–9 from evidence that can be gathered automatically from a GitHub repository.
-
Updated
Sep 1, 2026 - Python
Training-data memorization auditor for fine-tuned LLMs — Trainer/TRL plugin, canary MIA + regurgitation audit, Apache-2.0
-
Updated
Aug 30, 2026 - Python
This project demonstrates the process of fine-tuning the Qwen2.5-3B-Instruct model using GRPO (Generalized Reward Policy Optimization) on the GSM8K dataset.
-
Updated
Apr 7, 2025 - Jupyter Notebook
使用trl、peft、transformers等库,实现对huggingface上模型的微调。
-
Updated
Mar 21, 2025 - Python
Different post-training techniques for LLMs, including: SFT, DPO and Online RL
-
Updated
Sep 5, 2025 - Python
Vibe Innovation und Vibe Coding Workshop mit GitHub Codespaces
-
Updated
Apr 28, 2026 - Shell
Build task models that replace frontier API calls
-
Updated
Sep 18, 2026 - TypeScript
Config-driven LLM fine-tuning with safety evaluation, EU AI Act compliance, 6 alignment methods, and one-command bundled quickstart templates.
-
Updated
Jul 21, 2026 - Python
Add this topic to your repo
To associate your repository with the trl topic, visit your repo's landing page and select "manage topics."