Pinned Loading
-
AI-Can-Learn-Scientific-Taste
AI-Can-Learn-Scientific-Taste PublicWe propose Reinforcement Learning from Community Feedback (RLCF), a training paradigm that uses large-scale community signals as supervision, and formulate scientific taste learning as a preference…
-
Thinking-with-Video
Thinking-with-Video PublicWe introduce 'Thinking with Video', a new paradigm leveraging video generation for multimodal reasoning. Our VideoThinkBench shows that Sora-2 surpasses GPT5 by 10% on eyeballing puzzles and reache…
-
pjlab-sys4nlp/llama-moe
pjlab-sys4nlp/llama-moe Public⛷️ LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training (EMNLP 2024)
-
Awesome-Agent-RL
Awesome-Agent-RL PublicA curated list of awesome resources about reward construction for AI agents. This repository covers cutting-edge research, and practical guides on defining and collecting rewards to build more inte…
If the problem persists, check the GitHub status page or contact support.
