Skip to content
View billxbf's full-sized avatar
☕
☕

Highlights

  • Pro

Organizations

@Gentopia-AI

Block or report billxbf

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
billxbf/README.md

Welcome 🖐️ I'm a research engineer at Nvidia.

I make Agents good - from data, rl infra, algo, recipe to harness, collectively optimized as a single problem.


Proud to Present as First Author >>

🎨 Skill2Env | Superintelligence from and for humanity. [Code] [Dataset]

🔳 Linex | The One infra for Agentic RL. [Code]

⚡ FlashREINFOCE | Solving instability from async RL policy drifts and token credit mis-assignent. [Paper]

⭐ Polar | Agentic RL on ANY harnesses at scale (first in its field). [Code] [Paper]

🧠 NanoGPX | Clean collection of modern LLM architectures (RoPE, GQA, RMSNorm, MoE, SSM, etc.) in nanoGPT style. [Code]

🤖 Gentopia & GentPool | An Agent [Framework] & [Platform].

🚀 ReWOO | Token-efficient harness via decoupling reasoning from observation. [Code] [Paper]

Pinned Loading

  1. NVIDIA-NeMo/ProRL-Agent-Server NVIDIA-NeMo/ProRL-Agent-Server Public

    Agentic RL on Any Harness at Scale

    Python 843 92

  2. NVlabs/Skill2Env NVlabs/Skill2Env Public

    Democratizing Collective Intelligence

    Python 99 14

  3. Linex Linex Public

    The One & Minimal Framework for Agentic RL.

    Python 2

  4. ReWOO ReWOO Public

    Decoupling Reasoning from Observations for Efficient Augmented Language Models

    Python 943 84

  5. Gentopia-AI/Gentopia Gentopia-AI/Gentopia Public

    Build Hierarchical Autonomous Agents through Config. Collaborative Growth of Specialized Agents.

    Python 328 42

  6. nanoGPX nanoGPX Public

    Forked from karpathy/nanoGPT

    Clean implementation of modern LLM recipes in nanoGPT style (RoPE, GQA, RMSNorm, MoE, SSM, etc.)

    Python 3