Welcome 🖐️ I'm a research engineer at Nvidia.
I make Agents good - from data, rl infra, algo, recipe to harness, collectively optimized as a single problem.
Proud to Present as First Author >>
🎨 Skill2Env | Superintelligence from and for humanity. [Code] [Dataset]
🔳 Linex | The One infra for Agentic RL. [Code]
⚡ FlashREINFOCE | Solving instability from async RL policy drifts and token credit mis-assignent. [Paper]
⭐ Polar | Agentic RL on ANY harnesses at scale (first in its field). [Code] [Paper]
🧠 NanoGPX | Clean collection of modern LLM architectures (RoPE, GQA, RMSNorm, MoE, SSM, etc.) in nanoGPT style. [Code]
🤖 Gentopia & GentPool | An Agent [Framework] & [Platform].
🚀 ReWOO | Token-efficient harness via decoupling reasoning from observation. [Code] [Paper]




