Skip to content
View mturan33's full-sized avatar

Highlights

  • Pro

Block or report mturan33

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
mturan33/README.md

Mehmet Turan Yardımcı

Robot Learning | Reinforcement Learning | Humanoids

Incoming M.Sc. student in Computer Science at Karlsruhe Institute of Technology (KIT). Computer Engineering graduate, Çukurova University. I work on RL fine tuning of vision language action policies on consumer hardware, and hierarchical control for humanoid robots.

mehmet.yardimci@student.kit.edu | LinkedIn | Google Scholar | ORCID | Portfolio


Papers

Uncertainty-Gated Exploration Noise Suppresses Task Collapse in Online RL Fine-Tuning of a Flow-Matching Vision-Language-Action Policy
M. T. Yardımcı, Y. E. Çoğurcu. arXiv:2609.28838, 2026. Submitted to ICLR 2027; also submitted to the CoRL 2026 Workshop on Continually Self-Improving Robots (non-archival).

Critic Architecture Matters: Dual vs. Unified Critics for Humanoid Loco-Manipulation
M. T. Yardımcı. ICRA 2026 Workshop on Reinforcement Learning in the Era of Imitation Learning (RL4IL). arXiv:2606.11891

Benchmarking Local Path Planners in ROS using the BARN Dataset
M. T. Yardımcı, Y. E. Çoğurcu. Under review. Code


Projects

Online RL fine tuning of a flow matching VLA policy on a single 12GB consumer GPU, following the pi_RL approach. A verification first substrate: what trains is a logged fact, not an assumption. Project page | Paper

G1 Vision Language Action pipeline (in progress, repository not public yet)

RL to IL to VLA pipeline for the Unitree G1: expert demonstrations from trained RL policies, distilled into end to end visuomotor policies (ACT, Diffusion Policy, GR00T N1.6).

VLM task planning (Qwen3 VL) over PPO locomotion and arm policies on the Unitree G1: walk, reach, grasp, drawer, pick and place.

Multi stage PPO curriculum for whole body G1 locomotion in Isaac Lab: velocity tracking, terrain, torso, arm coordination. Basis of the dual vs. unified critic study (paper).


Open to research collaborations in humanoid robotics and robot learning.

Pinned Loading

  1. isaac-g1-hierarchical isaac-g1-hierarchical Public

    VLM-RL Hierarchical Loco-Manipulation For Long-Horizon Tasks With G1 robot in Isaac Lab/Sim

    Python 19 1

  2. isaac-g1-ulc isaac-g1-ulc Public

    Low Level RL Controller for G1

    Python 18 1

  3. smolvla_flow_rl smolvla_flow_rl Public

    Online reinforcement learning fine tuning for a flow matching vision language action policy, on a single 12GB consumer GPU.

    Python 3

  4. mujoco-ant-ppo mujoco-ant-ppo Public

    Training a MuJoCo Ant agent to walk using PPO from scratch.

    Python 6

  5. isaaclab-anymal-locomotion isaaclab-anymal-locomotion Public

    A legged locomotion project

    Python 3

  6. benchmark-local-path-planners-barn-challenge benchmark-local-path-planners-barn-challenge Public

    A Framework for BARN of classical and learning-based local path planners in BARN Challenge navigation benchmark.

    HTML 1