Robot Learning | Reinforcement Learning | Humanoids
Incoming M.Sc. student in Computer Science at Karlsruhe Institute of Technology (KIT). Computer Engineering graduate, Çukurova University. I work on RL fine tuning of vision language action policies on consumer hardware, and hierarchical control for humanoid robots.
mehmet.yardimci@student.kit.edu | LinkedIn | Google Scholar | ORCID | Portfolio
Uncertainty-Gated Exploration Noise Suppresses Task Collapse in Online RL Fine-Tuning of a Flow-Matching
Vision-Language-Action Policy
M. T. Yardımcı, Y. E. Çoğurcu. arXiv:2609.28838, 2026. Submitted to ICLR 2027;
also submitted to the CoRL 2026 Workshop on Continually Self-Improving Robots (non-archival).
Critic Architecture Matters: Dual vs. Unified Critics for Humanoid Loco-Manipulation
M. T. Yardımcı. ICRA 2026 Workshop on Reinforcement Learning in the Era of Imitation Learning (RL4IL).
arXiv:2606.11891
Benchmarking Local Path Planners in ROS using the BARN Dataset
M. T. Yardımcı, Y. E. Çoğurcu. Under review.
Code
Online RL fine tuning of a flow matching VLA policy on a single 12GB consumer GPU, following the pi_RL approach. A verification first substrate: what trains is a logged fact, not an assumption. Project page | Paper
RL to IL to VLA pipeline for the Unitree G1: expert demonstrations from trained RL policies, distilled into end to end visuomotor policies (ACT, Diffusion Policy, GR00T N1.6).
VLM task planning (Qwen3 VL) over PPO locomotion and arm policies on the Unitree G1: walk, reach, grasp, drawer, pick and place.
Multi stage PPO curriculum for whole body G1 locomotion in Isaac Lab: velocity tracking, terrain, torso, arm coordination. Basis of the dual vs. unified critic study (paper).
Open to research collaborations in humanoid robotics and robot learning.
