Training a robot arm to pick things up with a vision-language model + RL
I've been building a robot arm project where a vision-language model looks at a camera image and picks a subgoal (like "align over the block" or "close the gripper"), which then feeds into a small continuous-control policy that actually moves the arm. Stack: MuJoCo for simulation, Qwen2-VL-2B as th…
Read the full story at r/reinforcementlearning ↗
Timeline · 1 report
- 2026-09-18 11:17 · r/reinforcementlearning
Training a robot arm to pick things up with a vision-language model + RL