REFACTOR-VLA: Unsupervised Library Learning of Typed Motor Programs
This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.
Most current vision-language-action (VLA) models—such as OpenVLA, π0, RT-2, and RDT-1B—are “monolithic.” This means they generate raw motor commands or very short sequences of actions, without organizing behaviors into reusable, well-defined abstractions. As a result, these models perform poorly on…
Read the full story at Apple Machine Learning Research ↗
Timeline · 1 report
- 2026-09-02 00:00 · Apple Machine Learning Research
REFACTOR-VLA: Unsupervised Library Learning of Typed Motor Programs