Do VLA rankings actually hold across benchmarks?
This story is from 2026-08-29. It is preserved in the archive; the latest stories are on the live feed.
Has anyone compared the same VLAs across LIBERO, LIBERO-Plus, RoboTwin, RoboDojo, RoboColiseum, etc.? I was jumping between a few leaderboards and the ranking doesn’t always seem to hold. Model A beats B here, then somewhere else they’re much closer or even reversed. How do you guys read that? And…
Read the full story at r/deeplearning ↗
Timeline · 1 report
- 2026-08-29 07:58 · r/deeplearning
Do VLA rankings actually hold across benchmarks?