Mapping the failure boundary of a Go1 locomotion policy: 6,400 rollouts, survival statistics, and a live interactive map
This story is from 2026-08-20. It is preserved in the archive; the latest stories are on the live feed.
We froze a Go1 joystick-locomotion policy (MuJoCo Playground, Brax PPO) and swept a 20×20 grid of floor friction against lateral push, 16 trials per cell, using Kaplan-Meier survival per condition since trials that survive the window have to be censored rather than counted as failures. Things inter…
Read the full story at r/reinforcementlearning ↗
Timeline · 1 report
- 2026-08-20 18:38 · r/reinforcementlearning
Mapping the failure boundary of a Go1 locomotion policy: 6,400 rollouts, survival statistics, and a live interactive map