RL agent's internal state, fake-world detection goes from 50% to 73% after bad physics makes it miss food
Trained a RL agent to find food, then put it in a copy of its world with one physics rule changed (ground grip). At first, a probe on its internal state could only tell it was in the fake world about 50% of the time. Once the bad grip made it slip and miss a food, that jumped to about 73%. Nobody t…
Read the full story at r/reinforcementlearning ↗
Timeline · 1 report
- 2026-09-22 00:39 · r/reinforcementlearning
RL agent's internal state, fake-world detection goes from 50% to 73% after bad physics makes it miss food