Memory-State Critic for Asymmetric Actor-Critic with Application to Vision-Based Pursuit-Evasion
arXiv:2610.03830v1 Announce Type: new Abstract: In partially observable Markov decision processes, the optimal policy generally depends on the history of observations and past actions. Asymmetric actor-critic methods have become popular to learn such policies when additional information, such as th…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-10-06 04:00 · arXiv cs.LG
Memory-State Critic for Asymmetric Actor-Critic with Application to Vision-Based Pursuit-Evasion