Lost RL beginner trying to use PPO to fine tune a deterministic model
This story is from 2026-09-03. It is preserved in the archive; the latest stories are on the live feed.
I am working on training an agent to play Street Fighter III. This is my first time using RL for anything and even if the ideas mostly make sense to me, I often feel quite lost when it comes to implementing or modifying the algorithms themselves. The process I've been trying to follow is inspired b…
Read the full story at r/reinforcementlearning ↗
Timeline · 1 report
- 2026-09-03 14:12 · r/reinforcementlearning
Lost RL beginner trying to use PPO to fine tune a deterministic model