Does anybody want to give this blueberry ordering problem a shot?
This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.
I created a simulator, filter and controller for fresh produce ordering under varying observation scenarios. I showed that richer observations lead to better belief accuracy. But my controller sucks! It wasn't able to translate better beliefs into more profit. I think that RL would be a good fit he…
Read the full story at r/reinforcementlearning ↗
Timeline · 1 report
- 2026-09-04 04:35 · r/reinforcementlearning
Does anybody want to give this blueberry ordering problem a shot?