Where Reef Infra draws the boundary between an RL method and the service running it
Changing a reward or batching rule should not require redesigning how inference records survive a restart. Conversely, a service framework shouldn't quietly decide what counts as a useful update to your policy. Reef Infra's weight recipes expose that boundary fairly explicitly. The processor assemb…
Read the full story at r/reinforcementlearning ↗
Timeline · 1 report
- 2026-09-20 14:35 · r/reinforcementlearning
Where Reef Infra draws the boundary between an RL method and the service running it