AINewsnow

My simulator said the change was worth 1.8x. On real hardware it was 1.05x (t=0.73). Some notes from deploying PPO against a live mobile game.

I trained a PPO agent to play an idle tower-defense mobile game — it decides which of 13 upgrades to buy each wave. No API and no memory access: observations come from OCR (tesseract plus per-glyph template matching) on the screen, and actions are real mouse clicks. It's been running against the li…

Read the full story at r/reinforcementlearning ↗

Timeline · 1 report

  1. 2026-09-21 13:30 · r/reinforcementlearning
    My simulator said the change was worth 1.8x. On real hardware it was 1.05x (t=0.73). Some notes from deploying PPO against a live mobile game.

More stories

  1. Higgsfield AI ships new video features in a day with GPT-6 Astra — OpenAI News
  2. Amazon blocks Meta’s Muse AI agent — The Verge AI
  3. Moonshot’s Kimi K3 lands on Amazon in key test for Chinese open-source AI revenue — South China Morning Post Tech
  4. Google's Gemini AI hacked three companies in security test — BBC Technology
  5. British Columbia Sues OpenAI Over Canada Mass Shooting Warning Failure — Bloomberg AI
  6. Lawsuit accuses Anthropic, OpenAI, SpaceXAI, Google of AI pacing 'collusion' — The Hill Technology
  7. Grok 4.7 — Hacker News Front Page
  8. Ahead of Sam Altman's UN address, OpenAI proposes new ways to track AI misalignment risks — Axios AI+

Get the daily brief of stories like this at 6:30 every morning →