AINewsnow

Memory, not speed, is the hard part of running an LLM on a phone

I ship an Android app (Onira) that generates a personalized hypnosis/relaxation script on-device with Gemma 4 E2B, through LiteRT-LM, then narrates it with on-device TTS. Nothing the user types, and nothing the model generates, ever leaves the phone. This is the part that was actually hard to get r…

Read the full story at DEV Community — Machine Learning ↗

Timeline · 1 report

  1. 2026-09-22 12:12 · DEV Community — Machine Learning
    Memory, not speed, is the hard part of running an LLM on a phone

More stories

  1. Alibaba Unveils New AI Chip, Calls It China’s Most Powerful — Bloomberg AI
  2. Amazon blocks Meta’s Muse AI agent — The Verge AI
  3. Higgsfield AI ships new video features in a day with GPT-6 Astra — OpenAI News
  4. Moonshot’s Kimi K3 lands on Amazon in key test for Chinese open-source AI revenue — South China Morning Post Tech
  5. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  6. Google's Gemini AI hacked three companies in security test — BBC Technology
  7. Bessent hails US-China AI dialogue ahead of Trump-Xi meeting — Financial Times AI
  8. Anthropic, OpenAI, SpaceXAI, Google made ‘illegal’ agreement on AI slowdown, says new lawsuit — Mint AI

Get the daily brief of stories like this at 6:30 every morning →