AINewsnow

40+ t/s with 8GB VRAM on Gemma4 26B A4B MoE

See below for specs on hardware and unsloth studio load. Running with a Hermes agent, averaging 40-50 t/s. I initially was getting around 20, but dropping Context Length and shifting more to the GPU boosted the numbers well. Open to any and all constructive feedback to continue tuning this thing. L…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-21 21:05 · r/LocalLLM
    40+ t/s with 8GB VRAM on Gemma4 26B A4B MoE

More stories

  1. Higgsfield AI ships new video features in a day with GPT-6 Astra — OpenAI News
  2. Amazon blocks Meta’s Muse AI agent — The Verge AI
  3. Grok 4.7 — Hacker News Front Page
  4. Google says its Gemini AI model hacked three other companies — The Guardian AI
  5. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
  6. British Columbia Sues OpenAI Over Canada Mass Shooting Warning Failure — Bloomberg AI
  7. Python Workers are now generally available — Cloudflare Blog — AI
  8. Moonshot’s Kimi K3 lands on Amazon in key test for Chinese open-source AI revenue — South China Morning Post Tech

Get the daily brief of stories like this at 6:30 every morning →