AINewsnow

Finally!

So I have been tinkering. I think I got this V100 32gb dialed in (PG500-216 @ 185w, custom compile) Has Qwen3.8:27b - UD-Q4-KM (mtp) at 50 ish tps decode on as you can see a long string. Prefill dives fast on this, anyone have pointers on squeezing more prefill? Batch size at 4096 ubatch at 2048. (…

Read the full story at r/LocalLLM ↗

Timeline · 2 reports

  1. 2026-10-09 19:44 · r/GeminiAI
    It’s finally here!!! 😮
  2. 2026-10-08 18:51 · r/LocalLLM
    Finally!

More stories

  1. When will gemini 4 release? — r/Bard
  2. Welcome to Gemini at Work 2026: Introducing the Gemini agent — Google Cloud AI Blog
  3. Google Cloud introduces Gemini agent to change enterprise work — SiliconANGLE AI
  4. Is Gemini Pro model down? — r/GeminiAI
  5. The new update 😔 — r/GeminiAI
  6. Whatever happened to BABA is AI from 2024? [D] — r/MachineLearning
  7. Gemini 4 is coming today (or tmr depending on your time zone.) — r/GeminiAI
  8. What's the best AI agent orchestration setup in 2026? Hermes, Pi, OpenCode, Claude Code, or something else? — r/AI_Agents

Get the daily brief of stories like this at 6:30 every morning →