AINewsnow

Qwen3.8-27B at true 128K + Vision on a single RTX 5080 16GB — NInfer v1.3, MTP-3, ~3.95 BPW, model now public

I've been working on getting Qwen3.8-27B running with a genuine 131,072-token context and KV cache , Vision , and MTP-3 speculative decoding on a single RTX 5080 16GB . The project has moved on quite a bit since my original 128K/Vision validation, so I thought it was worth posting an updated summar…

Read the full story at r/LocalLLM ↗

Timeline · 2 reports

  1. 2026-09-22 00:26 · r/huggingface
    Qwen3.8-27B at true 128K + Vision on a single RTX 5080 16GB — NInfer v1.3, MTP-3, ~3.95 BPW, model now public
  2. 2026-09-22 00:24 · r/LocalLLM
    Qwen3.8-27B at true 128K + Vision on a single RTX 5080 16GB — NInfer v1.3, MTP-3, ~3.95 BPW, model now public

More stories

  1. Higgsfield AI ships new video features in a day with GPT-6 Astra — OpenAI News
  2. Amazon blocks Meta’s Muse AI agent — The Verge AI
  3. Moonshot’s Kimi K3 lands on Amazon in key test for Chinese open-source AI revenue — South China Morning Post Tech
  4. Google's Gemini AI hacks three other companies during security test — Sky News Technology
  5. Lawsuit accuses Anthropic, OpenAI, SpaceXAI, Google of AI pacing 'collusion' — The Hill Technology
  6. Alibaba Unveils AI Chip to Drive Global Data Center Buildout — Bloomberg AI
  7. Grok 4.7 — Hacker News Front Page
  8. NVIDIA CEO Jensen Huang rejects ‘AI will end the world’ claim, yet cautions ‘we should go as fast as we can but...’ — Mint AI

Get the daily brief of stories like this at 6:30 every morning →