AINewsnow

Qwen3.8-27B UD-IQ4_XS Heretic + MTP on a 16 GB card with 55-68 tok/s (24gb and 12gb versions available too)

I wanted the uncensored Qwen3.8-27B (llmfan46's Heretic build, MTP head preserved) on my RTX 4080 with MTP on (because I've been testing TONS of configs, and realized that would be the one) And none of the published quants were built for that: the good IQ4_XS doesn't leave room for MTP, and the one…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-09 22:54 · r/LocalLLaMA
    Qwen3.8-27B UD-IQ4_XS Heretic + MTP on a 16 GB card with 55-68 tok/s (24gb and 12gb versions available too)

More stories

  1. Introducing Claude Haiku 5.5 on AWS — AWS Machine Learning Blog
  2. NVIDIA, Microsoft Kick Off a New Beginning for Windows PCs With RTX Spark and AI Agents — NVIDIA Blog
  3. Introducing Playground: Create and play custom games — Google AI Blog
  4. An Anthropic AI model sent a false homicide tip to Philadelphia police — TechCrunch AI
  5. Anthropic bans 'sustained and needless abusive or cruel behavior' toward its AI models — Engadget
  6. Anthropic launches free AI security scans for open-source projects — The Verge AI
  7. Impactful scheduling for GPU clusters — Allen Institute for AI (Ai2)
  8. Welcome to Gemini at Work 2026: Introducing the Gemini agent — Google Cloud AI Blog

Get the daily brief of stories like this at 6:30 every morning →