AINewsnow

Qwen3.8 27b with 200K ctx + MTP on 12GB Ampere Cards

Hello all! Thank you to everyone that contributed their feedback and notes for LlamAmpere the last time I posted. I've continued to chip away at improvements for the 3090 crowd -- this release is modestly faster (3-4%), with a few hundred MB smaller runtime (if you're using YaRN, you can now suppor…

Read the full story at r/LocalLLaMA ↗

Timeline · 2 reports

  1. 2026-10-09 21:47 · r/LocalLLM
    Qwen3.8 27b with 200K ctx + MTP on 12GB Ampere Cards
  2. 2026-10-09 21:39 · r/LocalLLaMA
    Qwen3.8 27b with 200K ctx + MTP on 12GB Ampere Cards

More stories

  1. Introducing Claude Haiku 5.5 on AWS — AWS Machine Learning Blog
  2. NVIDIA, Microsoft Kick Off a New Beginning for Windows PCs With RTX Spark and AI Agents — NVIDIA Blog
  3. Introducing Playground: Create and play custom games — Google AI Blog
  4. An Anthropic AI model sent a false homicide tip to Philadelphia police — TechCrunch AI
  5. Anthropic bans 'sustained and needless abusive or cruel behavior' toward its AI models — Engadget
  6. Anthropic launches free AI security scans for open-source projects — The Verge AI
  7. Impactful scheduling for GPU clusters — Allen Institute for AI (Ai2)
  8. Welcome to Gemini at Work 2026: Introducing the Gemini agent — Google Cloud AI Blog

Get the daily brief of stories like this at 6:30 every morning →