AINewsnow

Qwen3.8 27B int4 with Dflash2 at 165t/s and 18M kv cache pool on dual 3090

This story is from 2026-08-28. It is preserved in the archive; the latest stories are on the live feed.

Hardware - 2× 3090 24GB - Ryzen 5 5600 - DDR4 96GB - WD_BLACK SN850X 2TB NVMe Software - vLLM 0.28.0 + patches (see recipe) - LMCache 0.5.4rc5 (54GB RAM L1 + 1.5TB NVMe L2) Model - cyankiwi/Qwen3.8-27B-AWQ-INT4 - DFlash2-W4A16 - fp8 KV - GPU KV cache size: 417,610 tokens The W4A16 is quant by me, y…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-08-28 16:44 · r/LocalLLaMA
    Qwen3.8 27B int4 with Dflash2 at 165t/s and 18M kv cache pool on dual 3090

More stories

  1. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
  2. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  3. NVIDIA CEO Jensen Huang rejects ‘AI will end the world’ claim, yet cautions ‘we should go as fast as we can but...’ — Mint AI
  4. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
  5. AI skills — r/AI_Agents
  6. AI hallucination of Chinese nuclear components almost led to US military attack — Ars Technica AI
  7. TypeSafe AI Releases Jev: A System One Model That Returns Typed, Calibrated Decisions Instead of Text — MarkTechPost
  8. AI's role in building AI surging? Anthropic says Claude now leads 26% of its R&D — Mint AI

Get the daily brief of stories like this at 6:30 every morning →