AINewsnow

Qwen 3.8 27B Performance Benchmarks Emerge Across Diverse Hardware

This story is from 2026-08-21. It is preserved in the archive; the latest stories are on the live feed.

Users report token generation speeds and optimization techniques for Qwen 3.8 27B on AMD, Apple, and Nvidia hardware, including a 7 tok/s Apple Neural Engine build and a 2.26x speculative decoding speedup via DFlash 2.

Read the full story at r/LocalLLaMA ↗

Timeline · 6 reports

  1. 2026-08-22 20:41 · r/LocalLLaMA
    I benchmark DFlash 2 (PR build) in llama.cpp on Qwen 3.8 27B against all speculative methods for 3 days. 2.26x on 100 real coding prompts, 4.68x with one n-gram drafter on top. Up to 8x on specific cases.
  2. 2026-08-22 12:16 · r/LocalLLM
    Any of you running Qwen 3.8 27B on an RTX Pro 4000 SFF Blackwell?
  3. 2026-08-22 04:48 · r/LocalLLM
    people running Qwen 3.8 27B on apple silicon… whats your best token generation speed and how did you attain it?
  4. 2026-08-21 19:58 · r/LocalLLaMA
    Strix Halo (8060S / gfx1151), Qwen-3.8-27B @ Q8 and Q6 UD v3, up to 256K ctx, llama.cpp, DFlash2, vision, real workloads quality and steady performances, optimized recipes, ...
  5. 2026-08-21 12:02 · r/LocalLLM
    Running Qwen 3.8 27b FP16 on the Apple Neural Engine - 7 Watts of power to run a FP16 model @ 7 tok/s
  6. 2026-08-21 09:25 · r/LocalLLM
    Anyone running Qwen 3.8 27B Q3/Q4 on an RX 9060 XT 16GB using llama.cpp?

More stories

  1. Which major chatbot apps work with CarPlay? — Engadget
  2. Apple’s new chief executive built up to unveiling the ideal AI device, then said it was the iPhone. — The Next Web
  3. A closer look at the upcoming Siri AI-powered home hub, a key pillar of Apple's strategy for the home; sources: Apple started cutting Fitness+ staff (Mark Gurman/Bloomberg) — Techmeme
  4. Gemini Joins the Hacker Club — Wall Street Journal Technology
  5. Apple’s Home AI Hub Details; Apple Fitness+ Layoffs and iPhone Duo Apple Pencil — Bloomberg AI
  6. ComfyUI on Apple Silicon: no MLX, no fp8, 600-second kernel builds. So I built my own launcher — a personal project I'm sharing in case it helps someone. — r/comfyui
  7. [Release] Nirvana Code: A single-binary Rust LLM engine built from the metal up for Apple Silicon (Metal 3, Persistent Prefix Cache, Speculative Decoding, Dual GGUF + MLX) — r/LocalLLM
  8. He’s the Face of AI Doomsday Fears — Wall Street Journal Technology

Get the daily brief of stories like this at 6:30 every morning →