AINewsnow

I made a custom llama.cpp build optimized for 7900xtx (one or two). for qwen 3.8 next and 27B. includes optimizations for PciE x4 and tensor parallel. read inside! (no AI slop)

This story is from 2026-09-08. It is preserved in the archive; the latest stories are on the live feed.

I found a lot of room on the table for these cards so I decided to make a specialized build to squeeze all I could. first The results: qwen 3.8 next Q3_K_XL: 920tk/s pp8192 (2 cards, ram offload), 24/27 tk/s on prose, 40+ tk/s on code with MTP but without MoE expert cache (which IS included if yuo…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-08 08:22 · r/LocalLLaMA
    I made a custom llama.cpp build optimized for 7900xtx (one or two). for qwen 3.8 next and 27B. includes optimizations for PciE x4 and tensor parallel. read inside! (no AI slop)

More stories

  1. qwen4exp: add hc ops by am17an · Pull Request #28901 · ggml-org/llama.cpp — r/LocalLLaMA
  2. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  3. Intel releases OpenVINO 2026.4 — r/LocalLLaMA
  4. dual 7900 xtx - some guy made a pretty optimized fork of lamacpp optimized for this setup Qwen 3.8 Q8 at 82 tokens / seconds decode — r/LocalLLaMA
  5. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
  6. Post-training image models for fandom — Character.AI Blog
  7. Testing Qwen 3.8 27B running locally on a single 5090 — r/LocalLLM
  8. US government website used Chinese model the FBI called "malicious" — Ars Technica AI

Get the daily brief of stories like this at 6:30 every morning →