Update on XTLLM: 8.5 tok/s Qwen3.8 Flash Next & 20 tok/s on Longcat (69B).
This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.
Quick background: I recently posted about XTLLM here , a project I started to help with a personal problem I was having with Local LLMs on my hardware (and windows). XTLLM is a pure Vulkan inference engine designed to run large MoE models on consumer AMD GPUs on Windows. I have the RX 6900 XT (12gb…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-04 19:07 · r/LocalLLM
Update on XTLLM: 8.5 tok/s Qwen3.8 Flash Next & 20 tok/s on Longcat (69B).