Built a Vulkan Inference Engine that runs enormous MoE models on consumer AMD GPUs.
This story is from 2026-08-20. It is preserved in the archive; the latest stories are on the live feed.
I got the DeepSeek V4 Flash 0731 (284-billion-parameter) DeepSeek model with 1M context window running locally on a sub-$1,500 AMD PC (sub $1,000 if you buy used!), and open sourced it. On an RX 6700 XT with 12GB of VRAM, ordinary system RAM, and an NVMe SSD. I started this project because running…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-08-20 22:16 · r/LocalLLM
Built a Vulkan Inference Engine that runs enormous MoE models on consumer AMD GPUs.