Qwen3.8-27B on a single 3090: 140 tok/s on code with a custom megakernel
I've been using Claude Opus 5.5 to speed up Qwen3.8-27B on my PC (rtx 3090), it wrote a CUDA megakernel that is 1.4-1.9x faster than llama.cpp depending on the task/context length. The results and code are below: Results (same hardware, my megakernel vs llama.cpp with MTP): - Writing code: 140 tok/…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-10 13:01 · r/LocalLLaMA
Qwen3.8-27B on a single 3090: 140 tok/s on code with a custom megakernel