HauhauCS' Qwen3.8 27B FastMTP is real
This story is from 2026-08-20. It is preserved in the archive; the latest stories are on the live feed.
Yes, DFlash and EAGLE-3 did reduced-vocab draft heads first — but this is the only Qwen3.8 27B repackaging with the MTP head separated into a sidecar GGUF and its vocab trimmed: output.weight [5120, 32768] plus a d2t tensor mapping draft rows back to the full 248,320-token vocab. The draft's logit…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-08-20 04:33 · r/LocalLLM
HauhauCS' Qwen3.8 27B FastMTP is real