Qwen3.5 0.8B on CPU
This story is from 2026-09-07. It is preserved in the archive; the latest stories are on the live feed.
Since the Qwen3.5 0.8B model is an interesting one for small specialized fine tunes, I was curious how fast it can run on CPUs. Why CPUs? Mainly because I want to use it as a local dictation cleanup model when I'm using the GPU for something else. Over the weekend, I let Codex build a small C++ eng…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-07 20:09 · r/LocalLLaMA
Qwen3.5 0.8B on CPU