AirLLM's layer-by-layer trick, but on CPU: running a 17.66 GB model on a 16 GB ARM board in pure C
Disclosure: I built Kestrel-LLM, the engine in this post. Solo side project, nobody's paying me to write this up. TL;DR: I've got a 16 GB ARM board (MemTotal reports 15.6 GiB) and a 17.66 GB q4 weight file for Qwen3-30B, so the model doesn't fit. What I did was take AirLLM's trick of holding one la…
Read the full story at r/machinelearningnews ↗
Timeline · 1 report
- 2026-09-30 07:21 · r/machinelearningnews
AirLLM's layer-by-layer trick, but on CPU: running a 17.66 GB model on a 16 GB ARM board in pure C