Made a tool that tells you which GGUF quants will fit your GPU/Mac, with the llama.cpp command to run them
This story is from 2026-09-17. It is preserved in the archive; the latest stories are on the live feed.
It's hard to tell from a model page whether a quant will fit once you add context, KV cache, and whatever layers or experts end up on CPU. So I made this: https://huggingface.co/spaces/LocalLLaMA/local-model-explorer Enter your GPU(s) and RAM, or a Mac / Strix Halo / DGX Spark and its unified memor…
Read the full story at r/LocalLLM ↗
Timeline · 2 reports
- 2026-09-18 01:27 · r/huggingface
Made a tool that tells you which GGUF quants will fit your GPU/Mac, with the llama.cpp command to run them - 2026-09-17 03:17 · r/LocalLLM
Made a tool that tells you which GGUF quants will fit your GPU/Mac, with the llama.cpp command to run them