Made a tool that tells you which GGUF quants will fit your GPU/Mac, with the llama.cpp command to run them
It's hard to tell from a model page whether a quant will fit once you add context, KV cache, and whatever layers or experts end up on CPU. So I made this: https://huggingface.co/spaces/LocalLLaMA/local-model-explorer Enter your GPU(s) and RAM, or a Mac / Strix Halo / DGX Spark and its unified memor…
Read the full story at r/huggingface ↗
Timeline · 1 report
- 2026-09-18 01:27 · r/huggingface
Made a tool that tells you which GGUF quants will fit your GPU/Mac, with the llama.cpp command to run them