RTX PRO 4000 Blackwell (24GB) - nvfp4 vs GGUF for long context (100k+)?
Hello everyone, I’m currently running an RTX PRO 4000 Blackwell SFF (24GB VRAM). I'm trying to optimize for a minimum context window of ~100k tokens using a Qwen 3.8-27B model. Current Setup: Running vLLM with the GGUF plugin using Qwen3.8-27B-UD-Q4_K_XL.gguf . This allows me to hit a context size…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-21 07:54 · r/LocalLLM
RTX PRO 4000 Blackwell (24GB) - nvfp4 vs GGUF for long context (100k+)?