What's your personal take on running dense models for code at nvfp4 vs fp8. I don't even bother running heavily quantized models unless its on brand new hw.
This story is from 2026-09-15. It is preserved in the archive; the latest stories are on the live feed.
Spent the last four months (learning) deploying nvidia nim models on an air gapped OCP environment ( 40X L40s cards across 10 dell r870's). The guardrails and limited model profiles available on ngc annoyed me at first until I realized you can do damn near anything with vllm and hugging face model…
Read the full story at r/LocalLLM ↗