AINewsnow

What's your personal take on running dense models for code at nvfp4 vs fp8. I don't even bother running heavily quantized models unless its on brand new hw.

This story is from 2026-09-15. It is preserved in the archive; the latest stories are on the live feed.

Spent the last four months (learning) deploying nvidia nim models on an air gapped OCP environment ( 40X L40s cards across 10 dell r870's). The guardrails and limited model profiles available on ngc annoyed me at first until I realized you can do damn near anything with vllm and hugging face model…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-15 13:13 · r/LocalLLM
    What's your personal take on running dense models for code at nvfp4 vs fp8. I don't even bother running heavily quantized models unless its on brand new hw.

More stories

  1. Amid growing AI fears, King Charles meets with industry leaders in Scotland — NPR Technology
  2. Deploy Hugging Face models on Amazon SageMaker AI with coding agents — AWS Machine Learning Blog
  3. King Charles to press Nvidia, OpenAI, Anthropic leaders on AI safety at summit — CNBC Technology
  4. Huawei details AI accelerator roadmap, pulls in next-generation Ascend NPUs by several quarters — FP4 performance of the Ascend 960PR doubles expectations — Tom's Hardware
  5. Flyweight: open-source C++/CUDA engine for running MoE models bigger than your VRAM on one GPU + system RAM. First PyPI release, looking for contributors. — r/LocalLLaMA
  6. A quick Minimax H3 news round-up - 17th September 2026 — r/comfyui
  7. Deployed Qwen 3.6 35B A3B on a single DGX Spark supporting 12 concurrent users at 262K context. Are there better ways to optimize this? — r/LocalLLM
  8. Beginner confused about Ollama vs LM Studio vs llama.cpp vs vLLM vs Unsloth — can someone explain? — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →