Red Hat Shrinks Nemotron 3.5 Lightning's 30B Agent Model by Half With FP8
Red Hat AI shipped an FP8 build of NVIDIA's hybrid Mamba-MoE Lightning model, halving memory while keeping the 1M-token agent workhorse on a single GPU.
Read the full story at AlphaSignal ↗
Timeline · 1 report
- 2026-10-03 22:00 · AlphaSignal
Red Hat Shrinks Nemotron 3.5 Lightning's 30B Agent Model by Half With FP8