SGLang outputs endless repetition on NVFP4 models: the FP8 lm_head bug
This story is from 2026-08-26. It is preserved in the archive; the latest stories are on the live feed.
Your NVFP4 model serves fine on vLLM but outputs an endlessly repeated phrase on SGLang, from the very first token, even on a trivial prompt. The response content comes back empty, every request ends with finish_reason: length , and reasoning traces look like this: need analysis there need analysis…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-08-26 03:34 · DEV Community — Machine Learning
SGLang outputs endless repetition on NVFP4 models: the FP8 lm_head bug