SGLang outputs endless repetition on NVFP4 models: the FP8 lm_head bug
Your NVFP4 model serves fine on vLLM but outputs an endlessly repeated phrase on SGLang, from the very first token, even on a trivial prompt. The response content comes back empty, every request ends
conatusai.hashnode.dev3 min read