Issue evidence · Serve · vllm-project/vllm
GitHub issue: https://github.com/vllm-project/vllm/issues/60998
Reproduced
A float max_tokens passes SamplingParams validation, then the engine fails to decode it and generate() never returns.
Ran the issue's script with Qwen/Qwen2.5-0.5B-Instruct on an RTX 4090 in the vLLM 0.30.0 image, capped at 240 seconds.
The failure is in the offline LLM API, so Badgr cannot reject the value before the engine; the receipt shows its response check catching the hang.
Checked 2026-10-11 in Badgr’s local development environment.
← All issue evidence