Issue evidence · Serve · vllm-project/vllm
GitHub issue: https://github.com/vllm-project/vllm/issues/60981
Reproduced, Badgr fix shipped
A speculative-decoding schedule on the ngram_gpu method kills the engine once the scheduler picks a different K, and every in-flight request returns 500.
Served Qwen/Qwen3-0.6B on an RTX 4090 with the same schedule on method "ngram" (the CPU proposer, which accepts a smaller K) and sent 8, 32 and 64 concurrent chat requests. The crash itself was not re-run live.
Badgr now refuses method ngram_gpu together with num_speculative_tokens_per_batch_size before it rents a GPU, and tells the caller to use "ngram" or drop the schedule (commit 59a63d9ef).
Checked 2026-10-10 in Badgr’s local development environment.
← All issue evidence