Issue evidence · Serve · vllm-project/vllm

ngram_gpu crashes with num_speculative_tokens_per_batch_size (assert num_speculative_tokens == self.k)

GitHub issue: https://github.com/vllm-project/vllm/issues/60981

Reproduced, Badgr fix shipped

A speculative-decoding schedule on the ngram_gpu method kills the engine once the scheduler picks a different K, and every in-flight request returns 500.

What Badgr ran

Served Qwen/Qwen3-0.6B on an RTX 4090 with the same schedule on method "ngram" (the CPU proposer, which accepts a smaller K) and sent 8, 32 and 64 concurrent chat requests. The crash itself was not re-run live.

What came back

  • Badgr verified the endpoint with that schedule.
  • 64 of 64 concurrent requests returned HTTP 200, past the 16-request boundary where the schedule changes K.

What Badgr changed

Badgr now refuses method ngram_gpu together with num_speculative_tokens_per_batch_size before it rents a GPU, and tells the caller to use "ngram" or drop the schedule (commit 59a63d9ef).

Related Badgr pages

Checked 2026-10-10 in Badgr’s local development environment.

← All issue evidence