Issue evidence · Serve · vllm-project/vllm
GitHub issue: https://github.com/vllm-project/vllm/issues/60970
Reproduced
With VLLM_USE_RUST_FRONTEND=1, a batch of unequal-length audio clips kills the engine and fails every in-flight request with HTTP 500.
Served Qwen/Qwen3-ASR-1.7B with VLLM_USE_RUST_FRONTEND=1 on an RTX 4090 through Badgr (vLLM 0.31.1rc1.dev273). The Rust frontend cannot start on the official checkpoint (see vllm-61057), so the run used the reporter's workaround there: a local copy with a generated tokenizer.json and --hf-overrides vocab_size=151936. Then sent one clip, then 8 concurrent chat requests with distinct audio, 4 near 5 s and 4 near 14 s.
Synthetic clips built from two public sample recordings, not the reporter's traffic. One run on one host.
Checked 2026-10-11 in Badgr’s local development environment.
← All issue evidence