Issue evidence · Serve · vllm-project/vllm

Qwen3-ASR with Rust frontend: EngineCore dies when a scheduler step batches unequal-length audio

GitHub issue: https://github.com/vllm-project/vllm/issues/60970

Reproduced

With VLLM_USE_RUST_FRONTEND=1, a batch of unequal-length audio clips kills the engine and fails every in-flight request with HTTP 500.

What Badgr ran

Served Qwen/Qwen3-ASR-1.7B with VLLM_USE_RUST_FRONTEND=1 on an RTX 4090 through Badgr (vLLM 0.31.1rc1.dev273). The Rust frontend cannot start on the official checkpoint (see vllm-61057), so the run used the reporter's workaround there: a local copy with a generated tokenizer.json and --hf-overrides vocab_size=151936. Then sent one clip, then 8 concurrent chat requests with distinct audio, 4 near 5 s and 4 near 14 s.

What came back

  • One request on its own: HTTP 200.
  • 8 concurrent mixed-length requests: all 8 returned HTTP 500.
  • Engine log: "ValueError: input_features has rank 3 but expected 2. Expected shape: ('nmb', 'tsl'), but got torch.Size([2, 128, 574])", then the Rust frontend received ENGINE_CORE_DEAD.

Synthetic clips built from two public sample recordings, not the reporter's traffic. One run on one host.

Related Badgr pages

Checked 2026-10-11 in Badgr’s local development environment.

← All issue evidence