Issue evidence · Serve · vllm-project/vllm

Rust frontend can't start with the official Qwen3-ASR checkpoints (no tokenizer.json, nested vocab_size)

GitHub issue: https://github.com/vllm-project/vllm/issues/61057

Reproduced, workaround proven

VLLM_USE_RUST_FRONTEND=1 vllm serve Qwen/Qwen3-ASR-1.7B fails at startup: the repo has no tokenizer.json, and once one is added vocab_size is not found.

What Badgr ran

Ran the serve on an RTX 4090 through Badgr (vLLM 0.31.1rc1.dev273, cu134 image), first on the official repo, then on a local copy with a tokenizer.json generated by AutoTokenizer.save_pretrained, with and without --hf-overrides vocab_size=151936.

What came back

  • Official repo: "does not expose a supported tokenizer file (tokenizer.json, tiktoken.model, or *.tiktoken)".
  • With tokenizer.json added: "the model config does not define `vocab_size`".
  • With tokenizer.json and vocab_size=151936: the server started and answered a transcription request with HTTP 200.

The workaround makes the server start, not stable: the same setup then crashes under mixed-length concurrency (vllm#60970). The Python frontend was not compared in this run.

Related Badgr pages

Checked 2026-10-11 in Badgr’s local development environment.

← All issue evidence