Issue evidence · Serve · vllm-project/vllm
GitHub issue: https://github.com/vllm-project/vllm/issues/61057
Reproduced, workaround proven
VLLM_USE_RUST_FRONTEND=1 vllm serve Qwen/Qwen3-ASR-1.7B fails at startup: the repo has no tokenizer.json, and once one is added vocab_size is not found.
Ran the serve on an RTX 4090 through Badgr (vLLM 0.31.1rc1.dev273, cu134 image), first on the official repo, then on a local copy with a tokenizer.json generated by AutoTokenizer.save_pretrained, with and without --hf-overrides vocab_size=151936.
The workaround makes the server start, not stable: the same setup then crashes under mixed-length concurrency (vllm#60970). The Python frontend was not compared in this run.
Checked 2026-10-11 in Badgr’s local development environment.
← All issue evidence