Issue evidence · Serve · vllm-project/vllm
GitHub issue: https://github.com/vllm-project/vllm/issues/61045
Reproduced, workaround proven
With cudagraph_mm_encoder on and video enabled, Gemma 4 engine startup runs out of memory during encoder CUDA graph capture on a 32 GB card.
Served nvidia/Gemma-4-26B-A4B-NVFP4 with the issue's exact command on an RTX 5090 (32 GB) through Badgr, using the vllm/vllm-openai nightly image (vLLM 0.31.1rc1.dev273, torch 2.13.0+cu130). Then tried the issue's two suggested workarounds and a run with encoder CUDA graphs off.
Leaving cudagraph_mm_encoder off is the workaround that worked here. One 5090 host, one run per variant.
Checked 2026-10-11 in Badgr’s local development environment.
← All issue evidence