Issue evidence · Serve · vllm-project/vllm

Gemma 4 encoder CUDA graph capture allocates max_frames_per_batch full-image rows per budget and OOMs

GitHub issue: https://github.com/vllm-project/vllm/issues/61045

Reproduced, workaround proven

With cudagraph_mm_encoder on and video enabled, Gemma 4 engine startup runs out of memory during encoder CUDA graph capture on a 32 GB card.

What Badgr ran

Served nvidia/Gemma-4-26B-A4B-NVFP4 with the issue's exact command on an RTX 5090 (32 GB) through Badgr, using the vllm/vllm-openai nightly image (vLLM 0.31.1rc1.dev273, torch 2.13.0+cu130). Then tried the issue's two suggested workarounds and a run with encoder CUDA graphs off.

What came back

  • Issue's command: engine died with "Tried to allocate 6.69 GiB", the figure in the issue, after EncoderCudaGraphManager initialized with max_frames_per_batch=1856.
  • Images only (video 0), and encoder_cudagraph_max_frames_per_batch=58: both still ran out of memory in encoder_cudagraph_forward, at 90%, 85% and 80% GPU memory utilization. The failed allocation was 644 MiB with the card full, so the suggested workarounds did not fix it on this build.
  • Without cudagraph_mm_encoder (otherwise identical command): the server became ready and answered a chat completion.

Leaving cudagraph_mm_encoder off is the workaround that worked here. One 5090 host, one run per variant.

Checked 2026-10-11 in Badgr’s local development environment.

← All issue evidence