Issue evidence · Serve · vllm-project/vllm

DeepSeek-V4.1 crashes in decoder-replay CUDA graph with --moe-backend flashinfer_moe

GitHub issue: https://github.com/vllm-project/vllm/issues/61040

Not run

DeepSeek-V4.1 serving crashes in the decoder-replay CUDA graph with the FlashInfer MoE backend.

What Badgr ran

Nothing was run: the report is for a 4x GB200 node, which is not rentable here.

What came back

  • No run attempted.

The suggested comparison (FlashInfer MoE against a supported backend) needs the same GB200 hardware.

Checked 2026-10-11 in Badgr’s local development environment.

← All issue evidence