Issue evidence · Serve · vllm-project/vllm
GitHub issue: https://github.com/vllm-project/vllm/issues/61040
Not run
DeepSeek-V4.1 serving crashes in the decoder-replay CUDA graph with the FlashInfer MoE backend.
Nothing was run: the report is for a 4x GB200 node, which is not rentable here.
The suggested comparison (FlashInfer MoE against a supported backend) needs the same GB200 hardware.
Checked 2026-10-11 in Badgr’s local development environment.
← All issue evidence