Issue evidence · Serve · sgl-project/sglang
GitHub issue: https://github.com/sgl-project/sglang/issues/43557
Not reproduced
An engine launched in decode mode crashes on its first decode after being switched to prefill and back to decode.
Ran SGLang 0.5.21 with the mooncake transfer backend and --enable-pd-role-switch on 2x RTX 4090 through Badgr, one engine per GPU, serving Qwen/Qwen3.5-0.8B, the issue's model. Followed the issue's five steps, and the workaround (both engines launched as prefill) for comparison.
Differences from the issue: requests went to the two engines directly rather than through a router, and decode_cuda_graph_memory_gb was not found in /server_info so a default of 1.0 was passed. The issue's 4-of-4 crash may depend on one of these.
Checked 2026-10-11 in Badgr’s local development environment.
← All issue evidence