Issue evidence · Serve · sgl-project/sglang

Runtime role switch: engine launched with --disaggregation-mode decode crashes with illegal memory access after decode -> prefill -> decode

GitHub issue: https://github.com/sgl-project/sglang/issues/43557

Not reproduced

An engine launched in decode mode crashes on its first decode after being switched to prefill and back to decode.

What Badgr ran

Ran SGLang 0.5.21 with the mooncake transfer backend and --enable-pd-role-switch on 2x RTX 4090 through Badgr, one engine per GPU, serving Qwen/Qwen3.5-0.8B, the issue's model. Followed the issue's five steps, and the workaround (both engines launched as prefill) for comparison.

What came back

  • The engine launched as decode served the request after both switches: all requests returned text and the engine was still alive after step 5.
  • The all-prefill workaround also served every request.

Differences from the issue: requests went to the two engines directly rather than through a router, and decode_cuda_graph_memory_gb was not found in /server_info so a default of 1.0 was passed. The issue's 4-of-4 crash may depend on one of these.

Checked 2026-10-11 in Badgr’s local development environment.

← All issue evidence