Issue evidence · Serve · sgl-project/sglang

PP>1 (GLM-5.3/DSA): intermittent runtime crash in RotaryEmbedding.forward_cuda

GitHub issue: https://github.com/sgl-project/sglang/issues/43588

Not run

With pipeline parallelism above 1, GLM-5.3 intermittently crashes in RotaryEmbedding.forward_cuda.

What Badgr ran

Nothing was run: the model is large enough to need a multi-GPU node, and the crash is intermittent, so a single run would not settle it.

What came back

  • No run attempted.

Needs a funded multi-GPU node and repeated runs.

Checked 2026-10-11 in Badgr’s local development environment.

← All issue evidence