Issue evidence · Serve · sgl-project/sglang
GitHub issue: https://github.com/sgl-project/sglang/issues/43588
Not run
With pipeline parallelism above 1, GLM-5.3 intermittently crashes in RotaryEmbedding.forward_cuda.
Nothing was run: the model is large enough to need a multi-GPU node, and the crash is intermittent, so a single run would not settle it.
Needs a funded multi-GPU node and repeated runs.
Checked 2026-10-11 in Badgr’s local development environment.
← All issue evidence