Issue evidence · Serve · sgl-project/sglang

DSV4.1 Flash + DSPARK: CUDA caching allocator exhausts GPU memory under sustained load

GitHub issue: https://github.com/sgl-project/sglang/issues/43486

Not run

Under sustained decoding the engine runs out of GPU memory every 15 to 25 minutes.

What Badgr ran

Nothing was run: the report is on 8x B200 and needs a soak of about two hours, roughly $80 of GPU time.

What came back

  • No run attempted.

Over the per-issue budget.

Checked 2026-10-11 in Badgr’s local development environment.

← All issue evidence