Issue evidence · Serve · ggml-org/llama.cpp

server: saving idle slots to the prompt cache stalls all requests for seconds with interleaved KV cells

GitHub issue: https://github.com/ggml-org/llama.cpp/issues/30252

Reproduced, workaround proven

With unified KV and many parallel slots, new requests wait seconds while idle slots are saved to the prompt cache.

What Badgr ran

Served Qwen3-8B Q4_K_M with 8 parallel slots and a unified KV cache on an RTX 4090 through Badgr, started 7 streams, and timed 8 new requests, twice per configuration. Second configuration: LLAMA_ARG_CACHE_IDLE_SLOTS=0.

What came back

  • Default: worst wait 4.0 s and 2.55 s (median 0.73 s).
  • Idle-slot caching off: worst wait 0.94 s and 0.87 s.

Two samples per configuration: this shows the stall disappears, not an exact speed-up.

Related Badgr pages

Checked 2026-10-10 in Badgr’s local development environment.

← All issue evidence