Issue evidence · Serve · ggml-org/llama.cpp
GitHub issue: https://github.com/ggml-org/llama.cpp/issues/30252
Reproduced, workaround proven
With unified KV and many parallel slots, new requests wait seconds while idle slots are saved to the prompt cache.
Served Qwen3-8B Q4_K_M with 8 parallel slots and a unified KV cache on an RTX 4090 through Badgr, started 7 streams, and timed 8 new requests, twice per configuration. Second configuration: LLAMA_ARG_CACHE_IDLE_SLOTS=0.
Two samples per configuration: this shows the stall disappears, not an exact speed-up.
Checked 2026-10-10 in Badgr’s local development environment.
← All issue evidence