Issue evidence · AI API · BerriAI/litellm

With BYOK, one client's 429 cools the shared deployment for everyone for the provider's full retry-after

GitHub issue: https://github.com/BerriAI/litellm/issues/45312

Reproduced, workaround proven

One caller's rate-limit error puts the shared deployment into cooldown for every caller, and the cooldown cannot be cleared.

What Badgr ran

Ran the issue's router on litellm 1.104.2 against a mock provider that answers 429 with a 24-hour retry-after, sending 4 requests per configuration.

What came back

  • Default (cooldown_time unset): 2 of 4 requests reached the provider, and a 86400 s cooldown was recorded.
  • cooldown_time=0: 4 of 4 requests reached the provider and no cooldown was recorded.

cooldown_time=0 is a workaround that turns cooldowns off for that router, not a fix for the shared-key case. One mock provider, one run per configuration.

Checked 2026-10-11 in Badgr’s local development environment.

← All issue evidence