Issue evidence · AI API · BerriAI/litellm
GitHub issue: https://github.com/BerriAI/litellm/issues/45645
Reproduced
A streaming 429 on the primary can put the healthy fallback deployments into cooldown.
Used the issue's config, callback and fake upstream on a local LiteLLM 1.104.2 proxy: two streaming /v1/responses calls to the primary, then a direct call to the backup group.
Checked 2026-10-10 in Badgr’s local development environment.
← All issue evidence