Issue evidence · AI API · BerriAI/litellm

Getting 500 error: routing NVIDIA NIM models as claude-* works for the first request, then fails

GitHub issue: https://github.com/BerriAI/litellm/issues/45825

Not run

Routing Nemotron on NVIDIA NIM behind claude-sonnet-4-5 and claude-haiku-4-5 aliases returns 500 from the second prompt onward.

What Badgr ran

Nothing was run against NIM: it needs an NVIDIA NIM API key, which Badgr does not have. Read the traceback the reporter attached (LiteLLM 1.104.2, Claude Code through /v1/messages).

What came back

  • The reporter's log shows the failure starting upstream: openai.APIError "Internal server error" raised while reading the NIM stream.
  • LiteLLM then wrapped it as MidStreamFallbackError, and the router logged "No fallback was attempted" because no fallbacks are configured.
  • No LiteLLM defect is visible in the log; the 500 starts at integrate.api.nvidia.com.

A reproduction would need a NIM key, or a mock NIM that returns this stream error. Adding a fallback model to the router config is the obvious mitigation, but that was not tested.

Checked 2026-10-11 in Badgr’s local development environment.

← All issue evidence