Issue evidence · Run · ollama/ollama

[Cloud] glm-5.3-flash / glm-5.3 hang the client on exit: stream never terminates

GitHub issue: https://github.com/ollama/ollama/issues/18926

Not run

Streaming from Ollama Cloud's glm-5.3-flash finishes the work but the client never exits (socket stuck in CLOSE_WAIT), and glm-5.3 is intermittently very slow.

What Badgr ran

Nothing was run: the models are served by Ollama Cloud (ollama.com/v1), which needs an Ollama Cloud API key that Badgr does not have. Read the issue and its log.

What came back

  • The reporter's run shows one client killed after about 57 minutes idle at 0% CPU with zero remote TCP connections.
  • A short supervisor task on glm-5.3 hit a 20-minute hard cap when it normally takes 2 to 3 minutes.
  • minimax-m2.7 and minimax-m3 on the same endpoint exit cleanly, which points at the GLM models or their serving path.

Not reproduced because the cloud service could not be called. A client-side guard (an overall deadline on the stream) is the mitigation a pipeline can apply without waiting for a server fix.

Checked 2026-10-11 in Badgr’s local development environment.

← All issue evidence