Issue evidence · Serve · ollama/ollama

qwen3.5 returns only thinking text and an empty message.content on a long prompt

GitHub issue: https://github.com/ollama/ollama/issues/18916

Reproduced, workaround proven

On a roughly 19.5k-token prompt, qwen3.5:4b finishes with done_reason "stop", HTTP 200, and nothing in message.content, only reasoning in message.thinking.

What Badgr ran

Served qwen3.5:4b with Ollama on an RTX 4090 through Badgr (64 s to ready, $0.0074 spend) and replayed the issue's request 3 times per variant: the original request, the same request with think=false, and sanitized and reduced copies.

What came back

  • Original 64k-context request: 3 of 3 returned HTTP 200 with 0 visible characters and 87 to 212 characters of thinking.
  • Same request with think=false: 3 of 3 returned 193 to 251 visible characters.
  • Sanitized and exact-replay variants: 6 of 6 empty. The reduced request was empty in 2 of 3, so the failure is frequent, not certain.

Badgr's outcome check treats a 200 with an empty message.content as a failed reply, so a pipeline using it retries or fails loudly instead of passing an empty string on. The workaround is think=false.

Checked 2026-10-11 in Badgr’s local development environment.

← All issue evidence