AI & LLMs

Ollama structured output bug skips schema when thinking models answer directly

Ollama versions since 0.34.4 fail to enforce JSON schemas on thinking models that skip the reasoning step, returning unstructured text with HTTP 200.

Illustration of a server rack with a JSON symbol partially hidden by fog
Illustration created for this article

A recent update to Ollama has introduced a subtle but critical bug affecting structured outputs from "thinking" language models. Since version 0.34.4, if a model decides to answer a prompt directly without entering its internal reasoning phase, it ignores the requested JSON schema entirely. This issue affects developers relying on strict data formats for automated workflows, as the server returns invalid structures with a successful HTTP status code.

What happened

The problem stems from how Ollama handles grammar constraints for models that support a "thinking" mode, such as Gemma 4. In an effort to apply structured output rules in a single pass, the software wraps the user-defined JSON schema with a grammar that expects a thinking block first. The system assumes any text appearing before the closing tag of the thinking block is part of the reasoning process and therefore unconstrained. The schema enforcement only kicks in after the thinking block closes.

When a model encounters a simple question and chooses to skip the thinking phase entirely, it never emits the opening or closing tags for that block. Consequently, the grammar wrapper sees the direct answer as "text before the closing" and allows the response to end immediately. Because the schema constraint is technically attached to the post-thinking section, and that section never arrives, the model’s raw output bypasses all formatting rules. The server logs show no errors, and the HTTP response remains 200 OK, making the failure difficult to detect without strict client-side validation.

Testing on Ollama 0.35.1 with the gemma4:e2b model confirmed this behavior. When asked simple factual questions with think: true, the model occasionally skipped reasoning and returned bare strings like "391" or "Paris" instead of the required JSON object. In contrast, every response that included a thinking block adhered strictly to the schema. Disabling the thinking mode entirely (think: false) also resulted in correct schema compliance, indicating the issue is specific to the interaction between the thinking grammar wrapper and direct answers.

Key details

  • The bug affects Ollama versions 0.34.4 and later, including the current 0.35.1 release.
  • It specifically impacts "thinking" models like Gemma 4 when the think parameter is enabled.
  • Responses that skip the thinking phase return unstructured text despite a valid JSON schema request.
  • The server returns HTTP 200 with done_reason: "stop", providing no error indication in logs.
  • An open pull request, #18783, proposes a fix by forcing the model to enter the thinking state.
  • Client-side validation is currently the only reliable way to detect these malformed responses.

Background

Structured output is a feature that forces large language models to reply in a specific format, such as JSON, rather than free-form text. This is essential for integrating AI into software pipelines where downstream code expects predictable data structures. Ollama implements this by converting JSON schemas into grammars that restrict the tokens the model can generate.

"Thinking" models are a newer class of AI that separate their internal reasoning process from their final answer. They typically wrap their chain-of-thought in special tags, allowing the system to distinguish between scratchpad work and the final output. Ollama’s recent change attempted to optimize how these two features work together by treating the thinking block and the final answer as a single continuous grammar sequence. However, this optimization assumed the thinking block would always be present, creating a blind spot for direct answers.

Why it matters

For teams running self-hosted AI infrastructure, reliability is paramount. This bug introduces a silent failure mode where applications receive data that looks valid at the transport layer but fails at the application layer. If your backend expects a JSON object with specific keys and receives a plain string instead, it may crash or behave unpredictably. Because the error does not appear in server logs, debugging can be time-consuming, especially if the issue only triggers on simple prompts that do not require complex reasoning.

Furthermore, this highlights the risks of relying on experimental or recently merged features in production environments. The change that caused this issue was intended to improve performance and consistency, yet it broke a fundamental contract of structured outputs. Teams using thinking models for classification, extraction, or simple Q&A tasks must now assume that schema compliance is conditional on the model’s internal decision to reason. This uncertainty complicates the design of robust AI agents and requires additional defensive coding practices.

What you can do

  • Validate all structured responses against the expected JSON schema in your client application before processing.
  • If reasoning is not required for a specific task, set think: false to ensure the schema is applied directly to the output.
  • Monitor for responses that lack the expected structure and log them separately for analysis.
  • Use the community toolkit script check-ollama-format-think.sh to test if your specific model and version are affected.
  • Consider pinning your Ollama version or avoiding thinking models for critical structured output tasks until a fix is released.
  • Review open pull requests, specifically #18783, for updates on the official patch status.

More news

All news