Back to blog
Engineering12 min read

The Structured Outputs API: response_format, JSON Schema and Strict Mode (2026)

How response_format, text.format and output_config.format actually work on the structured outputs API: schema rules, strict mode, refusals and limits. OpenAI, Anthropic, Google docs, Sept 2026.

NH
Nafiul Hasan
Founder, Prompt Architects

TL;DR: A structured outputs API is the provider parameter that forces valid, schema-matching JSON: response_format or text.format on OpenAI, output_config.format on Claude, and response_format again (a different shape) on Gemini's newer Interactions API. Each enforces the shape through constrained decoding rather than a prompt asking politely, but guarantees, refusal handling and schema limits differ sharply by vendor.

What is a structured outputs API, and how is it different from a JSON prompt?

A structured outputs API is a request parameter, not a prompting technique. You still write natural language telling the model what to extract or generate, but you also attach a JSON Schema to a dedicated field, and the provider constrains its own token sampling so the output can only be valid against that schema. That is a stronger guarantee than the technique covered in our guide to JSON prompts explained: asking a model to respond in JSON matching a schema pasted into the message body works in any chat window, but nothing enforces it there, so the model can still wrap the answer in a code fence, invent an extra field, or drop the instruction under load.

Structured outputs are the API-level version of the same idea, and this post is about that layer specifically: what each provider's parameter actually guarantees, what it refuses to do, and where a schema that validates cleanly on one vendor gets rejected outright on another. If you want ready-to-paste schemas and full request bodies for common jobs, that is the job of our JSON prompt generator; this page is the mechanism underneath it.

How does OpenAI's response_format enforce a schema?

OpenAI ships Structured Outputs on two surfaces with two different field names. Chat Completions takes response_format: { type: "json_schema", json_schema: { name, strict, schema } }. The newer Responses API takes the schema one level down, on text.format, with the same type: "json_schema" and strict fields but no wrapping json_schema key:

curl https://api.openai.com/v1/responses \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-2024-08-06",
    "input": [
      { "role": "system", "content": "Extract the event." },
      { "role": "user", "content": "Alice and Bob meet Friday at 3pm." }
    ],
    "text": {
      "format": {
        "type": "json_schema",
        "name": "event",
        "strict": true,
        "schema": {
          "type": "object",
          "properties": {
            "name": { "type": "string" },
            "date": { "type": "string" },
            "participants": { "type": "array", "items": { "type": "string" } }
          },
          "required": ["name", "date", "participants"],
          "additionalProperties": false
        }
      }
    }
  }'

That strict: true flag is the whole feature. Without it, json_schema mode still tries to follow your schema but does not guarantee it; with it, OpenAI's docs list "Explicit refusals: Safety-based model refusals are now programmatically detectable" as one of the direct benefits, alongside not having to retry malformed responses (fetched from platform.openai.com/docs/guides/structured-outputs, Sept 4, 2026). Model support matters here too: response_format: {type: "json_schema", ...} "is only supported with the gpt-4o-mini, gpt-4o-mini-2024-07-18, and gpt-4o-2024-08-06 model snapshots and later" — older snapshots fall back to plain JSON mode, which only guarantees parseable syntax, not your keys.

The schema itself is constrained, and the numbers are specific rather than vague. Every property must appear in required (there is no native optional field; emulate one with a nullable union type), additionalProperties: false is mandatory, and the root of the schema must be a plain object, not an anyOf union. Size limits are concrete: "A schema may have up to 5000 object properties total, with up to 10 levels of nesting", and separately, "A schema may have up to 1000 enum values across all enum properties." Composition keywords allOf, not, dependentRequired, dependentSchemas, if, then and else are not supported at all, on any model. A schema using any of them is rejected outright with strict: true, not silently ignored.

How is Anthropic's output_config.format different from OpenAI's approach?

Claude splits the same problem into two named features you can use separately or together: "JSON outputs (output_config.format): Get Claude's response in a specific JSON format" and "Strict tool use (strict: true): Guarantee schema validation on tool names and inputs" (platform.claude.com/docs/en/build-with-claude/structured-outputs, Sept 4, 2026). The first shapes what Claude says; the second validates the arguments it hands to your functions. A request can use either alone or both together.

response = client.messages.create(
    model="claude-opus-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Extract: John Smith, john@example.com, wants a demo."}],
    output_config={
        "format": {
            "type": "json_schema",
            "schema": {
                "type": "object",
                "properties": {
                    "name": {"type": "string"},
                    "email": {"type": "string"},
                    "demo_requested": {"type": "boolean"},
                },
                "required": ["name", "email", "demo_requested"],
                "additionalProperties": False,
            },
        }
    },
)

This shape is new enough that a lot of published guidance, including our own earlier posts, still describes Claude reaching JSON only through forced tool use. That is out of date: output_config.format is now generally available, no beta header required, though "beta headers are no longer required" is a change worth flagging precisely because the API still accepts the old shape too, for a transition period, and the Python SDK's .parse() helper still takes output_format=<PydanticModel> as a convenience that translates internally.

Anthropic's schema subset is more permissive than OpenAI's in one direction and stricter in another. Optional properties are allowed, but the output reorders them: required properties render first in schema order, then optional ones, which is a real behavioral difference from OpenAI (where reordering the same way is moot, since every field is required anyway). Numeric and string-length constraints like minimum or maxLength are rejected outright as unsupported, though the Python, TypeScript, Ruby and PHP SDKs auto-transform a schema that has them: they strip the constraint, move it into the field's description instead, and validate the real constraint client-side against your original schema. Recursive schemas, external $ref, and complex types inside an enum are not supported at all.

Anthropic is also the only one of the three that publishes hard numbers on schema complexity: a single request is capped at 20 tools with strict: true, 24 total optional parameters across every strict schema, and 16 parameters using anyOf or nullable unions, all combined. Exceed those and you get a 400 error; even under those ceilings, deeply nested or highly optional schemas can still hit an internal grammar-size limit and the same 400, with the message "Schema is too complex for compilation." Anthropic backstops this with a 180-second compilation timeout.

Where does Google's Gemini response_format fit, and why did the shape change?

Gemini's older generateContent endpoint took responseMimeType plus responseSchema (an OpenAPI-style schema subset) inside generationConfig. Fetching that reference page directly on Sept 4, 2026 shows both fields, plus the newer _responseJsonSchema alternative, marked with the same warning: "This item is deprecated!" (ai.google.dev/api/generate-content).

The documented path now is the Interactions API, a different endpoint (POST https://generativelanguage.googleapis.com/v1beta/interactions) with a response_format object. Unlike OpenAI or Anthropic, Gemini's response_format.type is not fixed to "json_schema"; it names the output modality, and only a "text" response format carries a schema at all:

curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
  -H "x-goog-api-key: $GEMINI_API_KEY" -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-3.8-flash",
    "input": "List 3 popular cookie recipes as JSON.",
    "response_format": {
      "type": "text",
      "mime_type": "application/json",
      "schema": { "type": "object", "properties": { "recipes": { "type": "array", "items": {"type": "string"} } } }
    }
  }'

An "image" response format instead takes aspect_ratio and image_size; an "audio" one takes sample_rate. That is a genuinely different design from OpenAI and Anthropic, where the schema-enforcement parameter only ever governs text.

On schema features, Gemini is the most permissive of the three in places OpenAI and Anthropic refuse outright: minimum, maximum, minItems and maxItems are documented as directly supported, and fields do not all have to be required. Gemini is the weakest of the three on precision about limits. Its own structured-output guide lists exactly two limitations. First: "Schema subset: Not all JSON Schema features are supported." Second: "Schema complexity: Very large or deeply nested schemas may be rejected" (ai.google.dev/gemini-api/docs/structured-output, Sept 4, 2026), with no property count, nesting depth, or enum ceiling published anywhere on the page. If you need to know exactly how large a schema Gemini will accept, there is currently no documented number to test against; you find the ceiling by hitting it.

How do the three compare, side by side?

OpenAIAnthropic (Claude)Google (Gemini)
Parameterresponse_format (Chat Completions) / text.format (Responses)output_config.formatresponse_format (Interactions API)
type value"json_schema""json_schema""text" (schema only applies here)
Every field required?Yes, no exceptionsNo — optional fields allowed, reordered after requiredNo — optional fields allowed
additionalProperties: falseRequiredRequiredNot mentioned as required
Published size limits5,000 properties / 10 nesting levels / 1,000 enum values20 strict tools, 24 optional params, 16 union params per request, 180s timeoutNone published ("may be rejected")
Refusal signalDedicated refusal field on the parsed responsestop_reason: "refusal", HTTP 200, tokens billedNot documented on the structured-output page
Works with streamingYes (general API feature)Explicitly listed as supportedNot addressed on this page

What happens when the model refuses, or the response is cut off?

This is where the three vendors diverge the most, and it is the part most comparison posts skip. On OpenAI, a refusal shows up as its own field on the parsed message, separate from your schema's fields, specifically so your code does not have to guess whether a null value was intentional or a refusal in disguise.

On Anthropic, "Claude maintains its safety and helpfulness properties even when using structured outputs." If it refuses, the response carries stop_reason: "refusal", an ordinary 200 status code, and you are billed for the tokens generated. Anthropic states the consequence directly: "The output may not match your schema because the refusal message takes precedence over schema constraints" — the grammar constraint does not silence the refusal. The same applies if the model hits max_tokens mid-generation: the partial output can be syntactically incomplete, and Anthropic's own guidance is to retry with a higher token budget rather than try to parse a truncated object.

Gemini's structured-output page does not document an equivalent refusal shape for schema-constrained responses at all. That absence is worth noting rather than guessing around: it means you should build your own detection (an empty or oddly-shaped result, an unexpected finish reason) rather than assume Google's refusal path mirrors either of the other two.

Does a valid schema mean the answer is correct?

No, on all three providers, and this is the mistake that shows up most in production incidents. A schema enforces the container: field names, types, and (on OpenAI and Anthropic) which fields must be present. It says nothing about whether price_usd: -50 makes sense, whether an extracted email is real, or whether a summary is faithful to the source text. Anthropic's own SDK behavior makes the boundary explicit: when a constraint like minimum: 100 is not supported natively, several of its SDKs quietly convert it into a plain integer field plus a description that reads "Must be at least 100" and validate the real constraint against your original schema client-side. The provider is telling you, directly in its own tooling, that shape and correctness are handled by two different systems. Run every parsed object through Zod, Pydantic, or a JSON Schema validator regardless of which vendor you use; if you want fully worked schemas for the common cases, that is what our JSON prompt generator collects in one place.

Can you combine structured outputs with streaming, caching, or tool calls?

On Claude, yes, explicitly. Anthropic's compatibility list names batch processing, token counting, streaming, and combining JSON outputs with strict tool use in a single request as all supported together. There is a cost wrinkle worth knowing before you budget for it: using output_config.format adds an automatic system prompt explaining the expected shape, so your input token count rises slightly on every call, and changing the format parameter invalidates any prompt cache tied to that conversation. If your bill has crept up since you switched a pipeline to structured output, that overhead is a real, if small, contributor; our guide to why your API bill is high covers the bigger levers.

Streaming structured output does not change how you should consume it in practice. A partial JSON fragment mid-stream is not valid against your schema by definition, so on all three providers the practical pattern is the same: stream for latency if you need it, but only parse and validate once the response is complete. Our streaming responses guide goes into when a half-finished response is actually usable and when it is not, which matters more for prose than it does for a JSON object you plan to feed straight into a database row.

Structured outputs solve a narrower problem than they are often sold as solving: they guarantee shape, on the vendor's terms, not correctness, and those terms are not interchangeable across OpenAI, Anthropic and Google. Pick the parameter name that matches your provider, respect its schema subset instead of assuming the other two's rules apply, and keep validating values downstream regardless of how confident the shape guarantee makes you feel.

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.

Create An Account

Frequently asked questions

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.

Create An Account