TL;DR: Stable Diffusion prompting in 2026 means writing for two different systems that share a name. Stability's hosted API splits prompt weighting across endpoints: documented on Core and Ultra, absent from the SD 3.5 endpoint itself, while local tools like ComfyUI use a separate convention on the same open checkpoints. Negative prompts are real and generous (10,000 characters) everywhere.
Is SD 3.5 Still the Current Stable Diffusion Model?
Yes, as of September 4, 2026. Stability AI's own model page lists Stable Diffusion 3.5 Large, Large Turbo, and Medium, with SD 3.5 Flash sold on the API as a distilled version of Medium. Its news feed's most recent posts cover a Series B funding round, a new way to work with Stable Audio 3.0, and Brand Studio. None of it is a new image model. If you've read that Stability shipped "SD4" or a new base model this year, that claim is not on Stability's own site; it traces back to aggregator blogs, not a primary source, and should be treated as unverified until Stability itself says otherwise.
The trade-off across the family runs in one direction: Large is documented as the most powerful of the four for quality and prompt adherence, Medium balances quality against speed, and Turbo and Flash are distilled specifically to cut generation down to four steps, trading some of both for a much faster result.
The family sits on the same license it shipped with: the Stability AI Community License, which is free for research, non-commercial, and commercial use below $1M in total annual revenue. Cross that line and you need an enterprise agreement directly with Stability. That license question is separate from the hosted API, which bills by credit regardless of your revenue.
Where Do You Actually Write a Stable Diffusion Prompt?
This is the split that most "SD 3.5 prompting" content skips, and it's the reason the same advice doesn't travel between guides. There are two genuinely different places a prompt gets typed:
Stability's own hosted API sits at three endpoints under /v2beta/stable-image/generate/: ultra, core, and sd3. Each takes a flat prompt string, capped at 10,000 characters, plus a handful of request fields that vary by endpoint. Ultra and Core each run one fixed, undisclosed underlying model; only sd3 lets you choose which SD 3.5 variant actually generates the image.
Local, open-weight tooling, meaning ComfyUI and the AUTOMATIC1111-descended UIs that grew up around earlier Stable Diffusion releases, runs the same SD 3.5 checkpoints from Hugging Face on your own hardware or a rented GPU. SD 3.5's own model card describes three fixed text encoders feeding the image model: CLIP-ViT/L, OpenCLIP-ViT/G, and T5-XXL. Local tools expose knobs the hosted API doesn't (steps, sampler, scheduler, and CFG scale on every model, not just one endpoint), and vice versa: none of the hosted endpoints require you to manage a checkpoint file, a VAE, or a text-encoder combination yourself.
Confusing the two is where most stale advice comes from. A parameter or a piece of syntax that's real on one side often doesn't exist on the other.
Does Stable Diffusion 3.5 Support Prompt Weighting Syntax?
This is the question the brief for this piece exists to answer, and the honest answer only holds up if you name the endpoint. Fetching Stability's live OpenAPI specification directly (api.stability.ai/v2alpha/openapi, checked September 4, 2026) shows the prompt field description on generate/core and generate/ultra spelling out an exact syntax: use the format (word:weight), where weight is a value between 0 and 1, and the spec's own worked example reads "The sky was a crisp (blue:0.3) and (green:0.8)" to convey a sky that's more green than blue.
That same explanation does not appear on generate/sd3, the endpoint that actually runs the SD 3.5 Large, Large Turbo, and Medium models. Its prompt field description is shorter and says nothing about weighting syntax at all. So the one hosted endpoint literally named for Stable Diffusion 3.5 is the one where Stability's own docs don't document a weighting mechanism.
Here's what that split looks like endpoint by endpoint:
| Feature | generate/ultra | generate/core | generate/sd3 |
|---|---|---|---|
| Documented (word:weight) syntax | |||
| Choose model variant | |||
| cfg_scale exposed | |||
| negative_prompt field | |||
| steps / sampler exposed |
Local tools sidestep this entirely, but they're not running the same rules either. Community-built UIs on the open checkpoints commonly support their own (word:weight) convention, usually running above 1.0 for emphasis and below for de-emphasis, a wider, differently centered scale than the 0-to-1 range Stability documents on Core and Ultra. If you copy a weighted prompt from a local workflow into Stability's API, or the reverse, expect the numbers to mean something different even when the syntax looks identical.
How Do Negative Prompts Work on the Hosted API?
Better than most competing guides suggest, and worse-documented than you'd expect for something this useful. negative_prompt is a real request field on all three endpoints, capped at 10,000 characters, the same limit as the positive prompt. Its own description text isn't consistent across endpoints, though: Core and Ultra's field description opens "A blurb of text describing what you do not wish to see", while the SD3.5 endpoint's version opens with "Keywords of what you do not wish to see" instead. Both describe the same field; Stability just wrote two different descriptions for it, which is worth knowing before you assume a phrase-style negative prompt behaves identically to a keyword list across endpoints.
Stability doesn't publish a canonical negative-prompt list the way some competitors do for their own models, so treat any "the 50 best Stable Diffusion negative prompts" list you find online as community convention, not vendor guidance. What the field description does confirm is that it's flagged as "an advanced feature" on every endpoint, which reads like Stability's own signal that a bare positive prompt with good, specific wording will usually get you further than stacking negative terms.
For a full cross-vendor breakdown of which models treat negative prompts as a real parameter versus which ones actively discourage the pattern, see our negative prompt support matrix. If you want ready-made negative prompt starting points across several image models including Stability's, our free negative prompt generator builds them from the same field definitions rather than a recycled list.
Which Parameters Actually Move the Output?
Beyond the prompt and negative prompt text, four fields do the rest of the work on generate/sd3, and their behavior is worth knowing precisely rather than by feel:
cfg_scale: how strictly the output follows your prompt, ranging 1 to 10. Stability's own default differs by model: 4 for Large and Medium, 1 for Turbo and Flash. Distilled models are tuned for a low-guidance, few-step regime, so raising cfg_scale on Turbo or Flash to "Large" levels fights the distillation rather than improving it. Our FLUX guidance scale piece covers how a structurally similar parameter behaves on a different vendor's model, if you want the contrast.
aspect_ratio: one of nine fixed values, namely 21:9, 16:9, 3:2, 5:4, 1:1, 4:5, 2:3, 9:16, and 9:21. Defaults to 1:1 if you don't set it, and only applies to text-to-image requests.
seed: an integer from 0 to 4,294,967,294. Pass 0, or omit it, for a random seed. Stability's spec makes no determinism promise here at all, which is worth flagging since several other vendors at least gesture at a similar-results promise for their own seed parameters.
mode and strength: mode switches the request between text-to-image and image-to-image; the latter requires an input image plus a strength value from 0 to 1, where 0 reproduces the input exactly and 1 behaves as if no image were supplied at all. If you're running SD 3.5 Flash specifically for image-to-image, Stability's own field description recommends staying in the 0.94-to-0.97 range for that model, a narrower window than the general 0-to-1 range implies.
style_preset: an optional enum of 17 fixed style names on every endpoint, including photographic, digital-art, cinematic, neon-punk, pixel-art, origami, and 3d-model. It nudges the model toward a rendering look without you having to spell that style out in the prompt text, and it stacks with whatever style language you already wrote rather than replacing it.
Here's what a full request body looks like with those fields in place:
{
"prompt": "a weathered lighthouse on a rocky coastline at golden hour, cinematic lighting, hyperdetailed",
"negative_prompt": "blurry, low contrast, extra towers, text, watermark",
"model": "sd3.5-large",
"aspect_ratio": "16:9",
"cfg_scale": 4,
"seed": 0,
"output_format": "png"
}
How Should You Structure a Stable Diffusion Prompt?
Stability's own guidance for the prompt field is short and worth repeating exactly because it's easy to over-complicate: a strong, descriptive prompt that clearly defines elements, colors, and subjects will lead to better results. That's the whole documented steer. Everything past it is convention rather than requirement, but the convention that's held up is a fixed order:
Subject and action → medium and style → lighting and color → composition and framing → technical modifiers
A filled-in example against that order:
a red fox mid-leap over a fallen log in a snow-covered forest,
oil painting style, warm golden-hour backlight,
low angle, shallow depth of field,
highly detailed fur texture
On Core or Ultra, you can lean on (word:weight) inside that same string to nudge specific terms: (warm golden-hour backlight:0.8) pulls the lighting term down slightly relative to everything else. On the SD3.5 endpoint, there's no documented equivalent, so the same emphasis has to come from word order, repetition, or a stronger adjective instead.
What About ComfyUI and Automatic1111-Style Local Setups?
If you're one of the large number of people still running Stable Diffusion locally rather than through Stability's own API, most of what's above still applies to the model itself, but almost none of it applies to the interface. Local, open-weight setups on SD 3.5 checkpoints:
- Give you direct control over steps and sampler choice, the exact parameters Stability's hosted endpoints have dropped entirely from their current spec. If you want the sampler-by-sampler and steps-by-model breakdown, that's covered end to end in our Stable Diffusion steps guide and our sampler reference.
- Commonly implement their own
(word:weight)parsing on top of the three text encoders SD 3.5 ships with, independent of whatever Stability documents for its hosted API. Treat these as two separate syntaxes that happen to share notation, not one feature with two names. - Run under the same Community License terms as the API-hosted checkpoints, free below the $1M annual revenue threshold and an enterprise agreement above it, since the license attaches to the model weights rather than to how you're calling them.
Because two of those three text encoders cap out at 77 tokens each, a much longer prompt doesn't automatically get read in full on the local side either. That ceiling comes from the encoders themselves, so neither the hosted API nor a local UI can raise it for you.
If you only take one habit from this piece: check which of the three hosted endpoints, or which local tool, a piece of Stable Diffusion advice was actually written against before copying it into your own prompt. A parameter name repeating across two guides doesn't mean the same value range, or even the same field, travels with it.
None of this is a knock on either path. The hosted API trades fine control for zero setup; local tooling trades setup time for exactly the dials the API removed. Knowing which one a piece of advice was written for is most of the battle.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.
Create An Account