TL;DR: Fill seven fields and the block below assembles a thumbnail prompt for any AI image model. Generate the plate, then set the text in a real editor, because models still garble words. YouTube recommends 3840 x 2160 at 16:9, a minimum width of 640 pixels, and custom thumbnails require a verified account.
What is a thumbnail prompt generator?
A thumbnail prompt generator is a fill-in-the-blanks template that turns seven decisions into one image prompt. There is nothing to install and nothing to sign up for. The value sits in the field list and the order, because a thumbnail almost never fails on rendering quality. It fails because the subject is too small, the background competes with the face, or the words were never going to be legible at the size a real viewer sees them.
Fill this in. Only video type, subject and text zone are mandatory.
VIDEO TYPE → tutorial | reaction | vlog | review | versus | listicle |
gaming | podcast | essay | news | short
SUBJECT → the one thing in frame, in one clause
EMOTION → the face or body posture a viewer reads in a quarter second
BACKGROUND → where it sits, and how far out of focus it is
PALETTE → two dominant colours plus one accent that fights them
TEXT ZONE → which third of the frame stays empty for type you add later
CONSTRAINTS → no text, no watermark, no logo, no border, no UI elements
Assembled, that becomes a prompt in this shape:
A [SUBJECT] photographed for a YouTube thumbnail, [EMOTION], set against
[BACKGROUND] rendered shallow and non-competing. Dominant palette
[COLOUR A] and [COLOUR B] with a [ACCENT] accent. Composition leaves the
[LEFT/RIGHT] third of the frame deliberately empty and low-contrast.
Subject occupies at least 40 percent of the frame height. Hard key light
from camera [LEFT/RIGHT], strong rim separation from the background.
16:9. No text, no letters, no numbers, no watermark, no logo, no border.
Two things in there do more work than the rest. The empty third is a text zone you are reserving for type you will set yourself. The "no text, no letters, no numbers" constraint is not decoration, and the next two sections explain why both matter more than any adjective you could add.
What size should a YouTube thumbnail be in 2026?
Bigger than the number most pages still quote. YouTube's Help Center page on custom thumbnails recommends thumbnails "Have a resolution of 3840 x 2160 pixels for videos and 2160 x 3840 for Shorts, with a minimum width of 640 pixels for videos and a minimum height of 640 pixels for Shorts" (accessed August 27, 2026). The old advice to export at 1280 x 720 is not wrong so much as superseded. It is now a floor, not a target.
Here is what that page actually publishes, all of it verified on the same date.
| Requirement | What YouTube publishes |
|---|---|
| Recommended resolution | 3840 x 2160 for videos, 2160 x 3840 for Shorts |
| Minimum | Width of 640 px for videos, height of 640 px for Shorts |
| Aspect ratio | 16:9 for videos, 9:16 for Shorts, 1:1 for podcast playlists |
| Formats | "image formats such as JPG or PNG" |
| File size, mobile upload | 2 MB for video thumbnails, 10 MB for podcasts |
| File size, desktop upload | 50MB for video, Shorts and podcast thumbnails |
| Account requirement | Upload your own "if your account is verified" |
Two footnotes matter. The formats line reads "such as JPG or PNG", which is an example list rather than an exhaustive one, so treat anything outside those two as untested rather than supported. And the YouTube Data API is stricter than the Help Center: the thumbnails.set reference publishes "Maximum file size: 2MB" and accepts only JPEG, PNG and octet-stream MIME types. If you upload through a tool or a scheduler rather than through Studio in a browser, 2MB is the number that will bite you.
Why design for the smallest render, not the export?
Because YouTube itself never shows anyone your 4K file. It shows one of five derivatives, and its Data API documents every one of them by name and pixel size (accessed August 27, 2026): default is "120px wide and 90px tall", medium is "320px wide and 180px tall", high is "480px wide and 360px tall", standard is "640px wide and 480px tall", and maxres, the largest, is "1280px wide and 720px tall".
Sit with the range. The biggest image YouTube serves is 1280 px across. The smallest is 120 px, which is roughly a postage stamp. Your composition has to survive a ten-times reduction without becoming a smudge, and the sidebar and mobile feed live down at the small end of that range, which is exactly where most impressions happen.
That single fact rewrites the prompt. It is why the skeleton above says the subject occupies at least 40 percent of the frame height, and why it asks for a background rendered shallow and non-competing. Detail costs you nothing at 3840 px and everything at 320. YouTube's own creator tips page says it plainly: "Thumbnails show up differently across devices, so make sure your thumbnail image is as large as possible."
The practical test takes ten seconds. Export, then view the file at 20 percent zoom, or drop it into a browser tab and shrink the window until it is the width of your thumbnail on a phone. If you cannot tell what the subject is, no amount of prompt engineering will save it. Go back and ask for fewer objects, higher contrast between subject and background, and a harder key light. If you need vocabulary for that last part, the lighting terms reference covers the forty terms image models actually respond to, and the background replacement guide covers separating a subject from a scene that is fighting it.
Why won't the model just put the text on for you?
Because image models draw letterforms as pixels rather than setting type, and they remain unreliable at rendering specific words legibly. The failure modes are familiar to anyone who has tried: a doubled letter, a plausible-looking word that is not the word you asked for, a second ghost line of nonsense below the real one. It is not a prompt problem and it does not have a prompt solution. We wrote a whole post on why the text in AI images comes out garbled, and the short version is that this is a property of how the models work, not a setting you forgot.
So the workflow that actually ships is two-stage:
- Generate the plate. Subject, emotion, background, lighting, palette. Explicitly forbid text: "no text, no letters, no numbers, no watermark". You want a clean image with a deliberately empty, low-contrast zone in one third of the frame.
- Set the type in a real editor. Figma, Canva, Photoshop, Affinity, GIMP, whatever you already have. The spelling is guaranteed, the kerning is yours, the font is consistent across your channel, and you can re-cut the same plate for three different titles when you A/B test.
That second stage is also the only way to hit YouTube's own advice on legibility. Its creator tips page says "If you add text, make sure to use a font that's easy to read." You cannot guarantee that from a prompt. You can guarantee it from a text layer.
One more reason to keep type in a layer: YouTube documents that "Vertical videos with 16:9 custom thumbnails will be replaced by an auto-generated 4:5 thumbnail on the home, explore, and subscription pages." Crops happen. Type that lives in a separate layer can be repositioned for a crop. Type baked into pixels cannot.
Thumbnail prompt templates by video type
The thumbnail job is not one job. A tutorial thumbnail has to promise a specific outcome, so it needs a legible artefact in frame. A reaction thumbnail has to promise a person's response, so it needs a face and almost nothing else. Presets that offer one "cinematic YouTube thumbnail" style serve neither. Below, 26 prompts grouped by what the video is. Swap the bracketed parts and keep the constraint line.
Tutorial and how-to
The subject is the outcome, not you. Show the finished thing, or the before and after, and leave the face out unless your face is the brand.
Close-up product shot of [FINISHED RESULT] on a clean [SURFACE], lit with
soft directional light from camera left, shallow depth of field, background
falling off to a flat [COLOUR] wash. Right third of frame deliberately empty
and low-contrast for a text overlay. Subject fills the left two thirds.
16:9. No text, no letters, no numbers, no watermark, no logo.
Split composition for a YouTube thumbnail: left half shows [BEFORE STATE]
in flat, desaturated, slightly grey lighting; right half shows [AFTER STATE]
in warm, high-contrast light. Hard vertical seam between the halves. Both
halves shot at the same angle and distance. 16:9. No text, no letters,
no numbers, no arrows, no watermark.
Overhead flat lay of [TOOLS OR INGREDIENTS] arranged on [SURFACE] with
generous negative space in the upper third. Crisp shadows, single overhead
key light, saturated [COLOUR] backdrop. Objects large enough to identify at
small size, maximum [N] objects total. 16:9. No text, no letters,
no numbers, no watermark.
Reaction and commentary
This is the one category where a face is the entire product, and it is also the one where you should not let a model invent one. Generate the environment and composite your own photo into it.
Empty reaction-video background plate: [SETTING, e.g. a blurred neon-lit
desk setup], rendered heavily out of focus, strong colour separation between
foreground plane and background. Left third of frame is a clean, evenly lit
area reserved for a cut-out subject. Right third is empty and low-contrast.
16:9. No people, no faces, no text, no letters, no watermark.
Background plate showing [THE THING BEING REACTED TO] as a large, softly
lit object or screen-shaped rectangle occupying the right two thirds of
frame, with a dark vignette on the left for compositing a person. Colour
palette [COLOUR A] and [COLOUR B], one [ACCENT] highlight.
16:9. No people, no faces, no text, no letters, no watermark.
Vlog
Vlogs sell a place and a feeling. The location does the work, and the person is small on purpose.
Wide environmental shot of [LOCATION] at [TIME OF DAY], one small human
figure at [LEFT/RIGHT] third for scale, dramatic sky, strong leading lines
toward the figure. Saturated but not neon. Upper third kept simple for a
text overlay. 16:9. No text, no letters, no numbers, no watermark.
Handheld point-of-view frame from [ACTIVITY]: [FOREGROUND ELEMENT] entering
from the bottom of frame, [LOCATION] filling the background, natural light,
slight lens flare. Clear focal subject readable at 200 pixels wide.
16:9. No text, no letters, no numbers, no watermark.
Review and unboxing
The product must be identifiable at 320 px. That means one product, one angle, and a background that does not argue with it.
Single [PRODUCT CATEGORY] centred on a seamless [COLOUR] backdrop, three-
quarter angle, studio softbox key from camera left with a subtle rim light
from behind. Product occupies 60 percent of frame height. Clean drop shadow.
Left third empty for a rating badge added later. 16:9. No text, no letters,
no numbers, no brand marks, no watermark.
[PRODUCT CATEGORY] surrounded by its unboxed contents arranged in a loose
arc on [SURFACE], shot from a low three-quarter angle, warm key light,
cool fill. Hero item noticeably larger and sharper than accessories.
16:9. No text, no letters, no numbers, no watermark.
Hands-only shot: two hands holding [PRODUCT CATEGORY] against a blurred
[SETTING], product tack sharp, hands slightly soft, high contrast between
product and background. Upper right quadrant empty and dark.
16:9. No text, no letters, no numbers, no watermark, no faces.
Versus and comparison
Symmetry is the whole idea. Both sides must be shot identically or the comparison reads as a verdict before the viewer has clicked.
Symmetrical two-up thumbnail plate: [OPTION A] on the left, [OPTION B] on
the right, identical camera angle, identical distance, identical lighting.
Narrow dark gap down the centre for a divider added later. Left side lit
[COLOUR A], right side lit [COLOUR B]. 16:9. No text, no letters,
no numbers, no versus symbol, no watermark.
Diagonal split composition: [OPTION A] occupying the lower left triangle,
[OPTION B] the upper right, hard diagonal seam, contrasting colour grade
across the seam, both subjects equal size. 16:9. No text, no letters,
no numbers, no watermark.
Listicle
The number goes in your text layer, not in the image. Ask for a repeating visual motif instead.
Grid of [N] identical [OBJECT CATEGORY] arranged in a loose [ROWS]x[COLUMNS]
pattern on a flat [COLOUR] background, even overhead lighting, consistent
scale and spacing, slight variation in colour between items. Bottom third
left empty. 16:9. No text, no letters, no numbers, no watermark.
One [HERO OBJECT] in sharp focus in the foreground, [N-1] similar objects
receding into soft focus behind it in a diagonal line. Shallow depth of
field, warm key light on the hero, cooler light behind. Right third empty.
16:9. No text, no letters, no numbers, no watermark.
Gaming
Gaming thumbnails compete against the most saturated feed on the platform. High chroma, hard rim light, one readable silhouette.
Dramatic hero shot of [CHARACTER OR VEHICLE DESCRIPTION] in [GAME-STYLE
ENVIRONMENT], strong rim light from behind in [ACCENT COLOUR], volumetric
haze, high contrast, cinematic colour grade in [COLOUR A] and [COLOUR B].
Silhouette must read clearly at 200 pixels wide. Left third empty.
16:9. No text, no letters, no numbers, no HUD, no UI, no watermark.
Wide establishing plate of [GAME LOCATION], dramatic sky, one small
silhouetted figure at the [LEFT/RIGHT] third for scale, strong atmospheric
perspective, saturated [COLOUR] key. Upper third simple for text.
16:9. No text, no letters, no numbers, no HUD, no UI, no watermark.
Extreme close-up of [GAME OBJECT OR ITEM] floating against a dark
[COLOUR] void with rim lighting and particle glow, single strong specular
highlight, heavy contrast. Object centred, edges of frame left dark.
16:9. No text, no letters, no numbers, no watermark.
Podcast clip
Clip thumbnails have to work as a still from a conversation. Generate the set, not the people.
Empty podcast set plate: two [CHAIR/STOOL DESCRIPTION] facing each other
across a [TABLE DESCRIPTION] with two microphones on boom arms, warm
practical lights in the background rendered as soft bokeh, moody low-key
grade in [COLOUR A] and [COLOUR B]. No people in frame. Centre area clear
for compositing. 16:9. No text, no letters, no numbers, no watermark.
Abstract audio-themed background plate: large soft waveform-like shapes and
a suspended studio microphone silhouette against a [COLOUR] gradient, heavy
vignette, shallow depth. Left half dark and empty for a cut-out subject.
16:9. No people, no faces, no text, no letters, no watermark.
Documentary and essay
Restraint reads as authority here. One object, one idea, and a palette that is not competing for attention.
Single symbolic object, [OBJECT], centred on a muted [COLOUR] background,
lit with a single hard light from above creating a long dramatic shadow.
Desaturated film-like grade, visible grain, negative space around the
object on all sides. 16:9. No text, no letters, no numbers, no watermark.
Archival-feeling composition: [SUBJECT MATTER] rendered as a faded,
slightly warm, low-contrast image with visible paper texture and soft
edges, generous empty margin at the bottom. Muted palette, no bright
accents. 16:9. No text, no letters, no numbers, no watermark.
News and update
Timeliness beats beauty. Recognisability beats both. The version number or date belongs in your text layer.
Clean product-announcement plate: [PRODUCT OR SUBJECT] rendered as a simple
geometric shape or device silhouette, centred on a flat [BRAND COLOUR]
background, single soft key light, subtle long shadow to the lower right.
Upper third empty. Minimal, high contrast, readable at 120 pixels wide.
16:9. No text, no letters, no numbers, no logos, no watermark.
Split-tone urgency plate: [SUBJECT] on the left in cool blue-grey light,
an empty high-contrast [ACCENT COLOUR] block occupying the right third for
a headline overlay. Hard-edged, graphic, poster-like rather than
photographic. 16:9. No text, no letters, no numbers, no watermark.
Shorts, the vertical variant
Shorts are a different canvas and a different upload path. YouTube recommends 2160 x 3840 at 9:16, with a minimum height of 640 pixels, and notes that custom thumbnails for Shorts can currently only be added in YouTube Studio on a computer. Every vertical prompt below reserves the centre band, because the top and bottom of a Short get covered by interface.
Vertical composition of [SUBJECT], full-bleed, subject centred and
occupying the middle 60 percent of frame height. Top 20 percent and bottom
20 percent kept simple and low-contrast because interface elements sit
there. Strong single-colour background in [COLOUR]. 9:16 aspect ratio.
No text, no letters, no numbers, no watermark.
Vertical plate: [SUBJECT] shot from a low angle against [BACKGROUND],
strong upward leading lines, dramatic key light from above. Composition
works when cropped to a centre square. 9:16 aspect ratio. No text, no
letters, no numbers, no watermark.
Vertical background plate with no subject: [SETTING] rendered heavily out
of focus with a bright pool of light in the centre band for compositing a
cut-out figure. Palette [COLOUR A] and [COLOUR B]. 9:16 aspect ratio.
No people, no faces, no text, no letters, no watermark.
What are the rules about faces?
Two of them, and they are policy rather than taste.
The first: do not generate someone else's likeness. YouTube's impersonation policy says impersonation "may also include using AI to copy the voice or likeness of an individual, to make it appear as if the channel is owned or authorized by that individual" (accessed August 27, 2026). Its GenAI disclosure page separately requires creators to disclose content that "Makes a real person appear to say or do something they didn’t do." An invented lookalike of a public figure reacting to something is squarely in that territory, and a thumbnail is content.
The second is the one people get backwards: your own face is a photo, not a generation. Shoot three or four reaction expressions once, in decent light, against a plain wall. Cut them out. Reuse them for a year. That is why every reaction and podcast template above generates a background plate with a deliberately empty compositing zone rather than a person. You get a consistent face across your channel, which is worth more than any single striking image, and you sidestep the likeness question entirely.
Will a better thumbnail get you more clicks?
We do not know, and neither does anyone selling you a preset pack. What we can do is point at what YouTube publishes.
Its A/B testing documentation is unusually blunt about the framing. Asked why watch time decides the winner rather than clicks, the page answers: "Great titles and thumbnails serve an important purpose beyond getting viewers to click." It then states the design choice outright: "we optimize tests for overall watch time over other metrics, like click-through-rate." The same page says "It's normal not to receive a" clear winner, and lists two causes: minimal difference between the options, and not enough impressions.
That is a company with the actual data telling you that a thumbnail comparison often has no measurable effect. Click-through rate is shaped by the title, the topic, the audience already following you, whether the impression happened on Home, in search, or in the suggested rail, and what else was on screen next to you. The thumbnail is one input into that, and the video's first fifteen seconds decide whether the click was worth anything. A strong hook and title is not a separate project from the thumbnail; it is the same project.
Read that figure carefully, because it is correlation stated as a "Fun fact" on a tips page. Videos that perform well are made by people who also bother to make thumbnails. It does not license anyone to promise you a percentage lift, and this post will not.
What a good prompt template genuinely buys you is speed and consistency. Twenty-six starting points means you are not staring at an empty box on upload day, and a fixed field order means your channel's thumbnails start to look like they came from the same place. That is worth having on its own terms.
Where is the line on misleading thumbnails?
At the point where the thumbnail promises something the video does not deliver. YouTube's Spam Policy names it directly: "Malicious clickbait: Using maliciously misleading titles, thumbnails, descriptions, or imagery to trick users into clicking on a video that does not deliver what was promised" (accessed August 27, 2026). The same page confirms the policy "applies to all types of content on YouTube, including unlisted and private content, comments, links, posts and thumbnails."
Separately, the custom thumbnails page carries its own content rules: thumbnails may be rejected and earn a strike if they contain nudity or sexually provocative content, hate speech, violence, or harmful or dangerous content, and it warns that "Repeat offenses may lead to the removal of your custom thumbnail privileges for 30 days or even account termination."
None of that makes a dramatic thumbnail a violation. A shocked face over a genuinely surprising result is fine. A shocked face over a video that contains no such result is the thing the policy describes. Read the policy pages rather than any blog's paraphrase of them, this one included, because they are updated more often than the posts that summarise them.
Which plan do you need for image prompts?
Ours, if you want the prompt written for you rather than assembled by hand. Prompt Architects generates the prompt, not the image. You still take the output to whichever image model you already pay for. That is a real limitation and worth saying before the pricing rather than after it.
Image prompt generation starts on the Pro plan. The Free plan does not include it. What Free does include, per our own FAQ page, is 5 prompt enhancements per day, forever, which is enough to paste one of the templates above and have it tightened, but not enough to run a thumbnail workflow through. If you are generating a few images a month, the templates on this page are free and you do not need us. If you are shipping two videos a week and want variants, presets and a saved library, that is the case for Pro. Our GPT Image 2 prompt generator covers the same ground for text-in-image work specifically.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.
Create An Account