Back to blog
Engineering11 min read

How Many Examples Should a Few-Shot Prompt Have?

How many examples does a few-shot prompt need? A sourced answer covering the 3-5 rule, what the original few-shot research actually tested, and when adding more examples helps or hurts.

NH
Nafiul Hasan
Founder, Prompt Architects

TL;DR: Two to five examples covers most few-shot prompts. Anthropic recommends three to five, OpenAI calls it "a handful", and Google refuses to name a number at all. The original GPT-3 research actually tested 10 to 100. More examples help until they don't, and where that ceiling sits depends on the task, not a fixed count.

How Many Examples Should a Few-Shot Prompt Actually Have?

For nearly every everyday few-shot prompt, the answer is two to five examples. That's the range every major vendor's own documentation converges on, and it's small enough to fit alongside your instructions without crowding out the actual task. The number that matters more than the count itself is whether those two to five examples are diverse enough to cover the shapes your real input will take.

That answer holds for the prompts you type into a chat window. It stops holding once you're running a prompt unattended, at scale, against a task with real category or edge-case complexity, which is where the rest of this guide gets more specific than just use five.

If you haven't already, read few-shot vs. zero-shot prompting first. That guide covers whether to add examples at all. This one assumes you've already decided to, and answers the narrower question: how many. If you're trying to decide which examples to use rather than how many, that's a different question, covered in how to choose good few-shot examples.

Where Does the "3-5 Examples" Rule Come From?

It comes from current vendor guidance, and the three vendors don't even agree with each other on how firm the number is. Anthropic's prompt engineering docs are the most specific: "Include 3–5 examples for best results." That line sits alongside advice to keep those examples relevant, diverse, and consistently formatted (verified from platform.claude.com, September 2026).

OpenAI's guide is deliberately looser. It describes few-shot learning as "including a handful of input/output examples in the prompt, rather than fine-tuning the model." Its own worked example shows exactly three product-review classifications, but the text never commits to a specific number (verified from developers.openai.com, September 2026).

Google goes further and declines to give a number at all. Its Gemini prompting-strategies page has a section literally titled "Optimal number of examples", and the actual advice inside it is: "Models like Gemini can often pick up on patterns using a few examples, though you may need to experiment with the number of examples to provide in the prompt for the best results." The same section warns that "if you include too many examples, the model may start to overfit the response to the examples" (verified from ai.google.dev, September 2026).

VendorStated guidanceSpecific number given?
Anthropic (Claude)"Include 3–5 examples for best results"Yes, 3-5
OpenAI (GPT)"a handful of input/output examples" (shows 3 in its own sample)No fixed number
Google (Gemini)Advises you to "experiment with the number of examples" rather than naming oneNo, explicitly declines

Three vendors, three different levels of specificity, and none of them cite a controlled experiment for the number they land on. That's worth knowing before you treat 3-5 as a research finding instead of a rule of thumb.

Anthropic's docs do offer one genuinely useful move if you're unsure your count is right: "You can also ask Claude to evaluate your examples for relevance and diversity, or to generate additional ones based on your initial set." In other words, if you have three examples and suspect you need five, you can hand the model your set and ask it to point out what shape of input isn't represented yet, rather than guessing at the count in isolation.

What Did the Original Few-Shot Research Actually Test?

Not 3-5. The paper that made in-context learning a mainstream technique, Language Models are Few-Shot Learners (Brown et al., 2020), the GPT-3 paper, describes its own methodology this way: it typically set the number of examples "in the range of 10 to 100 as this is how many examples can fit in the model’s context window" — 2,048 tokens, for GPT-3.

The paper's own benchmark data backs up the idea that more helped, up to a point, at least for that era of model. On the SuperGLUE benchmark, the researchers scaled the number of examples "up to 32 examples per task, after which point additional examples will not reliably fit into our context." Sweeping over that range, they found that GPT-3 "requires less than eight total examples per task to outperform a fine-tuned BERT-Large on overall SuperGLUE score." Eight examples, not three, and that's the number needed to beat a dedicated fine-tuned baseline, not just to see some improvement.

The paper also notes that larger models made increasingly effective use of additional in-context examples than smaller ones did, which is a detail worth sitting with: even in 2020, the right example count already depended on which model you were running, not on a single constant.

Does Adding More Examples Always Help?

No, and the answer changed meaningfully once context windows stopped being the bottleneck. In the classic few-shot regime described above, the pattern was diminishing returns: the first several examples did most of the work, and additional ones added less each time.

A 2024 NeurIPS paper from Google DeepMind, Many-Shot In-Context Learning (Agarwal, Singh, et al., 2024), tested what happens once that ceiling is removed. With today's much larger context windows, the researchers investigated in-context learning with "a large number of shots, for example, hundreds or thousands", what they call the many-shot regime. Their headline finding: "Going from few-shot to many-shot, we observe significant performance gains across a wide variety of generative and discriminative tasks."

That is not a universal green light to pile on examples, and the paper says so directly: "On some tasks (e.g., code verifier, planning), we did observe slight performance deterioration beyond a certain number of shots." So the honest picture is: some tasks keep improving deep into the hundreds of shots, a few tasks quietly get worse past some point, and which bucket your task falls into isn't something you can guess from the technique's name.

There's a practical cost attached to testing that, too. The same paper notes that inference cost rises linearly as you add shots, so chasing many-shot gains on a task that doesn't need them is not a neutral experiment. It costs real tokens whether or not it helps.

How Many Examples Should You Actually Use?

Start at the low end and only go higher when you can point to a specific reason. In practice, that means picking from three rough tiers:

SituationReasonable starting countWhy
One-off chat prompt, well-known task0-2Matches vendor default guidance; not worth the setup cost
Custom format, brand voice, or your own categories3-5The range every vendor's docs converge on for this exact case
Production pipeline, batch job, or agent stepTest 3, 5, 8, and beyondCost of testing is low relative to running the wrong count thousands of times
Huge context budget, complex task, examples are cheap to sourceScale toward dozens or more, and measureMany-shot research shows real gains here, but only on tasks that respond to it

Notice that the table above isn't really about picking one universal shot count. It's about matching the count to how many times the prompt will run and how much a wrong answer costs you. A prompt you'll type once tonight doesn't justify a sweep; a prompt an agent will run thousands of times a day absolutely does, because the difference between three and five examples, multiplied across thousands of runs, is either a rounding error or a real accuracy gap depending on the task.

For anything that runs more than a handful of times, the fastest way to find your real number is a direct sweep rather than a guess. Paste the same instruction, swap only the example count, and run it against a fixed set of test inputs:

[Your instructions, unchanged]

Examples (2):
[example 1]
[example 2]

Examples (4): add
[example 3]
[example 4]

Examples (8): add
[example 5]
[example 6]
[example 7]
[example 8]

[Your real input] →

Run your test inputs against the 2-example, 4-example, and 8-example versions and compare. If accuracy or formatting doesn't move between 4 and 8, you've found your ceiling for that task, and adding more is pure token cost with no return. If it keeps improving, that's your signal the task might reward pushing further, which is exactly the situation the many-shot research describes.

What Are the Signs You Have Too Many (or Too Few) Examples?

Too few examples usually shows up as inconsistency: the model handles your obvious cases fine but drifts on anything ambiguous, because it never saw an example resembling that edge case. Say you're routing support tickets into billing, bug, and feature-request buckets with two examples, one clearly billing and one clearly a bug. The first ticket that reads like both ("charged twice because of a bug in checkout") has nowhere obvious to land, because nothing in your example set showed the model what an overlapping case looks like. If your real inputs vary more than your examples do, that gap is exactly where the model improvises, and not always the way you want.

Too many examples shows up differently, and less predictably. Google's own guidance names the core risk directly: the model can start to overfit to the specific examples rather than the underlying pattern, matching surface details of your samples instead of the rule they were meant to illustrate. In the ticket-routing case, that might look like the model keying off an incidental word that happened to appear in three of your eight billing examples, and misrouting a genuine billing ticket that simply doesn't use that word. Beyond that, every additional example is pure token cost, and on some tasks the many-shot research found actual performance deterioration past a certain count, not just wasted tokens.

The reliable way to catch either problem is the same sweep described above: test a few counts against real inputs, not a hunch about what counts as enough.

Does the Right Count Change for Reasoning Tasks?

The count itself doesn't need to be higher, but what each example needs to show does change. For multi-step reasoning, examples that walk through the steps outperform examples that only show the final answer, because the model imitates the reasoning process, not just the output format.

That's the finding behind chain-of-thought prompting: worked examples that show intermediate steps, not just answers, produced large accuracy gains on math and logic benchmarks using single-digit example counts. So for reasoning-heavy prompts, spend your example budget on showing the path, not on adding more examples of the same shallow answer format. Three well-reasoned examples beat eight answer-only ones on this kind of task.

The same principle applies one level up, to how much detail you give the model about who it should be while reasoning. If you're also tuning the role or persona framing around a few-shot prompt, how specific your persona should be covers that adjacent decision, and the short version is similar: more detail isn't automatically better there either.

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.

Create An Account

The Short Version

There is no single correct number of examples, but there is a defensible default: start at two to five, matching what Anthropic, OpenAI, and Google's own docs each land on for everyday prompts. Push past that only when you have a specific, measurable reason, whether that's a production pipeline where testing 3 versus 5 versus 8 is cheap relative to running the wrong count at scale, or a task where a large context budget and genuinely diverse examples make many-shot prompting worth the extra tokens. Either way, test on your real inputs before deciding, rather than trusting a number from a blog post, including this one.


By Nafiul Hasan — Founder of Prompt Architects, building prompt tooling used across ChatGPT, Claude, and Gemini, writing from hands-on testing of production prompts.

Frequently asked questions

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.

Create An Account