TL;DR: Few shot example selection matters more than how many examples you use. Good examples match your real input's format and difficulty, cover the range of cases you'll actually see (including the awkward ones), and sit in an order that doesn't quietly bias the model toward one answer. One mislabeled or off-format example can hurt more than a good one helps.
What Makes a Few-Shot Example "Good"?
A few-shot example is not good just because the label is correct. It's good if it does three jobs at once: it looks like the input your model will actually see, it shows the model something the other examples in the set don't already show, and it sits where its position in the prompt won't quietly distort the answer. If you haven't already settled whether few-shot is even the right call for your task, our few-shot vs zero-shot comparison covers that decision; this post picks up once you've decided to use examples and need to choose the right ones.
Most guidance on in-context learning stops at recommending two to five examples — that's a real question, and our companion piece on how many examples a few-shot prompt needs answers it directly. This post assumes you already have a rough count in mind and answers the harder question: out of everything you could show the model, which specific examples actually earn their place?
Three properties do most of the work:
- Representativeness. Does this example resemble the input distribution you'll run in production — same length, same register, same rough difficulty?
- Diversity. Does this example teach the model something a different example in your set doesn't already cover?
- Position. Does where this example sits in the prompt help or hurt, independent of its content?
Get representativeness and diversity right and you can usually get away with a mediocre position. Get position wrong and even a well-chosen set of examples can mislead the model. We'll take each in turn.
Why Selection Beats Example Count
It's tempting to treat more examples as a free upgrade. It isn't — and the strongest evidence for that comes from a study that didn't touch example count at all, only example choice.
In What Makes Good In-Context Examples for GPT-3? (Liu et al., 2021), researchers held the number of examples constant and only changed which ones were selected. Instead of sampling randomly, they retrieved the examples semantically closest to each test input, using nearest-neighbor search over sentence embeddings — a method they named KATE. The paper reports that "the nearest neighbors, as the in-context examples, give rise to much better results relative to the farthest ones", and separately that this retrieval-based selection "consistently outperforms the random baseline" across the language understanding and generation benchmarks they tested, with especially large gains on table-to-text generation and open-domain question answering.
Read that finding carefully and it says something specific: the farthest, least-relevant examples measurably hurt, not just underperformed. Swapping in examples that don't resemble the real input isn't neutral padding — it's actively worse than showing fewer, better-matched ones. That's the core argument for spending your selection effort on relevance before you spend it on count.
How Do You Pick Examples That Match Your Real Input?
The KATE approach above is the automated version of a simple instinct: pick examples that look like what you're about to ask. In practice, most people building a one-off prompt do this by hand rather than with an embedding search, so here's what matching means concretely.
Match on four dimensions:
- Format. If your real input arrives as a full sentence, don't demonstrate with sentence fragments. If it's a messy pasted email, don't demonstrate with a clean, edited one.
- Length. An example one-tenth the length of your real input teaches the model the wrong expectations about how much reasoning or detail the task requires.
- Register. Casual input needs casual examples. Technical, jargon-heavy input needs technical, jargon-heavy examples. A mismatch here is easy to miss because both versions of the example are individually correct.
- Difficulty. If your real traffic includes hard, ambiguous cases, at least one example should be hard and ambiguous too — not just the two or three easiest cases you happened to have lying around.
If you're building a prompt template that runs against many different inputs (a support-ticket classifier, a content-tagging pipeline), a fixed, hand-picked example set stops matching every input equally well as the input distribution shifts. That's exactly the gap retrieval-based selection is built to close, at the cost of needing a labeled pool to retrieve from and some retrieval infrastructure. For a prompt you write once and run occasionally, hand-picking against the four dimensions above is enough.
| Selection approach | What it optimizes for | Main risk |
|---|---|---|
| Random sampling from what you have | Speed — zero selection effort | Ignores representativeness entirely; easy to end up with an unrepresentative set by accident |
| Hand-picked best examples | Curator judgment, control over labels | Drifts toward the curator's easy favorites; edge cases get left out |
| Similarity retrieval (nearest-neighbor to the real input) | Representativeness, per input | Needs a labeled pool and retrieval infrastructure; can over-fit to surface similarity and skip genuine diversity |
| Diversity-first selection | Coverage of the full input range | Can include cases so rare they're not representative of typical traffic |
How Do You Make Your Examples Cover the Edge Cases?
Representativeness and diversity pull in different directions, and that tension is the point. A set of examples that are all individually representative of the easy, common case will still fail on the hard, rare case if nothing in the set ever showed it.
Selective Annotation Makes Language Models Better Few-Shot Learners (Su et al., 2022) tackles a related but distinct problem: not which examples to show for one specific prompt, but which examples are worth labeling in the first place when you're building the pool you'll select from. Their method, vote-k, is built specifically to balance diversity and representativeness rather than optimize for either alone, and the paper reports it improving downstream task performance over randomly choosing which examples to annotate. The problem isn't identical to picking examples for a single prompt, but the underlying lesson transfers directly: representative-but-similar examples and diverse-but-rare examples are both incomplete on their own, and a good set needs both.
For a task you're building by hand, translate that into a simple mapping exercise before you write a single example:
1. List every distinct "shape" your real input takes:
- the common, clear-cut case
- the ambiguous or boundary case (could reasonably go either way)
- the malformed or incomplete case (missing info, wrong format)
- the adversarial or unusual case (sarcasm, contradictions, edge formatting)
2. For each shape you expect to see in production, either:
- include at least one example of it, or
- explicitly decide you're excluding it and accept the model will guess
3. Check your label distribution: if 4 of 5 examples share one label,
you've built in majority-label bias before the model even runs.
Step 3 matters more than it looks. A set that's accidentally lopsided on labels doesn't just fail to teach the rare label — it actively teaches the model to prefer the common one, which shows up again below.
Does the Order of Your Examples Change the Output?
Yes, and the size of the effect is larger than most people assume. Calibrate Before Use: Improving Few-Shot Performance of Language Models (Zhao et al., 2021) tested a two-example prompt on GPT-3's 2.7B-parameter model that reached 88.5% accuracy on a sentiment task. Reversing the order of those same two examples — nothing else changed — dropped accuracy to 51.3%, close to random guessing on a binary task.
The paper traces this instability to three specific effects: majority label bias (the model leans toward whichever label appears more often among the examples), recency bias (the model leans toward the label shown nearest the end of the prompt), and common token bias (the model leans toward answers that were common in its pretraining data, regardless of your examples). All three push in the same direction: toward whatever is easiest for the model to default to, not necessarily what your examples were trying to teach.
Two practical rules follow directly from the three biases above:
- Balance your labels, or put the rarer label last. If one label dominates your example count, either fix the imbalance or place an example of the less-common label closest to your real input, where recency bias works in its favor instead of against it.
- Don't cluster same-label examples together. A block of three positive examples followed by one negative teaches the model a pattern about position, not just about labels.
If your examples show step-by-step reasoning rather than direct answers, ordering interacts with the reasoning chain too — see our guide to chain-of-thought prompting for how that changes the calculus.
What Does One Bad Example Actually Do to Your Output?
It's worth walking through the mechanism, not just the effect. Say you're building a three-shot prompt to classify support tickets as billing, bug, or feature-request. Here's a clean set:
Ticket: "I was charged twice for my subscription this month."
Category: billing
Ticket: "The export button does nothing when I click it."
Category: bug
Ticket: "Could you add a way to export to CSV, not just PDF?"
Category: feature-request
Ticket: "My invoice shows a $4.99 charge I don't recognize."
Category:
Now swap the second example for one that's subtly off: same label, but wrong delimiter and a vague, non-representative ticket.
Ticket: "The export button does nothing when I click it."
Category = bug <- different delimiter (= instead of :)
Ticket: "Things are broken again I guess" <- vague, no real signal
Category: bug
Nothing about the label changed. But now the model has seen two different delimiters (: and =) and one example so vague it doesn't actually demonstrate what "bug" looks like as a category. Per the mechanisms above, this is exactly the kind of inconsistency that in-context learning is sensitive to: the model has to guess which delimiter you actually want, and it has one fewer genuinely informative example to anchor the bug category against. Multiply that across a batch of a few hundred tickets and you get a visibly inconsistent output format, not just an occasional wrong label. The fix isn't a better prompt instruction — instructions don't override what the examples themselves demonstrate — it's removing or replacing the one example that broke the pattern.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.
Create An AccountA Selection Checklist Before You Ship a Few-Shot Prompt
Run through this before you paste a few-shot prompt into production:
- Does every example match your real input's format, length, and register? If not, fix the example, not the instruction wrapped around it.
- Does your set include at least one boundary or ambiguous case, if your real traffic has them?
- Is your label distribution balanced, or is one label doing most of the work?
- Is your most representative or highest-priority example placed last, closest to the real input?
- Are your delimiters, formatting, and structure identical across every example, with zero exceptions?
- If you're running this prompt against varied inputs at scale, would retrieval-based selection from a labeled pool outperform your fixed set?
- Have you actually looked at outputs from edge-case inputs, not just the easy cases you tested first?
For a broader pass that covers more than just example selection, our prompt hygiene checklist covers the other checks worth running before you send anything important.
Selection is unglamorous work compared to writing a clever instruction, but it's where most of the actual accuracy in a few-shot prompt lives. Get the examples right and the instruction can stay simple.