Back to blog
Engineering12 min read

Recency and Primacy: Where to Put the Important Bit

Prompt position effects are real: models use information at the top and bottom of a prompt more reliably than the middle. Where to put documents, instructions, and your question, sourced.

NH
Nafiul Hasan
Founder, Prompt Architects

TL;DR: Prompt position effects are documented and real: models use information at the very start or end of a prompt more reliably than information placed in the middle. Put source documents near the top, state your governing instructions once near the start, and end every prompt with your actual question or task.

If a long prompt has ever come back missing one instruction buried in paragraph four, or an assistant answered a different question than the one you actually asked at the bottom of a wall of context, you have already met this effect without a name for it. It is not randomness and it is not the model "not reading carefully." It is a measured, published property of how these models weigh their own input, and it has a fix that costs nothing: move the important bit. If you are newer to structuring prompts deliberately in the first place, our beginner's guide to prompt engineering is the wider context this post assumes.

What are primacy and recency effects, and why do they show up in AI output?

Primacy and recency are borrowed terms from human memory research: people recall the first and last items in a list better than the ones in the middle. The same shape turns up in how language models use a context window, and it is not a coincidence of naming — researchers went looking for exactly this pattern once long-context models made it possible to test.

The mechanism is different from human memory, but the outcome looks similar. A transformer's attention is not a uniform scan of every token with equal weight; training data, positional encoding, and the way instructions get phrased all interact to make some positions easier for the model to draw on than others. The practical result for anyone writing prompts is the same regardless of the cause: content placed at the very beginning or the very end of an input gets used more reliably than content placed in the middle of a long one.

This matters most as prompts grow. A three-sentence request rarely has a meaningful "middle" to lose something in. A ten-page brief, a pasted transcript, a stack of retrieved documents, or a long system-plus-context-plus-question prompt is exactly where position starts to decide which of your instructions actually lands.

What did the lost-in-the-middle study actually find?

The specific, most-cited source behind this claim is a July 2023 paper from Stanford, UC Berkeley, and Samaya AI: "Lost in the Middle: How Language Models Use Long Contexts" (Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang; arXiv:2307.03172, posted July 6, 2023, revised November 20, 2023). It is worth naming precisely because it gets cited constantly and paraphrased loosely.

The authors tested a specific, dated set of models: OpenAI's GPT-3.5-Turbo and GPT-3.5-Turbo (16K), Anthropic's Claude-1.3 and Claude-1.3 (100K), and the open models MPT-30B-Instruct and LongChat-13B (16K). Their task put one document containing an answer among a set of distractor documents, then moved that answer's position from the start to the middle to the end of the input and measured accuracy at each point. The result, in their own words:

"We find that performance can degrade significantly when changing the position of relevant information, indicating that current language models do not robustly make use of information in long input contexts. In particular, we observe that performance is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle of long contexts, even for explicitly long-context models."

Their results section adds a concrete number: "GPT-3.5-Turbo's multi-document QA performance can drop by more than 20%" when the answer sits in the middle of a 20- or 30-document input, to the point that it performed worse than the model got with no documents at all. The paper also found something that undercuts a natural assumption: giving a model a longer context window did not fix the problem. As they put it, "extended-context models are not necessarily better than their non-extended counterparts at using their input context" — GPT-3.5-Turbo and its 16K-token sibling showed nearly identical position sensitivity once both models' inputs fit inside the shorter window.

Does this still happen on today's long-context models?

Partly, and the honest answer is more specific than "yes" or "no." A 2025 paper, "Positional Biases Shift as Inputs Approach Context Window Limits" (Blerta Veseli, Julian Chibane, Mariya Toneva, and Alexander Koller; arXiv:2508.07479, submitted August 10, 2025), revisited the question with a change in method: instead of measuring position as an absolute token count, they measured it relative to how much of a model's own context window the input actually fills.

Their finding reframes the original U-shape rather than simply confirming or denying it. In their own words, the classic effect "is strongest when inputs occupy up to 50% of a model's context window." Past that halfway point, something different happens: "the primacy bias weakens, while recency bias remains relatively stable. This effectively eliminates the LiM effect; instead, we observe a distance-based bias, where model performance is better when relevant information is closer to the end of the input." In plain terms, once a prompt genuinely fills most of a model's available window, the safest position for anything important stops being "start or end" and becomes just "end."

The same paper adds a mechanistic note worth knowing even if you never read past the abstract: "successful retrieval is a prerequisite for reasoning in LLMs, and that the observed positional biases in reasoning are largely inherited from retrieval." A model that cannot reliably find a fact buried mid-context cannot reason over it either, so the position problem is upstream of the reasoning problem rather than a separate one.

None of this means the effect has vanished on current models, and it does not license assuming any specific 2026 model has fixed it. It means the shape of the problem changes with how full the context is, which is a more useful thing to know than either "it's solved" or "it's exactly the same as 2023."

Where do the vendors themselves say to put things?

Independent of the research above, OpenAI, Anthropic, and Google each publish their own placement guidance directly, and it is worth reading as guidance rather than as confirmation of any specific mechanism, since none of the three explains why their rule works.

Each vendor's own current documentation, read directly and quoted verbatim. None of the three publishes the same rule, and none states a mechanism — treat this as practical guidance, not a settled standard.
FeatureLong documents / contextThe query or questionPublished magnitude
OpenAI (general prompt-engineering guide)Not addressed as its own blockcontext is "usually best positioned near the end of your prompt"Not published
OpenAI (GPT-4.1 prompting guide)N/A, guide addresses instructions, not documentsInstructions "at both the beginning and end" of long context, or above it if stated onceNot published
Anthropic (Claude prompting guidance)Place documents "near the top of your prompt, above your query, instructions, and examples"At the end, after documents and instructions"up to 30 percent" response-quality improvement in Anthropic's own tests
Google (Gemini long-context guidance)No single documented rule, differs by image, document, or video inputperformance is "better if you put your query / question at the end of the prompt"Not published; states accuracy "can vary to a wide degree" for multiple retrieval targets

The one point all three agree on, stated plainly and without hedging in every case above, is that the question itself belongs at the end of a long prompt. Where the disagreement shows up is documents and standing instructions: Anthropic is explicit that they go at the top; OpenAI's long-context guide says instructions specifically benefit from appearing at both ends; Google never states a single rule for context placement at all, because its own guidance differs by whether the input is an image, a document, or a video, and contradicts itself across those three cases.

So where do you actually put the important bit?

Combining the research above with what the three vendors currently publish gives a structure that holds up regardless of which model receives the prompt:

  1. Source material and background documents near the top. This is where Anthropic's guidance and the original lost-in-the-middle framing agree most directly: whatever the model has to read but not act on immediately goes first.
  2. Your governing instructions once at the start, and again right before the question if the prompt is genuinely long. This is specifically OpenAI's GPT-4.1 recommendation for long-context prompts, and it costs a few dozen tokens against the risk of an instruction landing in a position the model weighs less.
  3. The actual task or question, stated plainly, as the very last thing the model reads. Every vendor above agrees on this one, and Anthropic is the only one to publish a number for it: up to 30 percent better response quality in its own testing, specifically attributed to query placement.

A minimal skeleton that applies all three, adaptable to any model:

[SOURCE MATERIAL / DOCUMENTS — if long, this goes first]
Document 1: ...
Document 2: ...

[GOVERNING INSTRUCTIONS — state once here; repeat below only if the
 block above is genuinely long, e.g. multiple pages or documents]
- Constraint 1: ...
- Constraint 2: ...
- Output format: ...

[RESTATE THE KEY CONSTRAINT — only for long prompts]
Reminder: <the one rule most likely to get lost>

[THE ACTUAL QUESTION OR TASK — always last]
Given the above, do X.

If you already structure long prompts with XML-style tags to separate sections, this same three-zone order applies inside that structure too — see our guide to using XML tags in Claude prompts for the syntax specifics, which pairs naturally with placing documents and instructions in the order above rather than replacing it.

Does this matter for short, everyday prompts too?

Less than it sounds like it should, and that is worth saying plainly rather than turning a real effect into a universal anxiety. Anthropic scopes its own placement guidance to documents in the tens of thousands of tokens and up; the Liu et al. paper's shortest tested setting was already ten retrieved documents deep. A three-sentence request to summarize an email has no meaningful middle for anything to get lost in.

The habit is still worth adopting everywhere, for a reason unrelated to position bias specifically: ending a prompt with the actual ask, instead of opening with it and then qualifying it into the ground, is just clearer writing. It happens to also be the one placement rule every vendor above agrees on without exception. You are not paying a cost to build the habit into short prompts; you are only avoiding having to remember a different rule for long ones.

Where this becomes a genuine risk rather than a style preference is anything that accumulates length without you noticing: a chat that has scrolled past a dozen exchanges, a pasted support thread, a document you attached and then kept adding follow-up questions under. None of those feel like "a long prompt" while you are typing them, but the model is reading the same accumulated input either way.

A quick checklist before you send a long prompt

  • Is there a single instruction the output absolutely must follow? Put it last, immediately before your question, not buried three paragraphs up.
  • Are you pasting more than a page of reference material? Move it to the top, above your instructions, not interleaved with them.
  • Is your prompt long enough that you would call it "long" out loud? If yes, state your core constraint twice: once near the top, once right before the ask.
  • Does your actual question appear anywhere before the final few lines? Move it down. Every vendor's own documentation agrees on this one point.
  • If the output ignored something specific, check where that instruction sat before assuming the model failed to understand it. Position is a testable variable; move the sentence and try again before concluding the model can't do the task.

None of this requires memorizing which vendor said what. Prompt Architects' generator applies a consistent shape — background, instructions, then the ask — every time it drafts or restructures a prompt for you, which is the same skeleton described above without reconstructing it by hand from three different vendors' documentation each time you write a long one. It will not rewrite a short, single-line request into something more elaborate than it needs to be; the structure only matters once a prompt has enough content for position to start deciding what the model actually uses.

The underlying lesson holds even if every model eventually closes the gap the 2023 paper measured: a prompt is not a bag of instructions the model reads with equal attention regardless of order. Where you put something is part of what you are communicating, the same way it would be in an email, a brief, or a spec — the bottom line still belongs at the end, and the thing you most need followed still deserves a second mention if nothing about the rest of your file structure is doing the reminding for you.

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.

Create An Account

Frequently asked questions

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.

Create An Account