TL;DR: A literature review prompt generator builds the prompt, not the review. The safe pattern is narrow: paste the papers you already have, then ask the model to synthesise, compare and critique only what you pasted. Never ask it to find papers or supply citations from memory. Every citation-touching prompt below carries a verification instruction.
What does a literature review prompt generator actually generate?
It generates a prompt. Not a review, not a bibliography, not a set of sources.
That distinction sounds pedantic until you watch someone paste "write me a literature review on adolescent sleep and academic performance with 20 references" into a chat window and receive twenty beautifully formatted references, most of which do not exist. The output looked like the thing they asked for. It was not the thing.
A good literature review prompt generator does three things. It forces you to state what you already have, because the model's usable evidence is whatever you put in front of it. It picks the job you actually want done, because "help with my lit review" is six different tasks. And it attaches a verification instruction to anything that touches a citation, so that a gap in the evidence shows up as the words NOT IN SOURCE rather than as a confident sentence.
That third part is the whole product. Everything else is convenience.
Why is a literature review the worst place to trust a model's memory?
Because citations are the most fakeable thing a language model produces, and the most expensive thing to get wrong.
A reference is highly patterned text: surname, initial, year, plausible title, plausible journal, volume, page range, DOI. A model that has read millions of them can generate a new one that is structurally perfect and entirely invented. It reads as authoritative precisely because it follows the format.
That figure comes from a study that did the obvious experiment. The authors "used ChatGPT-3.5 and ChatGPT-4 to produce short literature reviews on 42 multidisciplinary topics", then checked all 636 references. Their finding: "55% of the GPT-3.5 citations but just 18% of the GPT-4 citations are fabricated". Among the references that were real, "43% of the real (non-fabricated) GPT-3.5 citations but just 24% of the real GPT-4 citations include substantive citation errors."
A parallel study in medical topics found worse. Of 115 ChatGPT-generated references, "47% were fabricated, 46% were authentic but inaccurate, and only 7% were authentic and accurate", and the most common single defect was a wrong identifier: "an incorrect PMID number was most common, listed in 93% of papers."
The obvious objection is that these are 2023 studies of 2023 models. Fair. So here is a 2026 one. A referencing audit of five current chatbots published in BMJ Open on 14 April 2026 reported that "Reference quality was poor, with a median completeness score of 40%", and concluded flatly that "Chatbot hallucinations and fabricated citations precluded any chatbot from producing a fully accurate reference list." Not one of the five produced a clean bibliography.
This is not a bug that got fixed. It got smaller and stayed.
It is not only an academic problem. In Mata v. Avianca the US District Court for the Southern District of New York sanctioned two lawyers and their firm who, in the judge's words, "submitted non-existent judicial opinions with fake quotes and citations created by the artificial intelligence tool ChatGPT, then continued to stand by the fake opinions after judicial orders called their existence into question." The same opinion is careful to say that "there is nothing inherently improper about using a reliable artificial intelligence tool for assistance". The failure was not the tool. It was skipping the check.
Your discipline has the same rule written down. The ICMJE recommendations say authors "should carefully review and edit the AI-generated content as the output can be incorrect, incomplete, or biased", and that "Referencing AI-generated material as the primary source is not acceptable." COPE's position is blunter: "Authors are fully responsible for the content of their manuscript, even those parts produced by an AI tool, and are thus liable for any breach of publication ethics."
What can a model legitimately do with literature?
Quite a lot, as long as the evidence is in front of it rather than in its weights.
| Feature | Sources you paste in | The model's memory |
|---|---|---|
| Summarise a study's method and sample | Reliable, with quote check | Unreliable |
| Extract themes across a set of papers | Reliable, with quote check | Unreliable |
| Build a comparison matrix | Reliable, with NOT IN SOURCE cells | Unreliable |
| Identify where two studies disagree | Reliable, with quote check | Unreliable |
| Draft a structure for the review | Reliable | Weak but harmless |
| Critique your own draft synthesis | Reliable, with quote check | Weak but harmless |
| Suggest database search terms | Useful, run them yourself | Useful, run them yourself |
| Find papers on a topic | Out of scope | Do not do this |
| Supply a citation for a claim | Out of scope | Do not do this |
| Tell you what the literature says | Out of scope | Do not do this |
The right shape of tool for this work is source-grounded: you supply documents, and answers are constrained to them with inline citations back into your own files. Google describes its notebook product as letting you "get grounded information based on your sources with clear in-line citations for accuracy, transparency, and trust", which is exactly the property you want. We covered the prompting style that suits it in how to prompt a source-grounded notebook. You can get most of the same discipline in a plain chat window, but you have to impose it in the prompt yourself, which is what the rest of this page does.
How do you build a literature review prompt from scratch?
Fill five slots. The generator prompt below is the skeleton every other prompt on this page is an instance of.
ROLE
You are a research assistant working strictly within supplied sources.
SOURCES
Below are [N] sources. Each is delimited and labelled S1, S2, S3...
You may use nothing outside these sources. You have no other knowledge
of this literature for the purposes of this task.
TASK
[One job: extract / synthesise / compare / map disagreement / outline /
critique my draft]
OUTPUT FORMAT
[Table with named columns / numbered themes / outline with H2 and H3 /
paragraph with inline source labels]
EVIDENCE RULES
1. Every substantive claim must be followed by the source label and a
verbatim quoted sentence from that source, plus its section or page.
2. If a source does not address a cell, row or question, write
NOT IN SOURCE. Do not infer, do not fill the gap, do not average.
3. If two sources conflict, report both with quotes. Do not resolve it.
4. Do not add citations, authors, years or DOIs that are not in the
supplied text.
SOURCES BEGIN
--- S1 ---
[paste full text or the relevant extract]
--- S2 ---
[paste]
Slot four and slot five are the ones people delete when they are in a hurry. They are the ones doing the work.
Prompts for synthesising sources you supply
Paste the sources first. Every prompt here assumes they are already in the conversation as S1, S2, S3 and so on.
2. Structured record from one paper
From S[n] only, fill this record. For each field, give the value, then a
verbatim quoted sentence from S[n] supporting it, then the section or page.
Write NOT IN SOURCE for any field the paper does not state.
Research question | Design | Setting | Population and N | Sampling |
Intervention or exposure | Comparator | Primary outcome and how measured |
Key result with direction and magnitude | Stated limitations | Funding
3. Same schema across the whole set
Apply the record schema above to S1 through S[n]. Return one row per source
in a markdown table, columns exactly as listed. Every populated cell carries
a quoted fragment and a page or section. Empty cells must read NOT IN SOURCE.
Do not harmonise terminology across papers; use each paper's own wording and
note in a final column where the wording differs.
4. Thematic synthesis
Working only from S1 through S[n], identify the themes that appear in two or
more sources. For each theme: a name, a two-sentence description, the list of
sources supporting it, and one verbatim quoted sentence per supporting source
with its page or section. List separately any theme that appears in exactly
one source. Do not name a theme that no source states.
5. Method and population map
Across S1 through S[n], map how the evidence was produced. Group the sources
by study design, then within each group note sample size, population and
setting. Quote the sentence that establishes each design classification.
Where a paper's design is ambiguous or unstated, write DESIGN NOT STATED and
quote the closest sentence. Finish with the two populations most and least
represented in this set.
6. Findings synthesis with direction
For the question "[your review question]", extract from S1 through S[n] every
finding that bears on it. For each: the source label, the direction of the
effect, the magnitude as the paper reports it including units, and a verbatim
quoted sentence. Do not convert units, pool results or compute an average.
If a source does not bear on the question, say so explicitly.
7. Gap statement grounded in the set
Based only on S1 through S[n], write three candidate gap statements. Each must
name what is missing and cite the sources that establish the boundary of what
has been done, with a quoted sentence each. Then, for each gap, state what a
reader could reasonably object to, given that these [n] sources are not the
whole literature. Do not claim a gap exists in the field, only in this set.
How do you build a comparison matrix across papers?
Build the empty matrix first, then fill it, then audit it. Three prompts, in that order, because a model asked to design and populate a table in one pass will design the table around what it can fill.
8. Design the matrix
From S1 through S[n], propose the columns for a comparison matrix that would
let a reader judge these studies side by side. Return only the column names
and a one-line definition of each. Do not fill anything in yet. Include at
least one column for a dimension on which these sources clearly differ, and
name that dimension explicitly.
9. Fill it
Fill the matrix agreed above for S1 through S[n]. One row per source. Every
populated cell must contain the value plus a verbatim quoted fragment from
that source and its page or section. Any cell the source does not address
must read NOT IN SOURCE. Do not carry a value across from another paper. Do
not use "similar to S2" as a value.
10. Audit it
Review the matrix you just produced. For every cell, check that the quoted
fragment actually supports the value in that cell, and that the fragment
appears verbatim in the source I supplied. List any cell where the quote does
not appear verbatim, or where the value goes beyond what the quote states.
Return that list first, before anything else, even if it is empty.
Prompt 10 catches the specific failure where a quote is real, the value is plausible, and the value is not what the quote says. That is the error that survives a careful read.
How do you find where the studies disagree?
Ask for the disagreement directly, because a model asked to synthesise will smooth it away by default.
11. Contradiction sweep
Across S1 through S[n], list every point on which two or more sources
disagree. For each: the point, the sources on each side, and a verbatim
quoted sentence from each source with its page or section. Do not resolve the
disagreement, do not say which is more likely correct, do not describe them as
"broadly consistent". If there are no direct contradictions, say so and list
the three points where the sources are closest to conflicting instead.
12. Explain the disagreement without resolving it
For contradiction [x] above, list the differences between the disagreeing
studies that could plausibly account for it: population, setting, measurement,
timeframe, comparator, analysis choice. Cite each difference with a quoted
sentence from the relevant source. Mark clearly which of these explanations
the sources themselves propose and which are your inference. Do not conclude.
13. Rank by how much it matters
Rank the contradictions you listed by how much each would change the
conclusion of a review answering "[your review question]". For each, state
in one sentence what the review would conclude under each side, and repeat
the verbatim quoted sentence from each source with its page or section so
the ranking can be audited. Flag any contradiction that sits in a claim I
would otherwise have written as settled background. If a contradiction
cannot be re-quoted, drop it from the ranking and say which one you dropped.
How do you draft and critique the review itself?
The drafting prompts are safe because they operate on structure. The critique prompts are the ones that earn their place, because they turn the model on your own text instead of on the literature.
14. Structure from the evidence you have
Propose an outline for a literature review answering "[question]", using only
S1 through S[n] as evidence. Give H2 sections and H3 subsections. Under each
subsection, list which sources would support it, what each contributes, and
one verbatim quoted sentence per source with its page or section. Mark any
subsection you cannot support from this set as EVIDENCE GAP rather than
including it for completeness, and write NOT IN SOURCE against any source
you listed but could not quote.
15. One synthesis paragraph, properly anchored
Write one paragraph of synthesis for subsection [x], using only S1 through
S[n]. Structure it as a claim about the state of the evidence, not a list of
study summaries. Mark every claim inline with its source labels. After the
paragraph, list each claim with the verbatim quoted sentence supporting it and
its page or section. If a claim cannot be supported by a quote, delete the
claim and say which one you deleted.
16. Critique a synthesis I wrote
Here is my draft synthesis and the sources it cites. For each sentence, decide
whether it is (a) supported by a quoted sentence in the supplied sources,
(b) an overreach beyond what the sources state, (c) unsupported by anything
here, or (d) a claim about the field that these sources cannot establish.
Return a numbered list, one line per sentence, with the quoted evidence for
every (a). Be harsh. Do not defend my wording.
17. Unsupported-sentence hunt
Read only my draft. List every sentence that carries no citation but makes a
factual claim about prior research, and every sentence whose citation would
have to be doing more work than a single study can do. Do not rewrite
anything. Just list the sentence numbers and, for each, what evidence it would
need.
How do you turn a research question into database search terms?
This is the one job where the model's broad vocabulary is an asset and its ignorance of your library is irrelevant, because you are going to run the search yourself.
18. Question into concept blocks
Take my question: "[question]". Break it into concept blocks. For each block,
list synonyms, alternative spellings, older terminology, and likely controlled
vocabulary headings, clearly labelled as candidates I must confirm in the
database's own thesaurus. Do not claim any term is an official subject
heading. Flag terms likely to produce high noise and say why.
19. Database-ready string
Assemble the confirmed terms into a Boolean search string for [PubMed /
Scopus / Web of Science / your database]. Use that database's syntax for
truncation, phrase search and field tags. Give me the string, then a
plain-English description of what it will and will not retrieve, then three
ways it is likely too narrow. I will run this myself and refine it.
Structured question frameworks help here, and their provenance is worth knowing. PICO's components are credited to Richardson and colleagues in 1995 by the Cochrane Handbook, though the original is an editorial without an abstract, so the acronym is widely attributed to that paper rather than quotable from it. SPIDER is cleanly traceable to Cooke, Smith and Booth in 2012. FINER traces to Cummings, Browner and Hulley in Designing Clinical Research, and guides cite different editions, which is why two dates circulate. PICo is widely attributed to JBI, but we could not confirm the expansion from an open primary source. We worked through all of this in the research question generator, including which frameworks apply outside health sciences, where PICO is not the default.
How do you actually check one citation?
Take a citation. Any citation, including one you are confident about. Four steps, five minutes.
- Resolve the DOI. Paste it after
https://doi.org/. If it does not resolve, the record does not exist as printed. A DOI with the right shape is not evidence of anything. - Search the exact title in a real database. Not a chat window. For biomedical work note that
pubmed.ncbi.nlm.nih.govreturned a cookie-wall shell rather than the record when we tested it on 27 August 2026; the NCBI eutils endpoint (eutils.ncbi.nlm.nih.gov/entrez/eutils/efetch.fcgi?db=pubmed&id=PMID) returns clean structured data instead. - Check the whole author list, the year and the journal. The most common near-miss is a real paper with a real first author and a wrong year, wrong journal or a co-author who was never on it.
- Open the paper and find the sentence. This is the step everyone skips and the only one that catches misattribution. If the paper exists but does not say what you cited it for, you have a real reference supporting a false claim, which is worse than an obvious fake because nothing flags it.
20. Verification pass over a draft
Here is my reference list and the passages that cite each entry. For every
entry, tell me exactly what I must check by hand: the DOI, the exact title
string to search, the full author list as I have written it, and the specific
claim in my text that this reference is supposed to support. Return it as a
checklist. Do not tell me whether any reference is real. You cannot know that.
21. Quote-and-page extraction for one claim
I claim: "[claim]". I cite S[n] for it. From S[n] only, quote every sentence
that bears on this claim, with its page or section. Then state whether those
sentences support the claim as written, support a weaker version, or do not
support it. If it is a weaker version, write the weaker version out for me.
What must you never ask a model to do here?
Three things, and they are the three things people most want.
Never ask it to find papers. Finding literature is an indexed retrieval problem. A chat model has no index. Web-connected modes retrieve real pages, which removes the invented-DOI failure but not the misattribution one, and a handful of retrieved pages is not a search strategy you could describe in a methods section.
Never ask it for a citation from memory. "Give me a source for this" is the prompt that generates fiction. The correct move is inverted: you bring the source, and the model tells you whether it supports your sentence.
Never ask it what the literature says. It does not know. It knows what text about the literature tends to look like. That is a genuinely different thing, and it is the difference between a review and a plausible essay about a review.
Everything else on this page sits inside those lines. If you want a broader set of research prompts built the same way, 30 AI prompts for literature review and research synthesis covers adjacent jobs, and the same generator-plus-verification pattern shows up in a different domain in the code review prompt generator.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.
Create An AccountWhere this leaves the tool, and us
Prompt Architects generates prompts. That is the honest scope. We are not a research database, we do not find papers, we do not index the literature, and we cannot verify a citation for you. Nobody can do that last one for you except a database and your own eyes.
What a generator is good for is the part that is tedious and easy to get wrong: writing the same five slots every time, remembering the evidence rules, and keeping the verification line attached to every prompt that touches a citation so it does not quietly fall off when you are tired at 1am. Save the skeleton, fill the variables per project, and the discipline survives contact with a deadline. Our free plan covers 5 prompt enhancements per day, forever, according to the FAQ page at prompt-architects.com/faq, checked 27 August 2026, which is enough to build and refine a working set.
The reason this page spends a third of its length on fabrication is that the alternative version of this article, the one that opens with "10 amazing prompts to write your literature review in minutes", is the reason people submit reference lists containing papers that do not exist. A literature review is an argument about evidence. A model can help you handle evidence you have gathered. It cannot gather it, and it cannot vouch for it, and a tool that pretends otherwise is selling you a retraction.