TL;DR: A research question generator is a prompt chain, not a button. Feed it a vague topic and it narrows scope, names your population and outcome, sorts descriptive from causal, fits a real framework such as PICO, PICo or SPIDER, then stress-tests answerability. You still own the question.
What does a research question generator actually do?
It does the drafting and the stress-testing. It does not do the thinking.
That distinction matters more here than in almost any other prompting task, because the work of turning "I want to study social media and teenagers" into something a committee will approve is genuinely hard, genuinely taught, and genuinely yours to defend. What a model is good at is producing eight variants in ten seconds, filling a framework table, and playing a hostile examiner on demand. What it cannot do is know that your ethics board will not approve interviews with under-16s, or that your department has three people who already did the comparative version.
So the useful shape of a research question generator is not a box that emits five questions. It is a sequence: narrow, name, expose, classify, fit, test, translate. Twenty-one prompts below run that sequence. Every one is free to copy and needs no account.
The frameworks they use are real and old. PICO comes out of evidence-based medicine, PICo and SPIDER out of qualitative synthesis, FINER out of a clinical research textbook. They are cited properly further down, because a generator built on an invented framework is just a confident wrong answer with a table around it.
What makes a topic vague, and a question researchable?
A topic is vague when you cannot name what would count as an answer.
"Social media and teenagers" fails that test instantly. Which teenagers, at what age, in which country? Which platforms, used how much, measured how? An answer to what: wellbeing, sleep, grades, friendship quality, self-reported anxiety? Compared with whom? Over what period? Every one of those is a decision that the topic silently defers, and a research question is simply the sentence in which you stop deferring them.
The Cochrane Handbook makes the cost of getting this wrong concrete in both directions. Too broad and you retrieve too much: the Handbook warns that "it is important not to ask a question that will result in retrieving unmanageable quantities of information; up-front scoping work will help authors to define sensible boundaries for their reviews." Too narrow and there is nothing to find: authors "should be aware of the possibility of asking a question that may not be answerable using the existing evidence".
A researchable question has four properties you can check without methodological training. You can name the population. You can name what you are measuring or exploring. You can name where the data comes from. And you can describe, in one sentence, what the results section would look like. If any of the four is missing, you have a topic.
Which research question framework fits your design?
There is no single correct framework, and any page that tells you otherwise is selling something.
Booth and colleagues ran a rapid review of the field for BMJ Global Health in 2019 and reported that, "Following review of 1481 references and 113 full-text citations, we identified 38 question formulation frameworks". Thirty-eight. The Cornell University Library evidence synthesis guide puts it more gently, noting "almost 40 different types of research question frameworks" and warning that while PICO helps for clinical questions, "it may not be the best choice for other types of research questions, especially outside the health sciences."
Here are the five you will actually meet, with where each came from.
| Framework | Expands to | Built for | Attribution |
|---|---|---|---|
| PICO | Population, Intervention, Comparison, Outcome | Quantitative effectiveness questions | Components credited to Richardson et al. 1995 in the Cochrane Handbook |
| PICo | Population, Interest, Context | Qualitative questions about experience | JBI meta-aggregation guidance, per the JCU library guide |
| SPIDER | Sample, Phenomenon of Interest, Design, Evaluation, Research type | Qualitative and mixed-methods synthesis | Cooke, Smith and Booth, 2012 |
| SPICE | Setting, Perspective, Intervention, Comparison, Evaluation | Service and practice evaluation | Booth, 2006, per the JCU library guide |
| FINER | Feasible, Interesting, Novel, Ethical, Relevant | A quality test, not a question shape | Cummings, Browner and Hulley, Designing Clinical Research |
Note the last row. FINER is not a template you fill in; it is a screen you run the finished question through. You can pass FINER with a question written in plain prose and no acronym anywhere near it.
Where do PICO, PICo, SPIDER and FINER actually come from?
Worth knowing, because framework attributions in the prompting world are frequently wrong. We traced three popular prompt frameworks to their claimed origins in the CO-STAR framework post and the popular attribution failed in every case. Methodology frameworks are better documented, so here is the paper trail, fetched and read on 27 August 2026.
PICO. The Cochrane Handbook's chapter on determining the scope of a review says the specification of a review question "requires consideration of several key components (Richardson et al 1995, Counsell 1997) which can often be encapsulated by the ‘PICO’ mnemonic". Its reference list gives the source as "Richardson WS, Wilson MC, Nishikawa J, Hayward RS. The well-built clinical question: a key to evidence-based decisions." The NLM record for that paper confirms it: ACP Journal Club, 1995, volume 123, issue 3, pages A12 to A13, filed as an editorial. Because it is an editorial, PubMed carries no abstract, so the acronym itself is widely attributed to that paper rather than quotable from it. What is verifiable is that Cochrane credits those authors for the components.
SPIDER. Attributable precisely, from the abstract itself. Cooke, Smith and Booth published "Beyond PICO: the SPIDER tool for qualitative evidence synthesis" in Qualitative Health Research in 2012, and the abstract states the expansion in full: "SPIDER (Sample, Phenomenon of Interest, Design, Evaluation, Research type)". The same abstract is honest about its own limits, closing with the line that "To constitute a viable alternative to PICO, SPIDER needs to be refined and tested on a wider range of topics." Fourteen years on, that caveat is still worth repeating to anyone who treats SPIDER as settled.
PICo. This one I could confirm only at one remove. The Cornell and Murdoch University library guides both give the expansion as Population, Interest and Context, and Murdoch defines the middle term as "a defined event, activity, experience or process". The James Cook University guide traces it to Lockwood, Munn and Porritt's 2015 paper on meta-aggregation in the International Journal of Evidence-Based Healthcare, and the NLM record confirms that paper exists and is JBI methodological guidance. The full text sits behind a publisher paywall, so I could not read the acronym's coinage there. Widely attributed to JBI, then, and consistently expanded across three independent university guides.
FINER. Doubly confirmed. The Cochrane Handbook states that "The FINER criteria have been proposed as encapsulating the issues that should be addressed when developing research questions", citing "Cummings SR, Browner WS, Hulley SB. Conceiving the research question and developing the study plan." A 2010 methods paper in the Canadian Journal of Surgery independently says that Hulley and colleagues "have suggested the use of the FINER criteria in the development of a good research question", and reproduces the criteria as "FINER = feasible, interesting, novel, ethical, relevant". The one wrinkle: guides cite different editions of that textbook, so you will see the attribution dated 1988 and 2007 both.
The generator: 21 prompts, in the order you need them
Run these in sequence on one topic. Each stage feeds the next.
Stage 1 — Narrow the scope
1. Make it interrogate you first
You are a research methods advisor. I will give you a vague research topic.
Do not propose a question yet.
Instead, ask me the 8 questions you most need answered before a researchable
question is possible. Cover at minimum: discipline and methodological
tradition, the population I can actually reach, the outcome or experience I
care about, the data I already have or could collect, my timeframe and word
limit, and whether I need a descriptive, comparative or causal answer.
Ask all 8 at once, numbered. Then wait for my answers.
TOPIC: [paste your vague topic]
2. The scope ladder
Build a scope ladder of 6 rungs for my topic, from the broadest version that
is not researchable down to the narrowest version a single researcher could
complete in [TIMEFRAME].
For each rung give: the question in one sentence, the population it implies,
the data it would require, and why it is or is not feasible at my level.
Mark the rung with the best trade-off and justify the choice.
TOPIC: [...]
LEVEL: [undergrad dissertation / master's thesis / PhD chapter / pilot study]
TIMEFRAME: [...]
3. Draw the boundary explicitly
For the topic below, write an in-scope / out-of-scope table with 8 rows.
Dimensions: population, setting, time period, exposure or intervention type,
outcome, language, study design, geography.
Each row: what is IN, what is OUT, and one sentence justifying the cut.
Flag any cut that would leave too little evidence to answer the question.
TOPIC: [...]
Stage 2 — Name the population and the outcome
4. Population specification
My topic is [...]. Propose 5 candidate populations, each specified to the
level of detail a methods section needs: age band, setting, inclusion
criteria, exclusion criteria, and how hard each would be to recruit or access.
Rank them by feasibility for a researcher with [ACCESS AND RESOURCES].
Name the one you would drop entirely, and say why.
5. Outcome specification
List 6 candidate outcomes for my topic. For each: the construct in plain
words, the type of instrument that operationalises it, whether it is
self-reported or observed, and its main validity threat.
Then name the two outcomes most often conflated in this literature and
explain the difference.
Do not cite specific papers or instruments by name. Describe the categories
and I will find the sources myself.
6. Exposure, intervention or mechanism
For my topic, distinguish the exposure, the intervention and the mechanism.
One paragraph each.
Then write three versions of the same question: one treating the thing as an
exposure people vary in naturally, one as an intervention someone
administers, and one as a mechanism to be explained.
For each, name the study design it demands.
Stage 3 — Surface the hidden assumptions
7. Assumption audit
Below is my draft research question. List every assumption it smuggles in:
causal assumptions, assumptions about direction of effect, assumptions that a
construct is unitary, assumptions about who counts as the population, and any
normative judgement dressed up as description.
For each, rewrite the clause so the assumption becomes explicit and testable,
or remove it.
DRAFT QUESTION: [...]
8. Contested-term audit
Identify every contested or under-defined term in my question.
For each: the two or three competing definitions used in this field, what
changes about the study depending on which I adopt, and the definition you
would defend, with the trade-off stated plainly.
Do not attribute definitions to named authors. Give me the positions and I
will find who holds them.
Stage 4 — Descriptive, comparative or causal?
9. Classify the claim
Classify my draft question as descriptive, comparative, associational,
causal, or evaluative. Quote the exact words that make it that type.
If the wording claims more than my design can support, say so directly.
Then rewrite it at the strongest claim level the design can actually deliver.
DRAFT QUESTION: [...]
DESIGN I CAN RUN: [...]
10. Demote an overreaching causal claim
My question implies a causal claim. My design is [DESIGN].
1. State plainly whether that design can support a causal claim.
2. If it cannot, rewrite the question as the strongest non-causal version
that keeps what I actually care about.
3. List what would have to be added to earn the causal version:
randomisation, longitudinal data, an instrument, a natural experiment.
4. Say which of those is realistic for me and which is not.
11. Three sibling questions
Give me three sibling questions on the same topic:
A. a descriptive question answerable with data that already exists
B. a comparative question requiring two clearly defined groups
C. an explanatory question requiring qualitative depth
For each: the design, the data, the analysis, the strongest finding it could
produce, and the most likely reviewer objection.
Stage 5 — Fit a real framework
12. Choose the framework, or refuse to
Here is my question and my intended design.
Recommend ONE formulation framework from PICO, PICo, SPIDER, SPICE, PEO or
PCC, or tell me that none fits and I should write the question in prose.
Justify the choice against my design specifically.
Name what the framework will force me to specify, and name what it will make
me leave out.
If my work is ethnographic, interpretive, historical or theory-building, say
so and do not force PICO onto it.
QUESTION: [...]
DESIGN: [...]
13. Fill a PICO table
Populate a PICO table for my question.
Columns: element, my content, what is still under-specified, and the search
concept that element becomes.
Then restate the question as one sentence in standard PICO order.
If an element is genuinely absent, and many questions have no comparator,
write "not applicable" and explain why rather than inventing one.
14. Fill PICo and SPIDER side by side
My work is qualitative. Populate BOTH tables for my question:
PICo (Population, Interest, Context)
SPIDER (Sample, Phenomenon of Interest, Design, Evaluation, Research type)
Then tell me which produces the more searchable question for my topic and why.
State explicitly where the two frameworks disagree about what my study is.
Stage 6 — Stress-test it
15. FINER scoring
Score my research question against FINER: Feasible, Interesting, Novel,
Ethical, Relevant.
For each: a score out of 5, the specific wording in my question that earns
that score, and one concrete change that would raise it.
Be harsh on Feasible. Assume one person with [TIME], [BUDGET] and [ACCESS].
If the question fails Feasible, say that before anything else.
16. Work backwards from the results table
Act as a sceptical methodologist. For my question, work backwards:
1. What would the finished results table or thematic map look like?
2. What variables or codes must exist to fill it?
3. What data would produce those, and does that data exist?
4. Can I get access to it, legally and practically?
5. How large a sample, or how many interviews, would it take?
If any step breaks, stop there, tell me the question is not answerable as
written, and propose the nearest version that is.
17. Shrink it to fit
My question is currently a multi-year question. Shrink it to something
completable in [TIMEFRAME] with [RESOURCES] without making it trivial.
Give me three shrink strategies: narrow the population, narrow the outcome,
narrow the design.
Show the resulting question for each, and state exactly what I lose.
Stage 7 — Turn the question into a search
18. Concepts to search terms
Turn my finalised question into search concepts.
For each concept: the concept name, then 6 to 10 synonyms and lexical
variants including British and American spellings, plural forms, and the
older terminology the field used before roughly [YEAR].
Flag any term that means different things in different disciplines and say
what it means in each.
Do not produce a finished Boolean string yet.
19. Draft the Boolean string
Using the concept lists above, draft a Boolean search string for [DATABASE].
Combine synonyms within a concept using OR, and concepts using AND.
Use the truncation and field-tag syntax that database actually uses, and tell
me which syntax you applied.
Give me a broad version and a narrow version, and predict which concept will
generate the most irrelevant results.
I will run this myself and refine from the real result count.
Stage 8 — Verify before you trust any of it
20. Novelty search plan, not a novelty verdict
Do NOT tell me whether this question has already been answered. You cannot
know that.
Instead, write me a plan to find out: 4 databases to search, the exact
strings to run in each, 3 types of publication where a prior synthesis would
live, and 5 signals in a retrieved paper that would mean my question is
already answered.
End with one line confirming that this plan must be executed by me.
21. Hostile examiner
You are the harshest examiner on my committee. Read my question and write the
three objections you would raise in a viva: one about scope, one about
measurement, and one about what the finding would actually add.
For each objection, write the strongest defence I could give, then say
whether that defence holds.
Do not be encouraging. Do not praise the question.
Three vague topics, turned into researchable questions
The prompts are abstract until you watch them bite. Here is what the sequence produces on three topics that arrive in supervisors' inboxes every September.
Topic one: "I want to study social media and teenagers."
Stage 1 forces the platform, the age band and the country. Stage 2 forces you to pick sleep over the vaguer "wellbeing", because sleep has established measures and a defensible causal story. Stage 3 catches the assumption that use causes harm rather than distress driving use. Stage 4 catches that a survey cannot support a causal verb. The result: Among 14 to 16 year olds in UK state secondary schools, is self-reported nightly short-form video use associated with shorter school-night sleep duration, after adjusting for bedtime routine and household screen rules? Associated, not causes. That is the whole difference between a question you can defend and one you cannot.
Topic two: "Something about remote work and productivity."
The interesting failure here is Stage 3, which flags that "productivity" is doing two incompatible jobs at once, output per hour and perceived effectiveness, and that most disagreement in this literature is definitional rather than empirical. Stage 5 refuses PICO, because there is no intervention and no comparator group, and routes you to SPIDER instead. The result: How do knowledge workers in mid-size software firms describe changes in their sense of professional visibility after a shift to fully remote work, and what practices do they adopt to manage it? Sample, phenomenon of interest, design, evaluation, research type. PICO would have made you invent an intervention that never happened.
Topic three: "AI in education."
Stage 1 kills three quarters of this in one pass: which level of education, which use of AI, whose perspective, and are you studying the technology, the policy or the people. Stage 6 does the rest, because working backwards from the results table exposes that the version you wanted needs institutional data no student gets access to. The result after shrinking: What reasons do first-year undergraduates give for disclosing or not disclosing generative AI use in assessed written work, under a policy that requires disclosure? Interviewable, ethics-approvable, and finishable in one semester.
What can a model not do for your research question?
Two things, and both of them matter more than everything the prompts above do well.
It cannot tell you whether the literature already answers your question. A model without live retrieval is reconstructing plausible-sounding claims from training data that has a cutoff. A model with retrieval sees only what its particular search returned, which is not your library's subscriptions, not the grey literature, and not the thesis from your own department that nobody indexed. When a model says this appears to be an underexplored area, it is producing the sentence that fits, not a finding. Prompt 20 exists because of this: it asks for a search plan and explicitly forbids a verdict. Run the plan in real databases. If you want a model reading actual sources rather than recalling them, that is a different tool and a different workflow, covered in the post on source-grounded prompting in NotebookLM. For the reading and synthesis stage that follows, the literature review prompt set picks up where this page stops.
It will confidently invent citations. This is the well-documented failure mode covered in why ChatGPT makes things up, and question-framing is the workflow where it does the most damage. A fabricated reference can make your question look novel when a real synthesis already exists, or look answered when it is wide open. Either way you spend weeks going the wrong direction, and the reference survives into your proposal because it looked like every other reference on the page.
That is why several prompts above tell the model not to name authors or instruments. Removing the opportunity is more reliable than catching the output.
How do you check the model did not invent the literature?
Run this, on every single reference, before it enters a document.
- Search the exact title in your library discovery layer or Google Scholar. Not the author, not the topic. The title in quotation marks. A fabricated paper usually returns nothing, or returns a real paper with a different title.
- Resolve the DOI. Paste it after
https://doi.org/and load it. A fabricated DOI 404s. A real DOI that lands on a different paper than the one you were given is the more dangerous case, and it happens. - Check the journal actually published that volume and issue in that year. Journals change names, merge and fold. A plausible pairing of a real journal with an impossible year is a common fabrication shape.
- Open the abstract and confirm it says what you were told it says. This is the step people skip. A real paper attached to a claim it never made is invisible to every check above.
- Never let a reference into your proposal that you have not opened. Not the PDF necessarily, but the abstract page, with your own eyes.
If that sounds slow: five references at two minutes each is ten minutes, and one fabricated citation reaching a submitted proposal costs far more.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.
Create An AccountHow do you keep the question you landed on?
Write the final question, the framework table behind it, and the three rejected siblings into one document, and keep it. Supervisors ask why you did not do the comparative version, and the honest answer, with the feasibility reasoning attached, is worth more than the question itself.
Then save the chain. These 21 prompts are worth more on your second topic than your first, because by then your level, timeframe and access constraints are already filled in. If you would rather have the model interrogate you before it drafts anything, the technique generalises well beyond research questions and is covered in getting ChatGPT to ask you questions first.
Drafting eight candidate questions and stress-testing each against FINER used to be a week of supervision meetings. It is now an afternoon, and the week goes into the reading instead. That is the honest version of what a research question generator does: it accelerates the drafting and the demolition of a question you still have to own, defend and answer yourself.