TL;DR: Rubric prompting means grading a draft against named, weighted criteria on a fixed scale, with evidence required for each score, before you ship it. Building one is straightforward. Trusting the number it returns is not: research on LLM-as-judge scoring documents position, verbosity, and self-enhancement bias, and self-grading removes even the checks that research relies on.
What Is Rubric Prompting?
Rubric prompting is the practice of attaching an explicit scoring rubric to a prompt, then asking the model to grade the draft it just produced against each criterion, on a numbered scale, before returning anything as final. Instead of make this better, you hand the model a fixed list of what better means for this specific piece of work: a handful of named criteria, a weight or rank for each, a scale used the same way across all of them, and a rule that every score needs a specific quoted line from the draft as evidence.
The output isn't a single polished answer. It's a scored table, a short justification per row, and a revision that touches only what scored below your threshold. That last constraint matters more than it looks. An open-ended make it better pass tends to rewrite everything, including the parts that were already fine, and every unnecessary rewrite is a fresh chance to introduce an error that wasn't there a moment ago.
Why Does a Named Rubric Work Better Than a Vague Instruction?
Because good, professional, and on-brand are not checkable. A model asked for a vaguer better version has to guess which axis you meant, and it usually guesses by making the draft longer, more hedged, or more confident, none of which is the same thing as more correct. A rubric turns each vague adjective into a specific, checkable claim: not professional tone but does this read like a peer wrote it to another peer, with no exclamation points and no résumé-style self-praise.
The other advantage is reuse. A rubric written once for a cold outreach email, or a code review comment, or a client status update, gets pasted into that same kind of prompt every time you produce that content type, which is worth more than it sounds. Two people scoring the same draft against the same five named criteria will disagree less than two people independently asked whether it's good. A model scoring against the same criteria across a hundred drafts applies a consistent, if imperfect, bar, rather than a fresh, unstated standard every single time.
How Do You Structure a Rubric You Can Actually Score Against?
A rubric that produces a trustworthy-looking table and a rubric that produces a checkable one are built differently, and the difference is five specific parts.
| Element | What it forces | What happens if you skip it |
|---|---|---|
| Named, checkable criterion | Turns a vibe into a yes/no test | The model scores against its own idea of good, which drifts between runs |
| Weight or rank | Tells the model which failures matter most | A stray typo and a factual error look equally bad on the report |
| One fixed scale, used everywhere | Makes scores comparable across drafts and runs | A 1-to-5 scale mixed with pass/fail produces a table nobody can total |
| Evidence requirement (quote the line) | Forces the model to point at something real, not assert a feeling | Tone: 4/5 with no reason is a number you can't check or trust |
| Pass threshold | Decides when to stop revising | The model improves a draft that already passed, adding new claims along the way |
The same five parts work whether the content is a blog draft, a cold email, or a pull-request comment. Only the first column, the named criteria themselves, changes by content type:
- Blog draft: Does the title match what the body actually delivers? Is every statistic attributed to a named, checkable source? Does each section answer its own question in the first two sentences?
- Cold outreach email: Does the first line reference something specific to this recipient, not a swappable template variable? Is there exactly one ask? Would a stranger reading only the subject line know what this is about?
- Code review comment: Does every claimed bug name the specific line or input that triggers it? Is a suggested fix included, not just the complaint? Is severity, blocking versus nitpick, stated explicitly rather than implied?
How Do You Turn a Rubric Into a Prompt?
The whole pattern runs in one prompt, as three instructions stacked in order: draft, score, revise.
TASK: [describe the actual task and paste any source material]
Write a first draft that completes the task above.
Then, score your OWN draft against this rubric. Use a 1-5
scale for every row. For each score, quote the exact
sentence from your draft that earns or costs the point. Do
not give a score with no quoted evidence.
RUBRIC:
1. [Criterion 1] (weight: high)
2. [Criterion 2] (weight: high)
3. [Criterion 3] (weight: medium)
4. [Criterion 4] (weight: medium)
5. [Criterion 5] (weight: low)
Present the scores as a table: Criterion | Score | Quoted
evidence.
Then revise the draft, but ONLY to fix rows that scored 3
or below. Do not touch rows that scored 4 or 5, and do not
add new claims, statistics, or sources while revising.
Return: the rubric table, then the final revised draft.
The instruction to quote evidence for every score is doing more work than it looks like. A model asked to grade without that constraint often returns a plausible table of 4s and 5s with nothing to check any of it against. Forcing a quote per row doesn't guarantee the score is right, but it gives the human reading the table something concrete to spot-check in ten seconds, instead of a number you either trust blindly or ignore outright.
Can You Trust the Score the Model Just Gave Itself?
Not as an independent check, no. This is the part of rubric prompting worth being honest about, because the technique's whole appeal is a number that looks like measurement.
The most cited study of models grading other models' work is Zheng and colleagues' 2023 paper, Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, which tested GPT-4, GPT-3.5, and Claude-v1 as judges scoring pairs of answers produced by other models. The paper defines position bias as "when an LLM exhibits a propensity to favor certain positions over others", and found it severe: with the identical two answers presented in swapped order: "Only GPT-4 outputs consistent results in more than 60% of cases." Verbosity bias, defined as "when an LLM judge favors longer, verbose responses, even if they are not as clear, high-quality, or accurate as shorter alternatives", was harder still for weaker judges to resist: padding a response with restated, content-free items produced a favorable verdict from Claude-v1 and GPT-3.5 in 91.3% of cases, against 8.7% for GPT-4.
A third finding speaks directly to rubric prompting. The paper adopts the term "self-enhancement bias" from social cognition literature, using it "to describe the effect that LLM judges may favor the answers generated by themselves", and measured GPT-4 favoring its own answers with a 10% higher win rate and Claude-v1 with a 25% higher win rate, though the authors are careful to add that with limited data and small differences, the study "cannot determine whether the models exhibit a self-enhancement bias" beyond doubt.
Read that caution carefully, because it cuts both ways for rubric prompting. That paper's self-enhancement test compared a judge's own answer against a different model's answer to the same question, a genuinely adversarial setup with something else to weigh it against. Rubric prompting, in the single-prompt form above, is a harder case than that: there is no competing answer at all, just one draft grading itself, in the same conversation, with the same beliefs that produced it in the first place. If a model didn't know a claim was unverified when it wrote it, asking it to grade its own accuracy row doesn't hand it information it didn't have.
The paper's finding on math and reasoning grading makes the same point from another angle. It found that even when a judge model "can solve the problem (when asked separately)", it "was misled by the provided answers" and graded an already-solvable problem wrong once the check was framed around someone else's work. Grading your own draft against a rubric is a structuring device, useful for catching gaps against criteria you named in advance. It is not a substitute for a source the model never had.
How Do You Make the Self-Score Less Worthless?
The same paper tested several fixes for its judges, and two of them translate directly into a solo rubric prompt.
Grade in a fresh conversation, not the one that wrote the draft. The paper's fix for position bias was to run the comparison twice with the order swapped, and accept a verdict only if it held both times, calling anything else a tie. You can't swap positions when there's only one draft, but you can approximate the idea: paste the finished draft into a new chat with no memory of writing it, and run the rubric there. It won't remove self-enhancement bias outright, since it's still the same model, but it does remove the specific contamination of a critique pass following its own draft in the same context, primed to defend what it just said.
Give it something to grade against, not just criteria to apply. The paper's biggest fix for grading errors wasn't chain-of-thought reasoning: "even with the CoT prompt," it found the judge model "makes exactly the same mistake as the given answers in its problem-solving process". What worked was reference-guided grading, generating an independent answer first and then using that as the reference the judge compares the draft against, which cut the failure rate "from 70% to 15%" over the default prompt. For rubric prompting, the nearest equivalent is few-shot anchoring: attach one worked example of a draft that would score a 5 on your hardest criterion, and one that would score a 2, so the model has something beyond its own unaided judgment of what a passing score looks like. Few-Shot vs Zero-Shot Prompting covers when that extra length earns its keep.
Run it more than once and watch for disagreement, not just the score. Self-consistency across three separate scoring runs won't tell you a score is correct. It will tell you whether the model's own judgment is stable, and a criterion that comes back a 5, a 3, and a 4 across three identical runs is telling you something a single confident-looking 4 never would: that this particular call is closer to a coin flip than a measurement.
Is Rubric Prompting the Same as Red-Teaming Your Own Prompt?
No, and the two are worth running together rather than picking one. Red-Team Your Own Prompt Before You Trust the Output is adversarial and unstructured on purpose: rephrase the question, invert the framing, plant a false premise, see what breaks. It's built to surface a failure you didn't anticipate.
A rubric is the opposite instinct: a fixed, repeatable bar you decide on in advance and apply the same way to every draft of that content type, so a Tuesday review and a Friday review use the same named criteria instead of whatever the reviewer happens to notice that day. Red-teaming finds problems you didn't know to look for. A rubric checks for the ones you already decided matter, every time, without forgetting one. For anything with real stakes, run the rubric first, since it's cheap and fast, then red-team whatever passes. The same discipline of catching a gap before you hit send is behind The Prompt Hygiene Checklist, which runs against the prompt itself rather than the output it produces. And when the failure you're worried about is a made-up fact rather than a structural one, Why Does ChatGPT Make Things Up? covers the mechanism a rubric can't reach: a model doesn't withhold a guess just because you asked it to grade its own honesty.
Where a Rubric Fits Next to a Built-In Score
Prompt Architects' own Quality Score, on the Advanced plan and above, is a related but different check: it scores a prompt's clarity before you run it, not the output the prompt produces afterward. A rubric like the ones above sits on the other side of that line, judging what actually came back against criteria specific to that piece of content. Neither replaces the other. A clear prompt scored well by Quality Score can still return a draft that fails your tone criterion on a bad day, and a vague prompt can occasionally get lucky anyway. Score the prompt going in, if the tool offers it. Score the draft coming out regardless.
If you'd rather compare two prompt variants than gate a single draft, that's a related but different job: A/B Test Your Prompts covers blind-scoring output from Prompt A against Prompt B to learn which one to keep, using a rubric of its own.
None of this makes rubric prompting worthless, it just narrows what it's honestly for. It replaces a vague make it better with a checkable, reusable bar, and gives you a quoted line to check instead of a number to trust blindly. It does not replace a second reader, and it does not create knowledge the model didn't already have.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.
Create An Account