Back to blog
Industries18 min read

20 AI Prompts for Performance Reviews and Feedback

20 performance review prompts for managers writing feedback for their reports, built around an evidence-first workflow so every claim stays yours.

NH
Nafiul Hasan
Founder, Prompt Architects

TL;DR: These performance review prompts are for managers writing reviews and feedback for their own reports, not for someone drafting their own self-review. The workflow keeps every fact yours: you supply the dated, specific evidence, and AI only turns it into clear, consistent wording. It never decides how well someone performed, and it doesn't catch its own biased phrasing, you still have to run that check yourself.

Is This the Right Post for You?

The distinction matters beyond who's typing. A self-review only exposes the writer's own work history. A manager's review exposes someone else's: their compensation trajectory, their standing on the team, sometimes their disciplinary record, all decided by a person who isn't the one being written about. That's a different, higher-stakes confidentiality problem, and it's why the workflow below spends more time on redaction and calibration than a self-review would need to.

What Actually Goes Wrong When Managers Use AI for Performance Reviews?

Two failure modes account for almost everything that goes wrong here, and neither is exotic.

The first is invented specifics, the same hallucination risk that shows up anywhere a model is asked for something "specific and impressive" with thin input. Ask it to "write a strong review" for someone without supplying real evidence, and it will manufacture a percentage, a project name, or a deadline to fill the gap, because sounding plausible is what it's optimizing for, not accuracy. That invented detail then sits in a document with your name on it as the evaluator, not the model's.

The second is biased phrasing, and it's less obvious because it doesn't look like an error, it looks like normal management language. Large language models are trained on enormous volumes of human-written text, including exactly the kind of personality-focused, gendered feedback that research on real workplace reviews has already documented.

Writing in Fortune, Snyder found: "Abrasive alone is used 17 times to describe 13 different women." Among comparable trait words, only "aggressive" showed up in men's reviews at all, and it appeared "twice with an exhortation to be more of it." That study is over a decade old, and the underlying pattern, personality-based criticism landing on women where men get skill-development suggestions for the identical behavior, is exactly the kind of thing a language model will happily reproduce if you ask it to make feedback sound more direct or more forceful, because that phrasing is common and unremarkable in its training data. It won't flag it as biased on its own. You have to ask it to check, specifically, every time.

None of this is a reason to avoid AI here. It's a reason to be precise about what job it's allowed to do: evidence comes from you, the model drafts wording, and every claim that reaches an HR file is yours, not the model's, whichever of you typed the sentence. And whatever produced the wording, the finished document is still subject to your company's own non-discrimination and HR policies. If your organization has a specific policy on using AI for personnel decisions, that policy governs here, not this post.

What's the Rule Before Any Prompt Goes Into Your Review?

There's a simple test for whether a given prompt is safe: could you point to the source of every fact in what comes back? "Write a strong review for a senior engineer who's had a good year" fails immediately, because the model has to invent the engineer's actual work to comply. "Turn these five dated notes into a STAR-format statement, using nothing beyond what's here" passes, because nothing new enters the text.

Every prompt below is written to fail loudly rather than fill a gap quietly: instructed to flag what it can't support instead of inventing something plausible in its place. That instruction has to be repeated in each prompt, not stated once and assumed to stick, because a model that behaves carefully on your first message tends to drift back toward filling gaps the longer a conversation runs.

Which Prompts Turn Your Notes Into Structured Evidence?

Before any of this touches wording, get the raw material down: what your report actually did, when, and what happened as a result. Dates and facts only, no framing yet.

1. Group scattered notes into themes

I'm going to paste my notes on {REPORT NAME}'s work from this
review period. Group them into 4-6 themes based on what they
have in common. Don't add anything I didn't give you, don't
rate or rank the themes, and don't guess at impact I haven't
described. Just organize what's here.

{paste your raw notes: projects, dates, specific incidents}

2. Pull the concrete outcome out of a vague note

For each item below, tell me what changed because of it, using
ONLY the details I've given you. If I haven't stated an outcome
or a number, do NOT invent one, write "outcome not specified"
instead. Flag anything where you genuinely can't tell what
happened.

{paste the themed notes from the previous prompt}

3. Flag evidence from outside the review period

This review period runs from {START DATE} to {END DATE}. Look
at the list below and flag anything dated outside that window,
or anything with no date at all, so I can decide whether it
still belongs in this review.

{paste your evidence list with dates}

4. Merge your notes with peer feedback

Below is my own list of {REPORT NAME}'s work, followed by
feedback I collected from their peers. Merge them into one list,
labeling each item "my observation" or "peer feedback from
{ROLE, not name}." Don't blend two separate comments into one
claim, and don't infer agreement between sources that weren't
describing the same thing.

{paste your notes, then paste peer feedback with names replaced
by role labels}

How Do You Write Specific, Defensible Praise?

Vague praise is forgettable and, worse, it's the version most likely to get inflated into an invented claim. Specific praise, tied to something you can point to, does more for the reader and survives a calibration meeting.

5. Turn one accomplishment into a specific statement

Turn this into a specific, 2-3 sentence recognition statement:
what {REPORT NAME} did, why it mattered, and what happened as a
result. Use only the facts below. Don't add a scale word like
"exceptional" unless the evidence here actually supports it.

{paste one accomplishment with its outcome}

6. Credit the behavior, not just the outcome

Rewrite this so it names the specific action {REPORT NAME} took,
not just the result. A reader should be able to tell what they
actually did differently, not just that something good happened
around them.

{paste a vague praise note, e.g. "the project went well because
of them"}

7. Translate a technical accomplishment for a non-technical reader

Rewrite this for a manager two levels up with no context on the
technical details. Keep every fact and number exactly as
written. Explain only what changed for the team, the customer,
or the business because of it.

{paste one technical accomplishment}

8. Draft promotion-readiness bullets

Using only the evidence below, draft 4-6 bullet points that
describe work at the next level up for {ROLE}. Don't claim
{REPORT NAME} "is ready for promotion," that's a leveling
decision, just state what the evidence shows and let me draw
that conclusion myself.

{paste the strongest 3-5 pieces of evidence from this cycle}

How Do You Write Constructive Feedback Without the Bias Traps?

This is where the discipline matters most, because trait-based language rarely feels biased while you're writing it. It just feels like normal, direct feedback.

9. Turn a personality note into a behavior note

Rewrite the note below so it describes one specific, observable
action and its impact, not a personality trait. Don't use any of
these words or close synonyms: abrasive, aggressive, bossy,
emotional, difficult, intense. If the note doesn't actually
contain a specific action, say so instead of inventing one.

{paste a personality-focused critical note, e.g. "comes across
as too aggressive in meetings"}

10. Audit your own draft, sentence by sentence

This is feedback I drafted myself. For each sentence, tell me
whether it critiques a specific action or a general trait or
tone. For anything that critiques a trait, suggest a version
that names the specific behavior and its effect instead, using
only the situation I describe below.

{paste your own draft feedback, plus a short description of the
actual incident it's based on}

11. Run the swap test

Read the feedback below. If the person's name and role were
swapped for someone of a different gender doing the identical
thing, would you expect this phrasing to change? Flag any
sentence where the answer is yes, and explain specifically what
would likely change.

{paste your draft feedback}

12. Turn a growth area into a forward-looking statement

Turn this growth area into a forward-looking statement: what
specifically {REPORT NAME} could do differently next quarter,
based only on the situation described. Don't add a diagnosis of
why they behave this way, and don't guess at a cause I haven't
given you.

{paste the specific incident or pattern, with dates if you have
them}

13. Write in plain language for a non-native English speaker

Rewrite this feedback in plain, literal language: short
sentences, no idioms, no sarcasm, no figurative phrases. Keep
every fact identical. This is for someone highly competent in
their field who reads English as an additional language, not
someone who needs the substance simplified.

{paste your draft feedback}

How Do You Keep Feedback Consistent Across Your Whole Team?

A single review can look fair in isolation and still be unfair in aggregate, if the same behavior gets described differently depending on who's exhibiting it. Calibration catches that pattern before your reports compare notes.

14. Compare tone across two drafts

Below are draft review comments for two people at the same
level, doing comparable work this cycle. Compare the tone and
directness of the language only, not the underlying performance.
Flag any place where similar behavior seems to be described more
harshly, or more gently, depending on who it's about.

{paste draft A, then draft B}

15. Check whether your ratings match your own language

Here are my draft written comments for {N} reports, each labeled
with the rating I gave them. Read the language only. Flag any
case where a comment reads more critically than the rating it's
attached to would suggest, or more glowingly than a lower rating
would suggest.

{paste each report's rating and comment, names replaced with
labels A, B, C...}

16. Standardize length and format, not substance

Rewrite each of these so they follow the same structure:
strengths first, then growth areas, then next steps, roughly
{X} sentences each. Don't add or remove any substance, just make
the format consistent across all of them.

{paste multiple reports' worth of draft comments}

What Do You Do With a Genuinely Underperforming Report?

This is the case where vague language does the most damage, in both directions: too soft and the person never understands the gap is serious, too personality-focused and it reads as a character attack instead of a specific, fixable problem.

17. State the gap directly and factually

Turn this into direct, factual language about a specific missed
expectation: what the expectation was, what actually happened,
and the gap between them. No softening language, no personality
judgment, and nothing beyond what I've described below.

{paste the specific expectation, the specific outcome, and any
dates}

18. Draft next steps without inventing a consequence

Draft a "next steps" section using only what's agreed below: the
specific behavior or output that needs to change, how it will be
measured, and the timeframe. Don't add a consequence or an
ultimatum I haven't specified.

{paste the agreed expectation, how you'll measure it, and the
timeframe}

What's the Last Check Before You Hit Submit?

19. Run a bias and trait-word scan on the whole draft

Read this full review draft. List every word or phrase that
describes {REPORT NAME}'s personality or character rather than a
specific action (for example: "not a team player," "too quiet,"
"hard to read"). For each one, suggest a version that names the
specific behavior instead, or flag that none was given.

{paste your complete draft review}

20. Run the defend-it-line-by-line check

Read this draft skeptically. List every sentence that states a
rating, a comparison to peers, or a specific claim about impact
that isn't backed by a concrete example above. Don't rewrite
anything, just flag what needs a specific example before I
submit this.

{paste your complete draft review}

What Does This Look Like End to End?

A short, illustrative walk-through, with a placeholder name standing in for anyone's actual report, not a real case.

Say your raw note from prompt 1 reads: "Fixed the export bug that was crashing for enterprise customers. Led the migration off the old billing schema. Two peers mentioned she was hard to work with during crunch." Prompt 1 groups that into themes: technical delivery, project ownership, and a peer-perception note that doesn't fit cleanly anywhere yet. Prompt 2 asks what changed because of each item, using only those facts, and correctly flags the peer-perception note as having no stated outcome, just a characterization.

That's where prompt 9 does its job. Fed the raw phrase "hard to work with," it can't invent the specific behavior behind it, so it asks you to supply one. You recall what actually happened: in a planning meeting on March 12, she pushed back directly when a deadline moved, and didn't raise it again outside that meeting. Prompt 9 turns that into "in the March 12 planning meeting, pushed back directly on a deadline change, without escalating outside the meeting," which describes one specific, checkable action instead of a trait.

Prompt 11, the swap test, is the check that catches what prompt 9 alone might miss: read against the same scenario with a man's name in place of hers, doing the identical thing, would "hard to work with" have been the first word reached for, or would it have read as direct and appropriately assertive? If the honest answer is that the framing would likely soften, that's the signal to keep the behavior-specific version and drop the trait label entirely, not just reword it.

Prompt 20, the final defend-it check, then reads the finished sentence against the evidence behind it: a specific date, a specific meeting, a specific action, nothing invented and nothing left as an unexamined personality judgment. That's the version that goes in the file.

What Mistakes Do Managers Actually Make Here?

Four show up repeatedly across teams that adopt this workflow, and none of them require bad intent, just a shortcut under deadline pressure.

The first is pasting a thin note and asking the model to "make the feedback sound stronger" or "more direct," with no instruction against inventing specifics. That single vague prompt is the fastest route to a fabricated example, because a model asked for something more forceful will manufacture the missing detail rather than admit it doesn't have one.

The second is running the workflow once and treating the first draft as final. The defend-it check in prompt 20 exists because a first pass routinely produces at least one sentence that reads more confident, or more critical, than the underlying evidence actually supports, even when the earlier prompts each included a no-invention instruction. Skipping that last check is how that sentence reaches an employee's file instead of getting caught by the person who's actually accountable for it.

The third is skipping the calibration pass across the team and assuming that because each individual review looks fair on its own, the set of them is fair together. A trait-based note that reads as a minor observation about one report can read as a defining criticism about another, described in near-identical language, and the only way to catch that is the side-by-side check in prompts 14 through 16, not a second read of any single review in isolation.

The fourth is pasting a report's real name, a peer's real name, or a compensation figure into a general-purpose chat window out of habit, because that's how the tool gets used for everything else, without stopping to ask whether this specific document, about someone else's job and pay, belongs in that tool at all.

What Should You Never Paste Into an AI Tool as a Manager?

Everything below is either your report's sensitive information or a third party's, and none of it is necessary for the prompts above to work. Every one of them runs fine on redacted, role-labeled evidence.

What you havePaste as-isRedact first
Your own dated notes on work and outcomesYes—
A colleague's name used as a peer-feedback source—Replace with a role label like "Peer A"
Compensation, pay-band, or bonus figures—Never paste
Disciplinary history or an active HR case—Never paste
Medical, leave, or accommodation details—Never paste
Confidential business metrics tied to the note—Replace with a general description

If you're not sure whether your company's AI policy allows any of this for the tool you're using, that's a question for whoever owns that policy, not a guess to make on your own. Redacting is a five-minute pass: swap the report's name and any peer's name for a role label, swap project codenames for something generic, and swap any comp or headcount figure for a round placeholder, in your own notes before any of it reaches a chat window. Un-redact only in your final document, after the drafting work is done and outside the AI tool entirely.

How Do You Keep This System Ready for Next Review Cycle?

The habit worth keeping year-round isn't a prompt, it's the running note in step 1: capturing what each report shipped, closed, or handled well as it happens, instead of reconstructing six months of it from memory the week reviews are due. Everything above works whether you have that note or you're piecing evidence together under deadline pressure, but the note turns an evening of recall into ten minutes.

The prompts themselves are worth keeping somewhere you'll actually find them next cycle, the same way turning a good ad-hoc conversation into something you keep beats rebuilding a prompt template from scratch every time you need it. A personal prompt library saved once and reused across your whole team's review cycle means next quarter starts with the workflow already built, not with remembering what worked last time.

None of this makes writing reviews faster the first time through: gathering real evidence, checking it for bias, and running the defend-it check still takes longer than asking a model to "write a strong review" and copying whatever comes back. What it buys instead is a set of documents you can stand behind, one report at a time, in front of the person who actually did the work and the calibration meeting that comes after.

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.

Create An Account

Frequently asked questions

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.

Create An Account