TL;DR: ChatGPT sometimes answers from a live search and sometimes from frozen training memory, and it rarely announces which one you are looking at. Fact-checking means isolating every specific, checkable claim and verifying each one at its primary source, not by asking the same chat to confirm itself. A federal court record shows exactly why that shortcut fails.
Why "are you sure?" is not a fact check
In May 2023, a lawyer named Steven Schwartz found himself explaining to a federal judge why his court filing cited cases that did not exist. His defense, laid out in his own affidavit, was that he had checked. He had opened the same ChatGPT conversation that generated the fake cases and asked it two direct questions: "Is Varghese a real case" and "Are the other cases you provided fake". Both times, according to Judge P. Kevin Castel's opinion, ChatGPT told him the cases were real and could be found on Westlaw and LexisNexis.
They could not. The opinion, Mata v. Avianca (S.D.N.Y., decided June 22, 2023, docket 1:22-cv-01461), is a real sanctions order against a real law firm, and it is worth reading in the primary source rather than a summary of it, because the detail that matters is easy to lose in retelling. The attorney did fact-check. He just fact-checked using the tool that had invented the problem, inside the conversation where it had already committed to the fabrication. It confirmed its own mistake twice.
Judge Castel's opinion is also more measured than the case's reputation suggests. It states plainly that "there is nothing inherently improper about using a reliable artificial intelligence tool for assistance." The sanction was not for using ChatGPT. It was for never checking outside it. That distinction is the entire subject of this post.
Is the answer in front of you grounded, or generated?
Before you can fact-check a ChatGPT answer, you need to know which kind of answer it is, because the two fail differently.
A grounded answer draws on something outside the model's training: a live web search, an uploaded file, a connected app. ChatGPT Search is a named, separate surface with its own trigger, and OpenAI's Deep Research mode explicitly works from the public web, uploaded files, and connected apps, proposing a research plan you can edit before it runs. When an answer comes with citations or a visible search step, at least part of it is checkable against something outside the chat.
A generated answer draws only on the model's trained parameters, patterns learned during training, frozen at a cutoff date, with no live lookup involved. Ask a plain chat question with no citations and no visible search, and you are getting the model's best guess from memory. That memory cannot know about anything that happened after training ended, and it cannot distinguish a well-supported fact from a plausible-sounding one from the inside, because both are stored the same way.
The uncomfortable part is that even a grounded answer is not automatically a checked one. OpenAI states only that browsing reduces the rate of fabrication, not that it eliminates it, and publishes no numeric accuracy figure for either mode. Grounding narrows where an error could have come from. It does not certify that there isn't one. If you want the fuller mechanical explanation of why answers get invented at all, we cover that separately in why your ChatGPT answers are bad and in why ChatGPT makes things up.
The same caution applies to an AI-generated search summary sitting above the results on Google or elsewhere. It looks like a citation because it often links a page. It is still a compressed paraphrase produced by a model, and a summary can drop a qualifying clause the source page actually contains. An AI Overview is a lead to follow, not a source to cite.
This is, incidentally, the same discipline professional fact-checkers and reporters already practice: a claim does not get published on the strength of one source repeating it confidently, however fluent that source sounds. It gets published once someone has traced it back to where it was actually produced. Treating a chatbot's answer the same way, as a lead rather than a verdict, is not a lower standard than professional practice. It is the same standard, applied to a new kind of source.
What actually needs checking
Not every sentence in an answer carries equal risk. The claims worth isolating are the ones with enough specificity to be wrong in a checkable way:
- A named person, case, organization, or paper
- A statistic, percentage, or dollar figure
- A direct quote attributed to someone
- A date, version number, or "as of" claim
- A claim about a product's current price, plan, or feature set
- Anything the model attributes to a named, findable source
Broad, well-established background, how a for loop works, what a mortgage is, carries far less risk than an obscure name paired with a specific detail, because there is little room for the model to have quietly guessed at the broad case.
A useful first move is making the model itself produce the checklist, without asking it to grade the checklist:
List every specific, checkable claim in your last answer: named
people, organizations, cases, papers, statistics, direct quotes,
dates, and version numbers. For each one, state it as a single
short line. Do not tell me whether each one is true — just extract
the list so I can check it myself.
This is not verification. It is triage. The model is good at finding its own claims. It is not a reliable judge of whether those claims are real, which is exactly what the Mata screenshots demonstrate.
Two habits of mind that let a false claim through
Beyond the mechanics, two ordinary reading habits do most of the damage before you even get to checking anything.
Specificity bias. A claim that says "usage grew substantially last year" reads as a guess. A claim that says "usage grew 34% last year" reads as a fact, because precision feels like evidence. It is not. A language model can generate a precise-sounding figure exactly as easily as a vague one; nothing about the number's exactness reflects how it was produced. Treat an unusually specific number as a claim that needs checking more, not less, than a vague one.
Repetition bias. If you have seen a similar claim somewhere before, whether from another chat, a search summary, or a blog post, a new instance of it feels pre-confirmed. It often is not. Aggregator sites frequently repeat each other's numbers without anyone tracing them back to an original source, so seeing the same figure five times can mean there is one unverified root and four copies of it, not five independent confirmations.
A worked example, to show the workflow in practice
Here is a hypothetical, invented for this walkthrough rather than a real reported figure: suppose a chat answer tells you, "A 2024 university study found that teams using structured prompt libraries save 11 hours per week." That single sentence contains four separate things to isolate: an institution ("a university"), a year (2024), a specific number (11 hours), and an implied causal claim (using the library caused the saving).
Applying the workflow: the institution is too vague to check as written, so the first real step is asking the model to name the specific university and the study's title, which is itself a test, because a real citation usually has a name attached and a fabricated one is often generic on purpose. If a name comes back, the next step is finding that study on the university's own research page or the publishing journal, not on a blog that mentions it in passing. If no study turns up at the primary source, or the primary source describes something different from "11 hours per week," the claim gets marked unconfirmed. It does not get repeated with the university's name attached just because a chat produced one.
Notice what this workflow does not require: no special tooling, no subscription, no expertise beyond patience and a second browser tab. The cost of checking a specific claim is almost always smaller than the cost of repeating a false one, and the four-part breakdown above (institution, date, number, causal link) generalizes to nearly any specific-sounding sentence you will run into.
The actual verification workflow
Once you have the list, checking it takes five repeatable steps.
- Isolate the claim, word for word. Copy the exact name, number, or quote rather than your memory of it. A citation that is subtly wrong, one digit off, one word changed, is often as damaging as one that is entirely invented.
- Go to the primary source, not a summary of it. A quote gets checked against the actual document, searched for as a literal string. A court case gets checked on a docket site. A study gets checked on the journal's own page. A product claim gets checked on the vendor's own documentation, not a comparison blog that may itself be wrong.
- Confirm the source is what it claims to be. A page returning successfully is not proof it is the right page. A fetched document can be a paywall wall, a cookie-consent shell, or, on at least one recorded occasion in this project's own research, a completely different article than the one requested. Read the title of what actually loaded before you trust the body text.
- Trace a statistic to where it was produced, not where it was repeated. Numbers travel through blog posts and social threads faster than their original context does, and the qualifying detail, sample size, date, which product version, usually gets dropped somewhere along that chain.
- Mark what you cannot verify as unconfirmed, not as false. These are different outcomes. A claim you could not check has an unknown status. Repeating it as settled fact, or repeating a denial you also could not check, both overstate what you actually know.
| Claim type | Where to verify | What "checked" looks like |
|---|---|---|
| Court case or legal citation | A docket site (e.g. CourtListener) or the court's own records | You read the actual filing or opinion, not a description of it |
| Academic paper or statistic | The journal, the paper's own page, or the funder's registry | The number appears in the primary study, with its sample and date |
| Direct quote | The original document or transcript | The exact string appears verbatim on the page |
| Product price, plan, or feature | The vendor's own current documentation | Dated, and matched against the live page, not a cached description |
| Current-events or "as of" claim | A dated primary report | The date on your source is close to the date of the claim |
What if two sources disagree with each other?
This happens more often than it should, including between a vendor's own pages. When it does, the honest move is to report both positions, each attributed and dated, rather than silently resolving the conflict by guessing which one sounds more authoritative. A confident-sounding number from an aggregator is not evidence against a hedged number from the primary source. It is usually evidence the aggregator skipped the same step you are trying not to skip.
If you cannot resolve a contradiction, say so in whatever you are writing or deciding, rather than picking a side to make the sentence read cleaner. A flagged contradiction is more useful to the next reader than a false resolution, and it is far less embarrassing to correct later than a confident claim that turns out to have been the wrong half of a disagreement nobody noticed.
The habit that actually protects you
None of this requires distrusting every sentence a model produces. It requires knowing which sentences are specific enough to be wrong, and treating "I asked it to check" and "I checked" as two different verbs. The first one is what happened in Mata v. Avianca, twice, in the same conversation, with the same fabrication both times. The second one is five extra minutes with a primary source open in another tab.
If you want shorter, more direct answers to begin with, which reduces how much there is to check in the first place, see how to get shorter, sharper answers from ChatGPT and why your ChatGPT answers might be too short already. Neither replaces checking a specific claim. Both just reduce how often you have to.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.
Create An Account