How to Check If ChatGPT Is Right: A Two-Minute Routine (And When It Is Not Enough)
TL;DR
- A seven-step, two-minute triage for one ChatGPT answer that tells you to use it, flag it or escalate - not a verdict that it is true.
You cannot prove a ChatGPT answer true in two minutes. You can decide, in two minutes, whether it is safe to rely on, needs a flag, or needs a proper check. Seven moves get you there: name what a wrong answer would cost, check whether your question planted a false premise, find the claim the answer stands on, rate its risk, see whether anything was retrieved, read laterally for one independent source, and decide.
This is for one answer you are about to act on. It is not a publication gate — that is the 42-check fact-check checklist — and not a full source trace. It is the front door to both, and it works on any chat assistant.
Key Takeaways
- Two minutes buys a triage decision, not a verdict. The budget is a design target we set, not a time we measured.
- Check your own question first. Models "often fail to correct a user's incorrect legal assumptions" when a question is built on one (Dahl et al., arXiv:2401.01301, retrieved 2026-09-17).
- Check the late claims, not just the first sentence. Factual precision was "significantly worse" in the later part of generated text across every model tested (Min et al., FActScore, EMNLP 2023, retrieved 2026-09-17).
- Obscure and specific is where errors cluster. Accuracy tracks how often a fact appeared in training data (Kandpal et al., arXiv:2211.08411, retrieved 2026-09-17).
- Leave the answer to check it. Professional fact checkers reached better conclusions faster than historians and students by opening new tabs instead of reading deeper (Wineburg & McGrew, Teachers College Record, 2019, retrieved 2026-09-17).
The routine at a glance
| Step | Budget | The question | If it fails |
|---|---|---|---|
| 1. Name the stakes | 10 s | What happens if this is wrong? | High stakes: triage only, then escalate |
| 2. Check your question | 15 s | Did I assume something in how I asked? | Re-ask neutrally in a fresh chat |
| 3. Find the load-bearing claim | 20 s | Which single fact does the answer depend on? | If you cannot find one, the answer is not checkable yet |
| 4. Rate the claim's risk | 10 s | Is it specific, obscure or recent? | Move straight to step 6 |
| 5. See what was retrieved | 15 s | Does it cite anything, and does the page say it? | Treat as unsourced recall |
| 6. Read laterally | 40 s | Does one independent source agree? | Escalate |
| 7. Decide | 10 s | Use, use with a flag, or escalate? | — |
The budgets add up to 120 seconds, and they are a design target, not a measurement — we have not timed anyone running this routine (ANALYSIS). They exist to stop two failures: skipping the check because it feels long, and sinking twenty minutes into a low-stakes answer.
Step 1 — Name the stakes (10 seconds)
Say what a wrong answer would cost. A wrong film release year costs nothing. A wrong dosage, filing deadline or published price costs a great deal.
The right depth of checking depends on purpose. Mike Caulfield, who developed the SIFT method for checking online information, puts it directly: for a high-level explanation "it's probably good enough to find out whether the publication is reputable", while for deep research "you may want to chase down individual claims" (Caulfield, SIFT (The Four Moves), retrieved 2026-09-17, SOURCED).
For anything medical, legal, financial or about to be published, the two minutes are triage. Run them to find the weakest point, then escalate at step 7 whatever the result.
Step 2 — Check what your question assumed (15 seconds)
Reread your own prompt. "When did the EU ban X?" assumes a ban. "Why is Y faster than Z?" assumes it is.
The measured version comes from legal research. In a study that ran more than 800,000 queries about federal court cases, the authors found that models "often fail to correct a user's incorrect legal assumptions in a contra-factual question setup" (Dahl et al., "Large Legal Fictions", arXiv:2401.01301, retrieved 2026-09-17, SOURCED). That work is on legal questions; the same pattern on general questions is our reading, not a finding (ANALYSIS).
Fix: if your question carried a premise, open a fresh chat and ask the neutral version — "Did the EU ban X, and if so, when?" If the two answers disagree, the first one was shaped by your wording.
Step 3 — Find the claim the answer stands on (20 seconds)
A long answer is a stack of small claims, and it can be mostly right and wrong where it matters. The FActScore evaluation was built on exactly this: it "breaks a generation into a series of atomic facts" because generations "often contain a mixture of supported and unsupported pieces of information, making binary judgments of quality inadequate" (Min et al., arXiv:2305.14251, retrieved 2026-09-17, SOURCED). Pick the one fact that, if false, makes the answer useless to you — usually a number, date, name or rule.
Then pick one more from the end of the answer. The same study found that "across all LMs, the later part of the generation has significantly worse precision" — the authors suggest because earlier facts tend to be better-known and errors propagate as generation continues (SOURCED). The study measured biographies written by 2023-era models; we have not confirmed the effect holds for current models (ANALYSIS).
Step 4 — Rate how risky that claim is (10 seconds)
| Marker | Why it raises the risk |
|---|---|
| Obscure subject — a small company, a niche study, a minor official | A model's ability to answer "relates to how many documents associated with that question were seen during pre-training" (Kandpal et al., retrieved 2026-09-17, SOURCED) |
| A precise figure — 37.4%, $2.3 billion | Numbers are the most checkable-looking and the easiest to misattach to the wrong claim; see why AI makes up statistics |
| A citation, quote or case name | Format is not evidence; see why ChatGPT makes up sources |
| Anything recent or changeable — prices, rules, office holders | May postdate the model's training data unless it searched |
FActScore found the rarity effect held even for the retrieval-augmented system it tested, whose score showed "a relative drop of 50%" at the atomic-fact level as entities became rarer (Min et al., retrieved 2026-09-17, SOURCED). Search access does not remove the obscurity problem. If the claim hits any marker, do not skip step 6.
Step 5 — See whether anything was actually retrieved (15 seconds)
Look for a link next to the claim. If there is none, the claim is recall from training, however current it sounds.
Retrieval is a decision the model makes, not a switch that forces a search. Google's API documentation describes the model analysing the prompt to determine "if a Google Search can improve the answer", and billing "for each search query that the model decides to execute" (Gemini API grounding documentation, retrieved 2026-09-17, SOURCED). OpenAI's equivalent pages returned HTTP 403, so we make no claim about how ChatGPT decides.
If there is a link, open it and find the sentence that says what the answer says. A real page that does not contain the claim is how a cited answer can still be wrong. The full pattern is in how to make ChatGPT cite real sources.
Step 6 — Read laterally for one independent source (40 seconds)
Leave the chat. Open a new tab, search the claim itself — the figure plus the organisation, or the rule plus the jurisdiction — and find one source that did not come from the answer.
This step has direct research behind it. Wineburg and McGrew watched 10 professional fact checkers, 10 PhD historians and 25 Stanford undergraduates evaluate live websites. The fact checkers read laterally: "leaving a site after a quick scan and opening up new browser tabs." They "arrived at more warranted conclusions in a fraction of the time" (Wineburg & McGrew, Teachers College Record 121(11), 2019, author manuscript, retrieved 2026-09-17, SOURCED).
The gap was large. On one task — working out who was behind an advocacy website — checkers found its parent organisation in an average of 51 seconds without prompting; historians averaged 3 minutes 40 seconds and students 5 minutes 18 seconds. (same source, SOURCED).
Two habits from the study transfer directly:
- Click restraint. Before clicking any one result, checkers evaluated the list of search results as a whole, rather than opening the top link.
- Do not read deeper to check. Historians and students tended to read "vertically", staying inside a site to judge it. Rereading the answer more carefully is the same move, and in the study it was the slower, less accurate one.
The study was about websites, not AI answers, and its checker group was ten people. Applying it to chat output is our reading (ANALYSIS) — but an AI answer presents everything in the same confident register, so there is nothing inside it to judge by.
What counts as independent: a primary source (the regulator, the court, the study, the vendor's own pricing page) or reputable coverage that clearly did not copy the same sentence. Five pages repeating one figure with no origin are one source.
Step 7 — Decide (10 seconds)
Three outcomes, and only three:
- Use it. Low stakes, no premise problem, and one independent source agrees with the load-bearing claim.
- Use it with a flag. Low stakes, but you could not confirm it in the budget. Say so wherever you pass it on — "unverified, from an AI answer".
- Escalate. High stakes, an independent source disagrees, or you found nothing. Move to a full trace, or cut the claim.
"It sounded right and I ran out of time" is outcome 2, not outcome 1.
What two minutes cannot catch
The routine is fast because it skips these.
- Misattribution. A real figure from a real source, attached to a claim the source never made. Catching it means reading the source's own sentence in context. That is a trace, not triage.
- Omission. An answer can be accurate in everything it says and wrong by what it leaves out — an exception, a jurisdiction, a newer rule.
- Arithmetic and reasoning. If the answer calculated something, redo the calculation. Checking the inputs does not check the maths.
- Genuinely contested questions. One agreeing source settles a date. It does not settle a question experts disagree on.
What not to spend the two minutes on
- Asking ChatGPT whether it is sure. The legal study above reports that models "cannot always predict, or do not always know, when they are producing legal hallucinations" (Dahl et al., retrieved 2026-09-17, SOURCED).
- An AI detector. It estimates whether text was machine-written, not whether it is true — see do AI detectors actually work.
- Regenerating five times. Samples that disagree are a real warning sign — the idea behind SelfCheckGPT (Manakul et al., arXiv:2303.08896, retrieved 2026-09-17, SOURCED). But samples that agree are not confirmation; one independent source tells you more (ANALYSIS).
To make fewer answers wrong in the first place, see how to stop ChatGPT from making things up.
How This Guide Was Sourced
Written and maintained by the LogicBalls editorial team (logicballs.com). Disclosure: LogicBalls builds AI writing tools, including a chat assistant. Every limitation described here applies to our tools as well, and nothing in the routine requires a product.
AI involvement. AI assisted with drafting, and AI performed the source checks: every quotation and figure above was matched against the linked source on 2026-09-17, and every link was resolved on that date. Editorial responsibility for this page rests with the LogicBalls editorial team.
Sources. Lateral reading: Wineburg and McGrew, read from the authors' manuscript in Stanford's repository, because the journal page returned HTTP 403. SIFT: Mike Caulfield's own description. Atomic facts, rarity and position effects: FActScore (EMNLP 2023). Long-tail knowledge: Kandpal et al. Premise acceptance and self-knowledge: Dahl et al. Sampling-based detection: Manakul et al. Conditional retrieval: Google's Gemini API documentation.
What could not be fetched. OpenAI's help centre, terms of use and product announcements all returned HTTP 403 to automated requests on 2026-09-17. No claim here is sourced to OpenAI material, including how ChatGPT decides when to search.
What is not claimed. The time budgets are a design target — we timed no one. We publish no rate for how often ChatGPT answers are wrong and no measure of how many errors the routine catches. The research cited used specific older models and tasks; results on current models may differ.
No LogicBalls telemetry is used in this guide. Every figure above is external and linked.
Frequently Asked Questions
What is the single fastest check?
Open a new tab and search the load-bearing claim with the name of whoever would have published it. That is lateral reading.
Does it help if ChatGPT gave a source?
Only once you open it and find the sentence that supports the claim.
Should I ask ChatGPT to double-check itself?
Not as your check. Models do not reliably know when they are wrong. A fresh chat with a neutral question is more useful than asking the same chat whether it is sure.
Conclusion
Two minutes will not tell you an answer is true. It will tell you whether you can act on it, and it breaks the habit of trusting an answer because it sounded certain. Name the stakes, check your question, find the claim that matters, and leave the chat to confirm it. When the routine says escalate, the 42-check checklist and the source trace are the next step.
Related reading
- How to Stop ChatGPT From Making Things Up
- How to Make ChatGPT Cite Real Sources (And Check the Ones It Gives You)
- How to Trace an AI Claim Back to Its Original Source
- Free AI Fact-Check Checklist: 42 Checks Before You Publish
- Verified AI Writing: How to Publish AI-Assisted Content You Can Stand Behind