How to Find Out What ChatGPT Says About Your Company

brand reputation AI search monitoring
Ankit Agarwal
Ankit Agarwal

Marketing Head

 
September 14, 2026
8 min read
How to Find Out What ChatGPT Says About Your Company

Ask in a clean session, ask the questions a buyer would ask rather than the question you want answered, record whether the answer cited anything, and repeat monthly. That last part is what makes it useful: a single answer on a single day tells you almost nothing, because these systems give different answers to the same question depending on whether they searched, what they retrieved, and what the session already knew about you.

This is the method we use, written for anyone who has to report on it. It works the same way across assistants. Nothing here is sourced to OpenAI's own documentation — their help pages returned HTTP 403 to our automated requests on 2026-09-06, so we do not quote them and do not describe named product features on their authority.

Key Takeaways

  • One answer is not data. Run the same questions monthly, record the date, and compare.
  • Ask in the cleanest session available, logged out or memory-off, so you are testing the model and not your own history with it.
  • Whether the answer cited a source is the most important thing you record. Cited means retrieval, which you can influence. Uncited likely means training data, which you cannot.
  • Search is conditional, not automatic. Google's own API documentation bills per search "the model decides to execute" — so some answers about you never touch a live page.
  • Ask buyer questions, not vanity questions. "What does this company do" is worth more than "is this company good".

Step 1 — Set up a session that tests the model, not your history

Assistants adapt to context. If you have spent a year discussing your own company in an account, that account is the worst possible place to test what a stranger sees.

Use the cleanest session the assistant offers: logged out where that is possible, or a fresh session with memory and personalisation turned off. Then hold these constant so the results are comparable month to month:

  • The account state — same logged-out or memory-off condition every time.
  • The model — record which one. Answers differ across models from the same provider.
  • Whether web search was available, and whether it was on or off.
  • The date and time.

Write those four things down before the first question. Without them you have anecdotes.

Step 2 — Ask the questions a buyer would ask

Twelve questions, in this order. The order matters: the open ones first, before you introduce any facts of your own that the model can echo back.

  1. What does [company] do?
  2. Who is [company] for?
  3. What does [company] cost?
  4. What are [company]'s pricing tiers?
  5. Is there a free plan for [company]?
  6. What are the limitations or downsides of [company]?
  7. Who founded [company], and when?
  8. Where is [company] based?
  9. What are the main alternatives to [company]?
  10. Is [company] legitimate and safe to use?
  11. What do reviews say about [company]?
  12. What integrations does [company] support?

Ask each one on its own, in a fresh session where practical. Asking all twelve in one conversation lets the model's earlier answers contaminate the later ones — and question 1's answer will shape questions 3 and 6 if they share a thread.

Do not correct it mid-session. The moment you say "actually our pricing is X", you have given the model the fact, and everything after is a test of its short-term memory rather than of what it knows.

Step 3 — Record four things per answer

For every answer, log:

field why it matters
The answer, verbatim Paraphrasing loses the specific wrong number, which is the thing you need
Did it cite anything? This determines whether the error is reachable at all
What it cited The page to go fix, or get outranked
Confidence language "I believe", "as of my last update", or flat assertion — the last is the dangerous one

Verbatim matters more than it sounds. "It got our pricing wrong" is not actionable. "It said our Pro plan is $19/month when it is $39" tells you which page to look for.

Step 4 — Sort what you find into two piles

This is the step that decides whether you have work to do or not.

Cited answers are retrieval. The model searched, read pages, and summarised them. Google's Gemini API documentation describes this workflow directly and bills per search query "that the model decides to execute" (Gemini API grounding documentation, retrieved 2026-09-06, SOURCED) — which confirms both that retrieval happens and that it is conditional rather than automatic. If a wrong answer cites a page, you have a target: get that page corrected, or publish something better and more current that outranks it.

Uncited answers are probably training data. No provider publishes the sources behind any individual company's facts, and there is no documented route to request a correction. Anthropic's own support documentation offers a feedback control and a general support address, describes no formal process for correcting facts about specific companies, and tells users not to rely on the assistant "as a singular source of truth" (Anthropic support documentation, retrieved 2026-09-06, SOURCED).

So for the uncited pile the honest answer is: you cannot fix it directly. What you can do is make the correct version so widely and clearly retrievable that future retrieval finds it. That is slow, and it is the only lever there is. We go through the full set of what is and is not documented in AI misinformation about your brand.

Step 5 — Repeat, and watch the direction of travel

Monthly is a reasonable cadence for most companies, plus an immediate re-run after anything material changes: pricing, positioning, a leadership change, a rebrand, a new product line.

What you are watching for over time is not a single wrong fact. It is:

  • Stale facts persisting long after your site changed. Old pricing is the most common by a wide margin.
  • A wrong source becoming the consistent citation. One outdated roundup that keeps getting retrieved is worth real effort to displace.
  • Category drift — the model describing you as something adjacent to what you are. This tends to come from your own site being inconsistent about it.

Three months of records makes all three visible. One month makes none of them visible.

The mistakes that make the exercise useless

  • Asking leading questions. "Why is [company] the best tool for X" will get you an answer built from your own framing, not from what the model holds.
  • Testing in your own logged-in account. You are the least representative user of the model's knowledge about you.
  • Recording a summary instead of the text. The specific wrong number is the deliverable.
  • Running it once before a board meeting. A single answer is noise. The trend is the finding.
  • Treating a cited answer as verified. Retrieval can find and repeat a wrong page with total confidence. Open the citation.

How This Guide Was Sourced

Written and maintained by the LogicBalls editorial team (logicballs.com). Disclosure: LogicBalls builds AI writing tools. We have an obvious interest in people taking AI accuracy seriously. The method above needs no product to run — it is a spreadsheet and an hour a month.

Sources. Retrieval behaviour and its conditional nature: Google's Gemini API grounding documentation. Provider feedback mechanisms and the absence of an entity-fact correction process: Anthropic's own support documentation. Both retrieved 2026-09-06 and linked inline.

What could not be fetched. OpenAI's help documentation returned HTTP 403 to automated requests on 2026-09-06. No claim here is sourced to OpenAI material, and we deliberately do not name or describe specific ChatGPT session, memory or search features on their authority. The setup guidance above is written in terms of the general principle — test in the cleanest session available — rather than a named feature, for that reason.

What is not claimed. We publish no measurement of how often assistants get company facts wrong. The 12 questions are ours, chosen to match what a buyer asks, and are not validated against any published research.

No LogicBalls telemetry is used in this guide.

Frequently Asked Questions

How often should I check?

Monthly for most companies, and immediately after a pricing, positioning or leadership change. Old facts persist in these systems long after your website changes.

Should I check every assistant?

Check the ones your buyers use. The method is identical across them, and the answers differ, so run them separately rather than assuming one represents the others.

The answer was right this month and wrong last month. Which is true?

Both were true answers from the system on those days. That variance is the finding — record it rather than resolving it. It usually means the answer depends on whether the model searched.

Can I automate this?

You can automate the asking and the logging. You cannot automate the judgement of whether a claim is wrong, and that is most of the work. Start manual for the first two or three cycles so you know what you are looking at.

It said something defamatory about us. What now?

Record it verbatim with the date, the model and the session conditions, because that record is the only evidence that exists. Then take legal advice — the position on third-party model output about a company is unsettled, and we found no decided case on it.

Related reading

Ankit Agarwal
Ankit Agarwal

Marketing Head

 

Ankit Agarwal is a growth and content strategy professional focused on building scalable content and distribution frameworks for AI productivity tools. He works on simplifying how marketers, creators, and small teams discover and use AI-powered solutions across writing, marketing, social media, and business workflows. His expertise lies in improving organic reach, discoverability, and adoption of multi-tool AI platforms through practical, search-driven content strategies.

Related Articles

Why Does ChatGPT Make Up Sources? 6 Causes and Fixes
AI hallucination

Why Does ChatGPT Make Up Sources? 6 Causes and Fixes

A study of 636 generated citations found 55% fabricated by GPT-3.5 and 18% by GPT-4. Here are the six mechanisms behind that, and what actually reduces each one.

By Ankit Agarwal September 13, 2026 9 min read
common.read_full_article
Why AI Makes Up Statistics (And How to Get Real Ones Instead)
AI accuracy

Why AI Makes Up Statistics (And How to Get Real Ones Instead)

Three ways an AI statistic goes wrong, only one of which is invention. The dangerous one is a real number attached to the wrong claim — and it survives every check except opening the source.

By Ankit Agarwal September 13, 2026 8 min read
common.read_full_article
Gemini Accuracy: What Google's Own Documentation Says
AI accuracy

Gemini Accuracy: What Google's Own Documentation Says

Google publishes factuality numbers for Gemini, and its own benchmark suite puts every model tested below 70%. Here is what those figures measure, and where an independent benchmark disagrees.

By Ankit Agarwal September 12, 2026 9 min read
common.read_full_article
Free AI Fact-Check Checklist: 42 Checks Before You Publish
fact checking

Free AI Fact-Check Checklist: 42 Checks Before You Publish

The 42-item checklist we run on every AI-assisted draft, grouped into nine gates in the order that catches the most for the least work. Copy it, no email required.

By Ankit Agarwal September 12, 2026 10 min read
common.read_full_article