9 Places AI Gets Your Company Facts From (And How to Fix Each)
TL;DR
- Nine sources an assistant can take a company fact from, each with the documented mechanism and the fix - three you edit yourself, four you request from someone else, two you cannot reach.
A wrong answer about your company came from somewhere specific. Nine sources account for almost all of them: your own site, live retrieval, the training corpus, Wikipedia and Wikidata, your Google business data, review and directory listings, news and press releases, your customers' and partners' pages, and your structured data. Three you edit directly. Four you can ask someone else to change. Two you cannot reach at all. Finding the source is the whole job, because the fix is different for each.
This is the inventory, sorted by how much control you have. It is written against Google Search Central, the Gemini API documentation, Anthropic's web search documentation, OpenAI's crawler documentation, Common Crawl's FAQ, Wikipedia's and Wikidata's own policies, Google Business Profile and Merchant Center help, and Trustpilot's published guidelines, all retrieved 2026-09-18. Crawler names and product behaviour change; pin your reading to that date.
Key Takeaways
- Retrieval is the source you cannot force. Google's documentation says the model "analyzes the prompt and determines if a Google Search can improve the answer" (Gemini API, Grounding with Google Search, retrieved 2026-09-18).
- Anthropic names your company as a search trigger. Claude searches when a request involves "Information about specific organizations, people, or products that might have changed" (Anthropic, Web search tool, retrieved 2026-09-18).
- Blocking training crawlers costs you nothing in Search. "Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search" (Google's common crawlers, retrieved 2026-09-18).
- Wikipedia will not take your word for it, and you are "strongly discouraged from editing affected articles directly" (Wikipedia:Conflict of interest, retrieved 2026-09-18).
- Markup is not a channel. Google states "You don't need to create new machine readable files, AI text files, or markup to appear in these features" (Google Search Central, AI features and your website, retrieved 2026-09-18).
The nine, and what you can actually change
| # | Source | Can you change it? | How fast |
|---|---|---|---|
| 1 | Your own website | Yes, directly | As soon as it is recrawled |
| 2 | Live retrieval at answer time | No — you change what it finds, not whether it looks | Not documented by any provider |
| 3 | The training corpus | No, for what is already in it | Never, for a trained model |
| 4 | Wikipedia and Wikidata | By request, with sources | Community-dependent |
| 5 | Google Business Profile and Merchant Center | Yes, directly | After Google's review |
| 6 | Review and directory platforms | Only within their rules | Platform-dependent |
| 7 | News coverage and press releases | By asking the publisher | Publisher-dependent |
| 8 | Your customers' and partners' pages | By asking them | Relationship-dependent |
| 9 | Structured data on your pages | Yes, directly | No AI-feature effect documented |
Rows 1, 5 and 9 are yours today. Rows 4, 6, 7 and 8 are somebody else's page and a polite email. Rows 2 and 3 are properties of the system, and recognising them saves the effort of trying.
1. Your own website
How it reaches a model. Two separate ways. Googlebot crawls it for Search — preferences addressed to that user agent "affect Google Search (including Discover and all Google Search features)" (Google's common crawlers, retrieved 2026-09-18, SOURCED) — and an assistant searching at answer time can retrieve the page directly.
Google also documents a query fan-out technique, in which its AI features issue "multiple related searches across subtopics and data sources" to build one response (AI features and your website, retrieved 2026-09-18, SOURCED). An answer about your company can therefore rest on pages that do not rank for your name at all.
The fix. Put the facts in plain text on a crawlable page, and stop your own pages contradicting each other. The page for the job is a canonical facts page — dated, plain, complete. If prices are the problem, the surface-by-surface sweep is in AI is quoting your old pricing.
2. Live retrieval at answer time
How it reaches a model. The model decides. Google's grounding documentation says it "analyzes the prompt and determines if a Google Search can improve the answer", and that a project "is billed for each search query that the model decides to execute" (Gemini API, Grounding with Google Search, retrieved 2026-09-18, SOURCED).
Anthropic documents the same conditional behaviour and is explicit that your company is a trigger: Claude "determines when to search based on the prompt" and searches when the request depends on "Information about specific organizations, people, or products that might have changed". Steering it is possible; requiring it is not — "For a hard constraint, use max_uses to cap the number of searches for each request", which sets a ceiling, not a floor (Anthropic, Web search tool, retrieved 2026-09-18, SOURCED).
The fix. None for the decision itself. What you control is what a search finds when it happens — rows 1, 5, 7 and 8. When you test, record whether the answer cited anything at all: that single field tells you whether you are looking at retrieval or at row 3. The method is in how to find out what ChatGPT says about your company; why answers go stale even with search available is in why AI gives outdated answers.
3. The training corpus
How it reaches a model. Text crawled before a cutoff and used in training. Common Crawl, one widely used public corpus, documents that "CCBot is an automated crawler, checking first the robots.txt, and if crawling a page is allowed, fetches pages using HTTP GET requests" (Common Crawl FAQ, retrieved 2026-09-18, SOURCED).
Providers publish separate crawler controls for training, and the distinction matters. Google-Extended lets publishers "manage whether content Google crawls from their sites may be used for training future generations of Gemini models", and it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search" (Google's common crawlers, retrieved 2026-09-18, SOURCED). OpenAI documents GPTBot as "used to make our generative AI foundation models more useful and safe" while OAI-SearchBot is "used to surface websites in search results in ChatGPT's search features" (OpenAI, Bots, retrieved 2026-09-18, SOURCED).
The fix. For future crawls, robots.txt directives for the training agents. For what a shipped model already absorbed, nothing: no provider documents a process for correcting a factual claim a model makes about a third-party company, and Common Crawl's FAQ says nothing about removal from completed crawls. Blocking the training agents is a policy decision about your content, not a correction (ANALYSIS) — and by Google's statement above it is one that costs you no Search visibility.
4. Wikipedia and Wikidata
How it reaches a model. Both are crawlable, heavily linked, and retrievable at answer time — Anthropic's own documentation example shows a search result and citation pointing at en.wikipedia.org (Anthropic, Web search tool, retrieved 2026-09-18, SOURCED). Wikidata's structured statements are designed for reuse elsewhere.
Whether you can change it, and how. Yes, but not by editing. "The burden to demonstrate verifiability lies with the editor who adds or restores material, and it is satisfied by providing one inline citation to a reliable source that directly supports the contribution", and self-published material about yourself is usable only where it is "neither unduly self-serving nor an exceptional claim" (Wikipedia:Verifiability, retrieved 2026-09-18, SOURCED).
On your own article you are "strongly discouraged from editing affected articles directly" and should "propose changes on talk pages" using the {{edit COI}} template. Paid editing must be disclosed — you "must disclose who is paying you, on whose behalf the edits are made, and any other relevant affiliation" (Wikipedia:Conflict of interest, retrieved 2026-09-18, SOURCED). Wikidata accepts an item that "refers to an instance of a clearly identifiable conceptual or material entity that can be described using serious and publicly available references" (Wikidata:Notability, retrieved 2026-09-18, SOURCED).
The fix. Publish the fact in a citable place first — your own dated page, or coverage that states it — then request the correction on the talk page with that citation and your affiliation disclosed. A correction with a source is a request an editor can act on; one without is an argument.
5. Google Business Profile and Merchant Center
How it reaches a model. Google lists "Checking that your Merchant Center and Business Profile information is up-to-date" among its best practices for AI features (AI features and your website, retrieved 2026-09-18, SOURCED), which makes this the one place Google names your business data directly.
Whether you can change it. Yes, with a review step: "We review your changes before we update them live on your profile." Anyone else can act too — "If you find something wrong with a Business Profile that you don't own or manage, you can suggest an edit" (Google Business Profile Help, retrieved 2026-09-18, SOURCED). An unclaimed profile is editable by strangers.
For product data, Merchant Center has a rule worth knowing if you generate copy: "All descriptions created using generative AI must be provided using the structured description [structured_description] attribute instead of the description [description] attribute" (Merchant Center, product data specification, retrieved 2026-09-18, SOURCED).
The fix. Claim the profile, correct the fields, and re-read them quarterly. The correction routes Google documents for business data are set out in AI misinformation about your brand.
6. Review and directory platforms
How it reaches a model. They are public, well-linked pages about your company, so they are both crawlable and retrievable. They also tend to state your category and pricing in their own words rather than yours (ANALYSIS).
Whether you can change it. Only inside the platform's rules, and being unhappy is not one of them. Trustpilot's guidelines list the reportable categories — hate speech, threats, obscenity, defamation, personal information, advertising, reviews not based on a genuine experience — then say: "Not liking a star rating or disagreeing with a negative review is not a valid reason for flagging it, and we don't remove reviews just because a business thinks they are unfair or critical" (Trustpilot, Guidelines for businesses, retrieved 2026-09-18, SOURCED).
The fix. Correct the facts on the listing — category, plan names, price, feature list — through the platform's profile route, and leave the opinions alone. A factual error in a directory entry is usually a form; a bad review is not a data problem.
7. News coverage and syndicated press releases
How it reaches a model. News pages are crawled for Search and retrieved at answer time, and a press release is deliberately republished across many domains at once. Where several URLs carry the same text, Google chooses: "If you don't specify a canonical URL, Google will identify which version of the URL is objectively the best version to show to users in Search" (Google Search Central, canonicalization, retrieved 2026-09-18, SOURCED). So one wrong number in one release becomes the same wrong number on twenty domains, and the copy that gets retrieved may not be the one you can edit (ANALYSIS).
The fix. Check the release before it goes out — the only cheap moment. Afterwards ask the publisher for a correction in writing with the primary source attached; most will. Where one will not, publish the corrected, dated version on your own domain so a correct copy is retrievable at all.
8. Your customers' and partners' pages
How it reaches a model. Same mechanism as any third-party page: crawled, indexed, retrievable. Case studies, integration directories, partner listings and conference bios routinely describe what your product does, often with a category or plan name you retired.
The fix. You cannot edit them, but with a customer or partner you have the relationship to make asking work (ANALYSIS). Keep a list of the third-party pages that describe you, ordered by how often they turn up as citations in your own testing, and work down it. The prompts that surface which pages are actually cited are in 40 prompts to check what AI says about your brand.
9. Structured data on your pages
How it reaches a model. Less directly than it is usually sold. Google's position is explicit: "You don't need to create new machine readable files, AI text files, or markup to appear in these features. There's also no special schema.org structured data that you need to add" (AI features and your website, retrieved 2026-09-18, SOURCED).
The fix. Add Organization markup because it is cheap, correct and useful for Search features — and do not buy anything sold as an AI-visibility markup product on the strength of it. Two specifics before you mark up a product: a SoftwareApplication rich result requires a rating or review, and a past priceValidUntil date can stop a merchant listing showing. The rich result Google removed this year is covered in the FAQ schema generator post.
Where to start
- Test first, and record what each answer cited. You cannot fix a source you have not identified, and much of what looks like a model problem is row 1.
- Fix rows 1, 5 and 9 — your site, your Google business data, your markup. Cheapest, fastest, entirely yours.
- Work rows 4, 6, 7 and 8 by request, most-cited page first, with a citation attached to every ask.
- Decide row 3 once, as policy: block the training agents or do not, knowing Google says it costs nothing in Search.
- Stop fighting row 2. Make the correct answer the easiest thing to find, then re-test on a schedule.
Why a wrong answer happens, as distinct from where it came from, is in 7 reasons AI describes your product wrong. What Google documents about mistakes in AI Overviews is in AI Overviews accuracy.
How This Guide Was Sourced
Written and maintained by the LogicBalls editorial team (logicballs.com). Disclosure: LogicBalls builds AI writing tools. Seven of the nine fixes above are editing a page you already own or sending an email. None of them needs our product.
AI involvement. AI assisted with research and drafting. Every quotation was checked against the linked documentation on 2026-09-18, and each crawler name, attribute and policy line was read on the provider's own page rather than in a summary.
Sources. Google Search Central (AI features, canonicalization) and Google's crawler documentation; the Gemini API grounding documentation; Anthropic's web search tool documentation; OpenAI's bots documentation; Common Crawl's FAQ; Wikipedia's Verifiability policy and Conflict of interest guideline; Wikidata's Notability policy; Google Business Profile and Merchant Center help; Trustpilot's guidelines for businesses.
What could not be verified. No provider publishes how long a corrected page takes to change an answer, so no timeline is given. Common Crawl's FAQ does not address removal from completed crawls. Providers' published data descriptions do not enumerate which sites are in a training corpus, so nothing here claims any particular page is in one — row 3 describes the documented mechanism, not your page's presence in it.
What is not claimed. No ranking of the nine by how often each causes an error, and no measurement of how often a fix changes an answer. The ordering is by how much control you have, which is a judgement.
No LogicBalls telemetry is used in this guide. Every figure above is external and linked.
Frequently Asked Questions
Which of the nine causes the most wrong answers?
We have no measurement and neither does anyone we could cite. The cheapest place to look first is your own site, because it is the one source where a contradiction between two of your pages is entirely your doing.
Can I stop my site being used for AI training?
Add robots.txt directives for the training agents — Google-Extended, GPTBot and CCBot among them. That governs future crawls only, and Google states Google-Extended has no effect on Search inclusion or ranking.
Will blocking AI crawlers make my company invisible to assistants?
Blocking training agents and blocking search agents are different decisions. OpenAI documents GPTBot for model training and OAI-SearchBot for its search features; blocking the latter is what removes you from those results.
How do I get a Wikipedia article about us corrected?
Not by editing it. Post an edit request on the article's talk page with a citation to a reliable published source, disclose your affiliation, and if you are paid to do it say who is paying you.
A review site has our pricing wrong. Can I get the review removed?
Two different things. Correct the factual fields through the platform's profile route; the review stays unless it breaches that platform's guidelines, and disagreeing with it is not a breach.