8 Accuracy Questions to Ask Any AI Vendor Before You Pay
TL;DR
- Eight questions for an AI vendor, each with what a real answer sounds like and what a deflection sounds like, grounded in NIST's measurement functions, the EU AI Act's documentation duties and the FTC's substantiation rule.
Ask for the measurement rather than the number, the model version it was measured on, what gets monitored after you sign, the failure modes in writing, whether sources are checked, what happens when the output is wrong, which documents your legal team may read, and which obligations land on you rather than on the vendor. Eight questions. Each one below carries what a real answer sounds like and what a deflection sounds like, because the deflection is the more common response.
This post is written against NIST AI 100-1 (AI RMF 1.0), NIST AI 600-1 (Generative AI Profile, July 2024), Regulation (EU) 2024/1689 as published in the Official Journal, the European Commission's current application timeline for that regulation, the FTC's advertising guidance for small business, and Microsoft's published application card for Copilot, all retrieved 2026-09-18. It is not legal advice.
Key Takeaways
- A demo is not a measurement. NIST asks that performance be "measured qualitatively or quantitatively and demonstrated for conditions similar to deployment setting(s)" (NIST AI 100-1, MEASURE 2.3, retrieved 2026-09-18).
- NIST tells builders not to do what a sales call does. Its generative-AI profile says to "avoid extrapolating GAI system performance or capabilities from narrow, non-systematic, and anecdotal assessments" (NIST AI 600-1, MS-2.5-001, retrieved 2026-09-18).
- The law that forces accuracy metrics into writing is not yet in force for high-risk systems. Those obligations now start 2 December 2027 and 2 August 2028 (European Commission, retrieved 2026-09-18).
- The vendor's marketing claim is already regulated. Advertisers "must have proof to back up express and implied claims that consumers take from an ad" (FTC, retrieved 2026-09-18).
- A vendor doing this well writes its failure modes down. Microsoft's own card says Copilot "may include information in its response that isn't present in its input sources" (Microsoft Learn, retrieved 2026-09-18).
Our AI accuracy rate post promised this list as a separate piece, with the follow-ups for when a vendor deflects. This is it. That post interrogates the number itself in six questions; question 1 below hands you straight back to it, and the other seven cover the ground a purchase decision needs and a benchmark discussion never reaches.
| # | The question | Ask for | The deflection to expect |
|---|---|---|---|
| 1 | Show us the measurement, not the number | Task, test set, metric, method | "Independent testing confirms 99%" |
| 2 | Which model versions was it measured on | Version list, change-notice terms | "We always use the latest models" |
| 3 | What is monitored after we sign | A named metric and a report we see | "We monitor quality internally" |
| 4 | What are the failure modes, in writing | A written limitations page | "Our system doesn't hallucinate" |
| 5 | Are sources shown and verified | Citation behaviour and who checks it | "It's trained on trusted data" |
| 6 | What happens when it is wrong | Logging, escalation, correction, liability | "A human should always review output" |
| 7 | What can our legal team read | Instructions for use, evaluation results | "We're aligned with ISO 42001" |
| 8 | Which obligations land on us | A written split of duties and dates | "The AI Act doesn't apply to this" |
Question 1 — Show us the measurement, not the number
Accuracy is not a property a product has. NIST's framework adopts the ISO/IEC TS 5723:2022 definition — "closeness of results of observations, computations, or estimates to the true values or the values accepted as being true" — so every accuracy figure is the result of one experiment on one task. The framework then sets the bar that figure should clear: performance "measured qualitatively or quantitatively and demonstrated for conditions similar to deployment setting(s)," with measures documented (MEASURE 2.3), and limits on generalisability "beyond the conditions under which the technology was developed" written down (MEASURE 2.5) (NIST AI 100-1, retrieved 2026-09-18, SOURCED).
A real answer names the task, the test set and its size, the metric, and who ran it — and says where the figure stops applying.
A deflection cites "independent testing" without naming the tester, or quotes a percentage with no task attached. The six follow-ups for this one, including contamination and class balance, are in what "99% accurate" actually means.
Question 2 — Which model versions was that measured on, and what happens when you change them?
Most AI products are assemblies. Microsoft's application card points readers to OpenAI's documentation for details of the underlying models it uses, and notes that "Microsoft 365 Copilot is now named Microsoft Copilot" (Microsoft Learn, retrieved 2026-09-18, SOURCED). If the product name moved this year, the model behind it can move next quarter.
A figure measured on one model does not survive a swap, and swaps arrive as improvement announcements rather than change notices (ANALYSIS). NIST puts the remedy in contract language: maintain "well-defined contracts and service level agreements (SLAs) that specify content ownership, usage rights, quality standards, security requirements" (NIST AI 600-1, GV-6.1-004, retrieved 2026-09-18, SOURCED).
A real answer lists the models in use, says which of your workflows touch which, and offers notice before a default changes.
A deflection is "we always route to the best available model" — intent, not something you can hold anyone to. Why model facts rot this fast is covered in Gemini accuracy from Google's own documentation.
Question 3 — What do you monitor after we sign, and what will we actually see?
Pre-sale numbers describe a laboratory. NIST asks that functionality and behaviour "are monitored when in production" (MEASURE 2.4), and its generative-AI profile tells buyers to update vendor due diligence to "address ongoing monitoring, assessments, and alerting, dynamic risk assessments, and real-time reporting tools for monitoring third-party GAI risks" (NIST AI 600-1, GV-6.1-009, retrieved 2026-09-18, SOURCED).
A real answer names one metric it tracks in production, says how often it is computed, and agrees that you see it — in a report, a dashboard, or a review call.
A deflection is "we monitor quality internally." Monitoring you never see cannot inform a renewal decision, and cannot tell you whether the tool got worse on your work specifically (ANALYSIS). Ask the follow-up too: what happens when their own monitoring flags a regression. A vendor that has not thought about it will describe an alert with nobody attached to it.
Question 4 — What are the documented failure modes, and may we have them in writing?
Every generative system has them. NIST's profile defines the central one plainly: confabulation is where systems "generate and confidently present erroneous or false content in response to prompts," including outputs that "contradict previously generated statements in the same context," and notes these are "colloquially also referred to as 'hallucinations' or 'fabrications'" (NIST AI 600-1, retrieved 2026-09-18, SOURCED).
The best answer we can point at is a vendor's own admission. Microsoft's card states that in summarising several sources Copilot "may include information in its response that isn't present in its input sources. In other words, it may produce ungrounded results," and lists the evaluations it runs for "groundedness, relevance, and coherence" (Microsoft Learn, retrieved 2026-09-18, SOURCED).
A real answer is a page you can read, listing limits and high-stakes uses to avoid.
A deflection is "our architecture prevents hallucination." Grounding reduces invention from nothing; it does not stop a correct source being summarised wrongly. Nine reasons teams stay sceptical are in the AI trust gap.
Question 5 — Does it show sources, and who verifies them?
A citation is a claim about a claim. NIST's profile asks builders to "review and verify sources and citations in GAI system outputs during pre-deployment risk measurement and ongoing monitoring activities" (NIST AI 600-1, MS-2.5-003, retrieved 2026-09-18, SOURCED). Note the scope: pre-deployment and ongoing.
A real answer says whether the product cites at all, whether citations are links to retrieved documents or generated text, and what the vendor has done to check they resolve.
A deflection is "it's trained on trusted sources." Training provenance says nothing about whether this specific answer's references exist. Give your evaluators the two-minute check in how to check if ChatGPT is right and the longer gate in the 42-check fact-check checklist, then run them on the vendor's own demo output during the trial.
Question 6 — What happens when it is wrong, and who carries the cost?
Ask four things: is the output logged, how does a user escalate, how is a wrong answer corrected once it has reached a customer, and what does the contract say about the consequences. NIST's guidance points at the paperwork: contracts and SLAs specifying "quality standards," plus after-action reviews of "GAI system incident response and incident disclosures, to identify gaps" (NIST AI 600-1, GV-6.1-004 and GV-1.5-002, retrieved 2026-09-18, SOURCED).
A real answer may well be "the liability sits with you." That is defensible and common. It is also a price, and it should be priced before signature rather than discovered afterwards (ANALYSIS). What one uncorrected error actually costs, across eight documented cases, is in what one wrong fact actually costs.
A deflection is "a human should always review the output," offered as a control rather than as a cost. Ask who that human is, how long the review takes, and whether the time saving survives it.
Question 7 — Which documents may our legal team read before we sign?
The EU's regulation is useful here even where it does not yet bind, because it defines what adequate documentation looks like. For high-risk systems, instructions for use must state "the level of accuracy, including its metrics, robustness and cybersecurity referred to in Article 15 against which the high-risk AI system has been tested and validated and which can be expected" (Article 13), and "the levels of accuracy and the relevant accuracy metrics of high-risk AI systems shall be declared in the accompanying instructions of use" (Article 15) (Regulation (EU) 2024/1689, retrieved 2026-09-18, SOURCED).
For general-purpose models, providers must keep technical documentation including "the results of its evaluation," and must make information available to the providers who build on top of them (Article 53) (same source, SOURCED).
A real answer hands over a document: a card, a limitations page, an evaluation summary, instructions for use.
A deflection is a certification claim with nothing readable behind it. ISO/IEC 42001 is a paid standard; we have not read it and make no claim about its contents, and neither can a buyer who has only been told a vendor "aligns" with it. For the framework landscape itself, see 8 AI trustworthiness frameworks.
Question 8 — Which obligations land on us, and from when?
Buyers assume vendor duties are their protection. Often the duty is theirs. Under the EU regulation, providers of systems generating synthetic text "shall ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated" (Article 50) (Regulation (EU) 2024/1689, retrieved 2026-09-18, SOURCED) — and if you publish that text, your disclosure practice is your own problem.
The timeline has moved, which is exactly why this question needs a date attached. The Commission states the regulation "entered into force on 1 August 2024 and became applicable on 2 August 2026," that governance and general-purpose-model obligations applied from 2 August 2025, and that high-risk rules were pushed back by the amendment package known as the AI Omnibus: Annex III use cases to 2 December 2027 and Annex I embedded products to 2 August 2028 (European Commission, retrieved 2026-09-18, SOURCED).
Meanwhile the vendor's own claim is regulated today. The FTC states that advertisers "must have proof to back up express and implied claims," and that a money-back guarantee "is not a substitute for substantiation" (FTC, retrieved 2026-09-18, SOURCED).
A real answer splits the duties in writing. A deflection is "that regulation doesn't apply to us," with no reasoning about which category the system falls into.
What we can and cannot answer about our own product
We are a vendor, so these questions point at us too.
On question 1 our published answer is a negative: we quote no accuracy figure, because we have not run a benchmark that would survive the test-set question (our accuracy-rate post says so, and that has not changed). On question 2, our own pricing page sells model access by tier — 3 models on the free plan, 10, 15 and 31 on the paid ones (LogicBalls pricing, retrieved 2026-09-18, SOURCED) — so the model behind a given tool depends on the plan and can change.
What we publish instead of a number is our own error record: the audit of our own blog, the 14 error types in our own AI-assisted posts, and the checks that catch them. Ask us questions 3 through 8 in writing and hold the answers to the standard above.
How This Guide Was Sourced
Written and maintained by the LogicBalls editorial team (logicballs.com). Disclosure: LogicBalls builds AI writing tools. Every question here can be turned on us, and the section above does exactly that.
AI involvement. AI assisted the research and drafting. Every quotation was matched against the source document text on 2026-09-18: both NIST publications were read as PDFs, the regulation in its Official Journal text, and the Commission, FTC, Microsoft and pricing pages fetched and searched directly.
Sources. NIST AI 100-1 (AI RMF 1.0) for the accuracy definition and the MEASURE subcategories; NIST AI 600-1 (Generative AI Profile, July 2024) for the confabulation definition and the MS and GV suggested actions; Regulation (EU) 2024/1689 for Articles 13, 15, 50 and 53; the European Commission's regulatory-framework page for the application timeline; the FTC's advertising guidance for substantiation; Microsoft's Copilot application card as an example of a vendor documenting its own limits.
What could not be verified. ISO/IEC 42001 sits behind a paywall and returned HTTP 403 to our request; we did not read it and make no claim about its contents. The AI Omnibus amending text itself was not read — the dates above are the Commission's own description of it, and a consolidated version of the regulation was not retrievable at the URL we tried. No provider's help centre was used as a source.
What is not claimed. No ranking of these eight questions by how often they get answered, no measurement of how many vendors deflect, and no legal advice about which category any system falls into.
No LogicBalls telemetry is used in this guide. Every figure above is external and linked.
Frequently Asked Questions
What is the single most useful question if I only get one?
Question 1, in its narrow form: what task was measured, on what data, with what metric. A vendor who can answer that in one sentence has done the work. One who cannot has a number rather than a measurement.
Is a vendor legally required to tell me its accuracy rate?
Not generally. The EU regulation requires accuracy metrics in the instructions for use of high-risk systems, and those obligations now begin on 2 December 2027 and 2 August 2028 depending on category. Outside that, asking is your own diligence.
What if the vendor answers everything well but has no benchmark?
That can be the honest position, and it is ours. Published limits, production monitoring and a correction process are worth more than an unreproducible percentage.
Where do general software buying questions fit?
They still apply — fit, adoption, support, exit. Those are covered in our B2B software buying questions guide. The eight above are the accuracy-specific layer on top.
Related reading
- AI Accuracy Rate: What "99% Accurate" Actually Means When You Buy AI
- The AI Trust Gap: 9 Reasons Teams Still Do Not Rely on AI Output
- What One Wrong Fact Actually Costs: 8 Real Consequences
- Gemini Accuracy: What Google's Own Documentation Says
- 8 AI Trustworthiness Frameworks Every User Should Know About