AI Buying

How to Read an AI Vendor's Claims: Three Questions Before You Buy

A polished vendor deck can compress your thinking time. Three questions give you room to examine the claim — and the same discipline that lets you say a calm no is what frees you to say a fast yes.
By Bruno Oliveira • 15 min read • September 15, 2026

The Triage, in Numbers

3 Questions That Triage an AI Capability ClaimHas it shipped · has it been reproduced · does it change the job
8pp Largest Accuracy Drop When a Benchmark Was Mirrored by Fresh ProblemsGSM1k study, Scale AI, 2024 (arXiv 2405.00332) — many frontier models tested showed minimal signs of overfitting
2 Columns a Vendor Deck Splits Into When You AskAvailable on our account now · roadmap
80% time savings on AI-assisted tasksAnthropic
6% of orgs are AI high performersMcKinsey

Somewhere in your inbox this week there is probably a vendor deck. It is beautifully produced. It cites a benchmark, shows a demonstration that borders on magic, and implies — politely, confidently — that the way your firm currently works is about to become obsolete. The demo call is already in the diary.

This is the page for that moment. It sets out the reading discipline I apply before any claim is allowed to change what a firm does, written for owners, finance leaders and the sponsors who carry an AI adoption through a larger organisation.

It is deliberately calm, deliberately practical, and written to be forwarded to whoever will sit beside you on the call.

One premise underneath it: the lasting value in AI is rarely the single tool in the deck. It is the operating system a firm builds around the models — the context they hold, the routines they run, the judgement about what runs where. Tools come and go. The discipline for choosing them compounds.

The argument in 60 seconds

  • A vendor deck is built to persuade, and a polished one can compress your thinking time. The polish, the benchmark, the quiet suggestion that your current way of working is already obsolete. The discipline below is how you decline the rush without declining the technology.
  • The posture that works is sceptical-curious. Cynicism misses real progress; credulity chases vapour. The middle path treats every claim as a hypothesis, then moves quickly and gladly when the hypothesis survives testing.
  • 3 questions do the triage. Has it shipped, or has it only been announced? Has anyone independent reproduced it? Does it change the actual job, or just move a leaderboard?
  • A vendor's own numbers are a starting point, not proof. A polished demonstration is not a shipped product, and a self-reported score is not an independent one — treat both as hypotheses until they hold up somewhere you trust.
  • Clearing the 3 questions earns a review, not a signature. What survives goes on to the security, confidentiality and procurement work that decides adoption. The triage exists so that review is spent on claims worth the cost of it.
📋

Get the AI Opportunity Spotter™

The 1-page framework I use to identify highest-impact AI use cases in any business

In this article:

  • Generating table of contents...

Why Do AI Vendor Claims Need a Reading Discipline?

Because the announcement cycle moves every week, and much of it is noise dressed as signal. A reading discipline separates the claim worth acting on from the many that are not, so a firm keeps building on solid ground — and keeps the attention to move decisively when something real arrives.

The trap is two-sided. The cynic, burned once by an over-promised tool, dismisses everything, and can end up years behind people who were no smarter, only more open.

The true believer does the opposite: re-platforms the firm on every keynote, and compounds little, because nothing stays still long enough to deepen.

The posture that works will be familiar from investing and from science, and it has two halves that are easy to separate and hard to hold together. Sceptical — treat every claim as a hypothesis until it is tested. Curious — be genuinely delighted, and quick to move, when the test is passed.

It is, deliberately, the posture of someone who is hard to sell to and easy to genuinely impress.

What follows are the 3 questions that turn the posture into practice. Ask them in order, before the demo call if you can.

Question 1 — Has It Shipped, or Has It Only Been Announced?

A claim has shipped when the capability is available to your firm today — on the account and terms you would actually use, without a waitlist and without a staged environment. Someone whose judgement you trust using it on real work is good corroboration. It is not a substitute for availability to you.

A keynote demonstration is staged by design. A research preview may sit behind a waitlist or an access gate. "Rolling out over the coming weeks" is not "available on your account this morning".

The gap between a beautiful demonstration and a generally available product is where a great deal of disappointment lives. A demo is built to show the best case; your work is whatever it is on an ordinary Tuesday.

The version to ask in the room is plain: which of the capabilities in this deck can we use this afternoon, on our own account, on our own work — and which are roadmap? Ask it kindly, and watch the deck reorganise itself into two columns: now, and later.

The now column is the product you are actually evaluating. The later column is a direction of travel — worth hearing, not worth deciding on.

Until a capability sits in the now column, treat the claim as a note in your diary rather than a line in your plans. That is not a rejection. Record it, set a review date, and carry on with what already works. Announcements age quickly. Set the review date, and let the product prove it is still there when the date arrives.

Question 2 — Has Anyone Independent Reproduced It?

Vendors report their own numbers, on their own chosen tests, framed in their own best light. That is ordinary marketing. It is not evidence of bad faith, and it is not proof either. A claim earns trust when it is tested outside the vendor.

There is a documented version of this problem in machine-learning research, and the honest reading of it cuts both ways. In a 2024 study, a team at Scale AI commissioned GSM1k — 1,205 newly written grade-school mathematics problems, designed to mirror the widely used GSM8k benchmark in style and difficulty — and evaluated the leading open and closed models on both.

They observed accuracy drops of up to 8 percentage points, and found several families of models showing evidence of systematic overfitting across almost all model sizes. The same paper is equally clear about the other half: many models, especially those on the frontier at the time, showed minimal signs of overfitting, and all of them broadly demonstrated generalisation to novel problems guaranteed not to be in their training data.

Note carefully what that study does and does not show. The researchers ran both benchmarks themselves, under their own conditions, so it is not a measurement of vendors overstating their published scores. It shows something a buyer can use: evidence that a strong result on a public benchmark can overstate performance on fresh problems of the same kind — for some models, and not for others.

Which is precisely why the question is worth asking, and precisely why the answer is sometimes reassuring. A vendor grading its own homework is not the same as proof, and it is not the same as a lie either.

For a buyer, independence takes 3 practical forms. An independent evaluation the vendor did not commission. A practitioner in your field reporting results on real tasks. And the one that tests the claim on the work you actually care about: a controlled trial on work like your own, with a pass mark agreed before it starts.

Note the sequencing trap in that third form. Running a trial on real work is itself a decision about data, access and confidentiality, so it needs its own approval before it starts, separate from the approval to adopt.

There is also a question a deck cannot answer but a vendor can: may we speak to a customer in our sector who uses this capability for the workflow you are showing us? Not a logo on a slide — a conversation.

It is a reasonable thing to ask, and a supplier's willingness to arrange it tells you something either way. The reference will have been chosen by the vendor, so treat what you hear as testimony rather than as independent evidence — and use the half hour on results, conditions and limits together.

💡 The Sentence That Decides the Meeting

A vendor grading its own homework is not the same as proof — and it is not the same as a lie either.

Both errors are expensive. Treating every self-reported figure as dishonest can cost a firm real progress it should have moved on early. Accepting one without checking can leave a buying decision resting on a number nobody outside the vendor has ever reproduced.

The discipline is not suspicion. It is the ordinary professional habit of asking who ran the test, on whose data, and whether anyone else has since.

Question 3 — Does It Change the Actual Job, or Just Move a Leaderboard?

The most impressive-sounding claim in a deck is often not the most relevant one. A record on a mathematics-contest benchmark may change nothing about how your firm drafts an advice note or prepares for a difficult client meeting. Before a claim earns your attention, ask what it changes about the specific work in front of you.

A capability can top a coding leaderboard, or clear an academic benchmark that once looked untouchable, and still do nothing for how your firm turns a messy set of documents into a clear recommendation. The claim can be entirely true and entirely beside the point.

So make relevance concrete before the call ends: name the workflow that changes, and name the person who owns it. What stops, what gets shorter, what gets better?

If nobody at the table can answer, you are in a technology briefing rather than a buying decision. Both are useful; only one of them should end with a signature.

Relevance is the filter that turns a flood of announcements into the handful of claims that deserve real time. If the honest answer today is "nothing yet", admire the capability as you would a fast car you have no road for, and move on, unembarrassed. Admiration is free; adoption is not.

This is also the question that depends most on work you should already have done. A firm that has mapped where its highest-value AI work actually lives answers question 3 quickly, because it already knows which workflows matter and what each one demands. A firm that has not has less to test the deck against.

The same discipline that lets you say a calm no is what frees you to say a fast yes.

How Do You Run the Three Questions on a Live Vendor Deck?

In order, with the deck open, before the demo call if you can. Each question is a gate: a claim that fails stops there, and gets a note and a review date.

Here is the working checklist — the version to forward to whoever joins the call.

Question 1: Has It Shipped?

  • Can we use this today, on our own account and terms, on our own work?
  • Which claims in the deck are available to us now, on the account and terms we would use, and which are roadmap?
  • If access is gated or restricted in a way that does not match how we would use it: record it, set a review date, and revisit.

Question 2: Has Anyone Independent Reproduced It?

  • Who outside the vendor has tested this claim, and where is the result?
  • May we speak to a customer in our sector who uses it for this exact workflow — understanding that a vendor-chosen reference is testimony, not independent evidence?
  • What would our own controlled trial look like, what data access would it need, and what pass mark, agreed in advance, would justify going further?

Question 3: Does It Change the Actual Job?

  • Which specific workflow in our firm changes, and who owns it?
  • What stops, shortens or improves if the claim is true on our work?
  • If the honest answer is "nothing yet": admire it, note it, move on.

A claim that passes all 3 has not earned adoption. It has earned the next stage — the security, confidentiality and procurement review that decides adoption, with the firm's own rules written down before the tool arrives rather than assembled afterwards.

The 3 questions decide whether that review is worth commissioning at all. The gate is deliberately steep, which is exactly what frees the diary for the claims that clear it.

Working on this inside your firm?

GustoMind works with expert-led firms on exactly this — from a readiness diagnostic to a full AI operating model. No pitch, just a conversation about where you are.

When Does the Discipline Let You Move Fast?

When a claim is available, independently supported and relevant, the same discipline that filtered out everything else lets you move without hesitation into the review that decides it. Attention not spent on the claims that did not matter is attention still available for the one that does.

The failure mode these questions exist to catch is a specific one, and it is worth naming because it can be an expensive one. Picture the situation rather than any particular firm.

A workflow is finally working, after a long stretch of unglamorous effort. Then an announcement lands that appears to make the whole thing redundant, and the pull is to stop, tear it up, and wait for the thing in the deck.

Run the 3 questions before you do. A capability sitting behind a waitlist, resting on the vendor's own figures, and aimed at a problem your firm does not have is not a reason to abandon something that already works. It is a note in the diary with a review date on it.

That is the deeper reason the discipline pays. It protects the thing you have already built — and what a firm builds around the models is the asset that actually compounds, while the tools inside it change on a schedule you only partly control.

Scepticism that hardens into blanket refusal is simply the other failure mode wearing a more serious face. The firms that win with AI are not the ones that adopt everything earliest, and they are not the last hold-outs. They are the ones with a discipline for telling the difference — which is what lets them be early to the things that turn out to be real.

The frontier genuinely is moving quickly, and that is good news. A reading discipline is not how you resist the frontier. It is how you ride it without being thrown.

If a deck has just landed and you would like a sharper version of these 3 questions for your firm's own buying decisions, see how the engagements work and send a note describing the claim you are trying to read. And if this way of reading the frontier is useful, the thinking continues in The AI Operating System — the fortnightly LinkedIn newsletter this article grew from.

Frequently Asked Questions

Does a pilot count as independent evidence?

A pilot counts when it is a controlled trial: tasks you chose, a baseline you measured, and a pass mark agreed before it starts.

A pilot that is really an extended demonstration — vendor-chosen tasks, vendor-run sessions, no baseline — tells you how the tool performs under conditions the vendor selected, which is not the question you asked. Design the pilot around work your firm already does, agree what data it may touch before it starts, and decide in advance what result would justify going further.

Are benchmark scores worth anything at all?

Yes — as a signal of direction and pace, not as proof of fitness for your work. A benchmark measures performance on its own chosen tasks, and your work may make different demands.

When one widely used benchmark was mirrored by a fresh set of problems in 2024, measured accuracy drops reached 8 percentage points and several model families showed systematic overfitting, while many of the frontier models tested showed minimal signs of it. Both halves matter, and neither tells you whether the capability helps with the work your firm actually does. The score that answers that is the one from your own quiet trial on a task you already run.

When is early adoption the right call?

When a capability is available to your firm, has held up somewhere you trust, touches a bottleneck you actually have, and has cleared the review that adoption requires. Being early to something real can be a genuine advantage. The discipline exists to make that yes safe, not to prevent it.

What if a competitor adopts before we do?

Watch what changes in their client work, not what appears in their marketing. A competitor adopting an announced-but-unshipped capability has volunteered to run the experiment on your behalf. Set a review date, and let the evidence arrive.

Waiting has a cost. So does rebuilding around a claim that did not hold. The review date is where you weigh the two.

Who should run the three questions?

The person who will own the workflow if the claim turns out to be true should lead, alongside whoever signs. The questions are deliberately non-technical, because the decision they inform is not principally a technical one: it is a decision about which piece of the firm's work changes, and at what cost.

A specialist can confirm that a capability does what the deck says, and the people who do the work daily will see things the owner does not. The owner is the one who can weigh whether it matters.

Dr Bruno Oliveira — PhD · Associate Professor, University of Bath. Founder of GustoMind.ai. Builds and installs AI operating systems for expert-led firms, running the same system daily in his own work.

✅ The First Cut, for Thursday Afternoon

This is the screen, not the whole framework — an availability-and-relevance cut that decides what is worth taking to the call, which is where question 2 belongs. It takes one sitting.

Take the deck already in your inbox and split it into two columns. Now, and later. Nothing else — no scoring matrix, no weighted criteria, no procurement template.

Whatever survives that cut is the only part you were ever evaluating. Then ask the question that costs a sentence: which workflow in this firm changes if all of this is true?

If the answer is a real workflow with a real owner, book the call and bring question 2 to it. If the answer will not come, you have learned something worth knowing before you spent an hour on it.