When Does the Discipline Let You Move Fast?
When a claim is available, independently supported and relevant, the same discipline that filtered out everything else lets you move without hesitation into the review that decides it. Attention not spent on the claims that did not matter is attention still available for the one that does.
The failure mode these questions exist to catch is a specific one, and it is worth naming because it can be an expensive one. Picture the situation rather than any particular firm.
A workflow is finally working, after a long stretch of unglamorous effort. Then an announcement lands that appears to make the whole thing redundant, and the pull is to stop, tear it up, and wait for the thing in the deck.
Run the 3 questions before you do. A capability sitting behind a waitlist, resting on the vendor's own figures, and aimed at a problem your firm does not have is not a reason to abandon something that already works. It is a note in the diary with a review date on it.
That is the deeper reason the discipline pays. It protects the thing you have already built — and what a firm builds around the models is the asset that actually compounds, while the tools inside it change on a schedule you only partly control.
Scepticism that hardens into blanket refusal is simply the other failure mode wearing a more serious face. The firms that win with AI are not the ones that adopt everything earliest, and they are not the last hold-outs. They are the ones with a discipline for telling the difference — which is what lets them be early to the things that turn out to be real.
The frontier genuinely is moving quickly, and that is good news. A reading discipline is not how you resist the frontier. It is how you ride it without being thrown.
If a deck has just landed and you would like a sharper version of these 3 questions for your firm's own buying decisions, see how the engagements work and send a note describing the claim you are trying to read. And if this way of reading the frontier is useful, the thinking continues in The AI Operating System — the fortnightly LinkedIn newsletter this article grew from.
Frequently Asked Questions
Does a pilot count as independent evidence?
A pilot counts when it is a controlled trial: tasks you chose, a baseline you measured, and a pass mark agreed before it starts.
A pilot that is really an extended demonstration — vendor-chosen tasks, vendor-run sessions, no baseline — tells you how the tool performs under conditions the vendor selected, which is not the question you asked. Design the pilot around work your firm already does, agree what data it may touch before it starts, and decide in advance what result would justify going further.
Are benchmark scores worth anything at all?
Yes — as a signal of direction and pace, not as proof of fitness for your work. A benchmark measures performance on its own chosen tasks, and your work may make different demands.
When one widely used benchmark was mirrored by a fresh set of problems in 2024, measured accuracy drops reached 8 percentage points and several model families showed systematic overfitting, while many of the frontier models tested showed minimal signs of it. Both halves matter, and neither tells you whether the capability helps with the work your firm actually does. The score that answers that is the one from your own quiet trial on a task you already run.
When is early adoption the right call?
When a capability is available to your firm, has held up somewhere you trust, touches a bottleneck you actually have, and has cleared the review that adoption requires. Being early to something real can be a genuine advantage. The discipline exists to make that yes safe, not to prevent it.
What if a competitor adopts before we do?
Watch what changes in their client work, not what appears in their marketing. A competitor adopting an announced-but-unshipped capability has volunteered to run the experiment on your behalf. Set a review date, and let the evidence arrive.
Waiting has a cost. So does rebuilding around a claim that did not hold. The review date is where you weigh the two.
Who should run the three questions?
The person who will own the workflow if the claim turns out to be true should lead, alongside whoever signs. The questions are deliberately non-technical, because the decision they inform is not principally a technical one: it is a decision about which piece of the firm's work changes, and at what cost.
A specialist can confirm that a capability does what the deck says, and the people who do the work daily will see things the owner does not. The owner is the one who can weigh whether it matters.
Dr Bruno Oliveira — PhD · Associate Professor, University of Bath. Founder of GustoMind.ai. Builds and installs AI operating systems for expert-led firms, running the same system daily in his own work.