What Is the Difference Between Validated Capacity and Realised Recovery?
Validated capacity is the credible saving available from workflows that are built, tested and ready but not yet habitual. Realised recovery is time that has actually left the week. Both are worth reading, and reading only one of them is how a working installation can get written off.
Capacity can be present well before recovery shows up, because a workflow can be ready long before it is habitual.
Read the two side by side, on the same set of workflows, and the distance between them becomes a question rather than a verdict. Habit, documentation and coaching are the first places I would look, before looking at the technology. Measure only realised recovery at day 30 and a working system can look like a failed one.
There is a second reason to hold both numbers. A gap that closes slowly is worth investigating as a training and habit question. A gap that does not close at all, over months, is worth investigating as a design question — whether the workflows chosen were the workflows that mattered. Both are hypotheses to test rather than diagnoses the numbers hand you, and neither question is even available from a single blended figure.
The Verdict
The first 30 days are not where a return is proved. They are where the ability to prove one is either built or lost.
Spend them buying a baseline and an honest adoption signal, and the conversation three months later has something to stand on. Skip them, and that conversation compares against whatever records happen to exist and, where they do not, a vague memory of how busy things felt before — which is not evidence, and cannot be argued with either.
Do that, and when a number does arrive you will at least be able to say what it was measured against.
If your firm is about to start measuring an AI installation and you would rather not spend the first month generating a figure you cannot defend, see how the engagements work and send me a note — happy to share the baseline sheet I work from. And if this way of thinking about measurement is useful, it continues fortnightly in The AI Operating System, the LinkedIn newsletter this article grew from.
Frequently Asked Questions
How soon can an AI installation honestly show a return?
It depends entirely on which layer you mean. Delivery can be shown within days, because it is observable. Adoption becomes readable within weeks, once there is something to adopt.
Outcome is the slowest of the 3 to become readable, in my experience, because a saving tends to show up only once the new way of working is genuinely being used, and then only once enough time has passed for it to be more than noise. In my own practice a read at around 90 days is the first one I would put in front of anyone, and only if a baseline exists to read it against.
Is hours saved a bad measure?
No. Hours saved is a perfectly good measure and a poor survey question. The trouble is not the quantity, it is the instrument: asked at day 30 and answered from recall, it produces a memory rather than a measurement.
Recorded bottom-up against a baseline captured before the build, workflow by workflow, the same quantity becomes a record rather than an impression — an estimate still, but one you can point at. The fix is in how it is collected, not in abandoning it.
We have already built the system and never took a baseline. What now?
Reconstruct one now, and label it honestly as a reconstruction. Start with whatever records already exist — calendars, job files, ticket counts, billing narratives — and only then ask the people doing the work what each workflow used to cost them. Record their estimates as estimates, and note which ones they are confident about.
What it is worth depends on what you are reconstructing from, not on the date at the top of the document. Built on good records it can be stronger than a baseline of rough recollections written down in advance; built on recollection alone it is weaker. Either way it is a great deal better than nothing, because it at least forces the conversation to be specific about workflows rather than general about feelings. Then start the monthly bottom-up tracking immediately, so that from this point forward the record is written down as you go rather than assembled from memory at the end.
Who should own the measurement, the builder or the team?
Split it the way the layers split. Delivery is the builder's to evidence, because the builder controls it. Outcome is largely the team's, because the team controls whether the time released gets used or reabsorbed. Adoption is genuinely shared, which is why it is the most useful number in the first month and the one most worth reviewing together.
A measurement framework that makes the builder solely responsible for the team's adoption is measuring the wrong party — which is the specific mistake that got my own first version rewritten.
What is a reasonable adoption signal at day 30?
Something plain enough to be counted without a dashboard. How many of the priority workflows have been used at all, by how many of the people they were built for, and how recently.
A handful genuinely in weekly use is a healthier day-30 position than a long list built and barely touched. If adoption is low, ask which of several possible causes is operating before concluding anything: the workflow was not the bottleneck, the habit has not formed, nobody was ever told it was theirs — or the thing genuinely is not good enough yet. The adoption number tells you to ask. It does not tell you the answer.
Dr Bruno Oliveira — PhD · Associate Professor, University of Bath. Founder of GustoMind.ai. Builds and installs AI operating systems for expert-led firms, running the same system daily in his own work.