AI Adoption

Before You Claim a Return: What to Measure in the First 30 Days

A return figure produced at day 30 and answered from recall tells you more about the measurement than about the system. The timing, the instrument and the unit of analysis are each working against you at once — and the first month can buy you something far more useful instead.
By Bruno Oliveira • 21 min read • September 16, 2026

The Measurement, in Numbers

3 Layers a Single Return Figure Collapses Into OneDelivery · adoption · outcome — different owners, different timescales
20% Speed-Up Developers Still Believed In After a Trial That Measured 19% LongerMETR randomised trial on early-2025 tools, July 2025; the same team’s February 2026 point estimates on later tools suggest a speed-up, and the authors call that data an unreliable signal
15% Average Gain in Issues Resolved per Hour, With Two Groups Moving DifferentlyGenerative AI at Work, Brynjolfsson, Li and Raymond, Quarterly Journal of Economics, 2025 — 5,172 customer-support agents
80% time savings on AI-assisted tasksAnthropic
6% of orgs are AI high performersMcKinsey

An earlier piece in this series ended on a test: an installation is not finished when it works, it is finished when the team can run it without the person who built it. The question that follows is always the same one.

You own the system now. What tells you it is working?

In my experience the answer that arrives first is nearly always a number — hours saved per week, a percentage, something that can go into a board pack by the end of the month. It is a fair question and a reasonable instinct.

It is also the wrong instrument for the first 30 days, and the reason is not that measurement is difficult. The reason is that 3 separate things are working against you at once: the timing, the instrument, and the unit of analysis.

The argument in 60 seconds

  • A return figure produced at day 30 and answered from recall tells you more about the measurement than about the system. Three separate bodies of evidence break three different parts of that calculation.
  • The timing is working against you. The established economics of general purpose technologies predicts that early measured productivity understates real value by construction, because the complementary investment is not captured in the numbers.
  • The instrument is unreliable. In one narrow randomised trial on early-2025 tools — software only, and its authors explicitly decline to generalise it — experienced developers expected a 24 per cent speed-up, took 19 per cent longer, and still believed afterwards that they had been sped up by 20 per cent. The lesson is about self-report, not about AI, and the same team's later data on later tools points the other way while they call it an unreliable signal.
  • The unit of analysis is wrong. A 15 per cent average gain in issues resolved per hour sat alongside movement in two directions: better speed and better quality for the less experienced, small speed gains and small quality declines for the most experienced. One average can be entirely true and still teach nothing.
  • Hours saved, answered from recall rather than recorded, is not a measurement. It is a memory — and in my experience, a memory shaped by whatever number the team was given at the start.
  • The first 30 days should buy a baseline and an adoption signal. The baseline is a short exercise, and it is the only thing in this list that stops being observable the moment the system is in place.
📋

Get the AI Opportunity Spotter™

The 1-page framework I use to identify highest-impact AI use cases in any business

In this article:

  • Generating table of contents...

Why Is a Return at Day 30 the Wrong Question to Ask?

Because the number being asked for is an outcome figure, and the instrument reaching for it is, in my experience, somebody's recollection of how the month felt. The system has barely been used and the habits around it have not formed. A recollected outcome number produced under those conditions is weak evidence in either direction.

None of which means an early number is impossible. Where the work already leaves records — matters opened, tickets closed, documents produced — an early reading can be perfectly real. The customer-support study below measured from system records rather than from anybody's recollection, which is exactly why its numbers mean something. The trouble arrives when there are no records, and the figure has to come from memory.

It helps to separate 2 things that the word return quietly welds together. Measuring is the act of recording what happened. Proving is the act of settling an argument about why.

A firm at day 30 is being asked to prove, from recollection, an effect the relevant economics gives good reason to expect will be understated at this stage. Three problems, stacked — and each one is enough on its own to produce a confidently wrong answer.

One thing to be clear about before going further. The 30-day and 90-day structure in this article is the practice I recommend and work to, not a schedule any of the research below establishes. The studies show why an early self-reported figure is fragile. Where you draw the lines afterwards depends on the workflow, the exposure and the design of the evaluation — a judgement rather than a finding.

None of that is an argument for not measuring. It is an argument for measuring the things that are observable now, and for understanding that the lasting value of an AI installation does not sit in the tool at all. It sits in the operating system a firm builds around the models — the shared context, the routines, the judgement written down. That is the thing eventually worth a number, and it is not the thing a month of use will have finished building.

Why Does the Timing Work Against an Early Measurement?

Because the value is being created and paid for well before it becomes visible in output. Brynjolfsson, Rock and Syverson described the productivity J-curve in a 2018 working paper revised in 2020. It is about general purpose technologies broadly, with AI named among them, and the date is worth naming rather than hiding, because the pattern predates this cycle entirely.

Their argument is structural. When a general purpose technology arrives, the organisation invests in complements — new processes, new products, new business models, human capital — and those investments are often intangible and poorly measured in the official numbers, even when they create genuinely valuable assets.

The consequence is precise, and it runs in both directions: productivity growth is underestimated in the early years of a new general purpose technology, and later, when the benefits of those intangible investments are harvested, it is overestimated.

Read that as an instruction rather than as an excuse — and note what it does and does not say. It says an early reading is likely to understate, not that nothing is visible; a firm can and should observe plenty at day 30. What it warns against is treating an early productivity figure as a settled verdict, because it risks pointing a decision in whichever direction the noise happens to fall.

The same model, incidentally, is why a return claimed too late can flatter as badly as an early one disappoints. The discipline is not optimism or pessimism. It is knowing which part of the curve you are standing on.

Why Can a Self-Reported Saving Not Be Trusted?

Because people have been measured misjudging their own speed — and not correcting themselves afterwards. The sharpest available evidence is a randomised controlled trial run by METR and published in July 2025.

16 experienced developers worked through 246 real issues drawn from large open-source repositories they had contributed to for years, with each issue randomly assigned to allow or forbid the use of AI tools.

Three numbers matter. Before starting, the developers expected AI to speed them up by 24 per cent. Measured, they took 19 per cent longer on the issues where AI was permitted. And afterwards, having lived through it, they still believed they had been sped up by 20 per cent.

The caveats belong here in the body rather than in a footnote, because the finding is easy to misuse. The authors are unusually explicit about what they are not claiming. They do not claim evidence that AI fails to speed up many or most software developers, noting that their developers and repositories may not represent the majority of software work. They do not claim it beyond software development, because software development is all they studied. They do not claim that near-future systems would behave the same way in the same setting. And they do not claim that there is no more effective way of using the same tools that would produce a speed-up in the very setting they tested. They describe the result as a snapshot of early-2025 capability in one relevant setting.

That snapshot has since been followed up, and the follow-up matters more than the original for a firm thinking about its own measurement. In February 2026 the same team published a note explaining that they are changing the design of the experiment.

A second study, begun in August 2025 with a larger and more varied group, produced raw estimates that pointed the other way: for the developers who had taken part before, a point estimate of an 18 per cent speed-up, with a confidence interval running from a 38 per cent speed-up to a 9 per cent slowdown; for newly recruited developers, a 4 per cent speed-up, with an interval running from a 15 per cent speed-up to a 9 per cent slowdown. Both of those intervals cross zero, and the original result carried a wide interval of its own.

What the authors did with that finding is the part to borrow. Rather than announcing a reversal, they said the data gives an unreliable signal of the current effect, and explained why. Developers increasingly declined to take part because they did not wish to work without AI. Between 30 and 50 per cent of participants said they had chosen not to submit certain tasks for the same reason. They also believe a lower rate of pay than the first study offered contributed to who took part. And time measurement became unreliable for developers running several agents at once.

Their considered view is that developers are probably genuinely more sped up now than in early 2025, and that their own data is only very weak evidence for the size of that change.

So 2 things moved between the studies, and it is worth keeping them apart. METR's own view is that developers are probably genuinely more sped up now than in early 2025 — a statement about outcomes rather than an isolation of any single cause. What changed alongside that is that the measurement stopped being readable. The first experiment worked: it caught a perception gap that self-report alone would never have revealed. The second ran into conditions that made the same design hard to interpret at all.

And the line in that note which should stop any firm about to write a 30-day number into a board pack is this one: some developers self-report very high speed-ups, and as the same team documented earlier, those estimates can be quite unreliable.

So the point is emphatically not that AI makes experts slower. The point is that self-report did not correct itself even after the experience, and that a research group who measure this for a living have twice had to work hard to get a readable answer. That is the exact instrument a 30-day self-reported return depends on. It is the same reason a self-reported number in a vendor deck earns a question rather than a decision — and your own team's recollection is a self-reported number too.

I have been on the wrong side of this myself. An early version of my own measurement framework led with hours saved per week, and I took it out before it reached anyone. It measured the client's behaviour rather than my delivery, and it leaned on a figure that anchors on whatever number you supply.

Tell a team to expect a particular saving and, in my experience, the reported figure tends to land near it — at which point the number is no longer independent evidence of anything. That is an observation from my own practice rather than a research finding, and it is the reason the framework was rewritten before it was ever used.

💡 The Line That Decides the First Month

Asked for at day 30 and answered from recall, hours saved is not a measurement. It is a memory.

Hours saved is a perfectly good quantity. METR recorded it in a randomised comparison, and this article later recommends tracking it. The failure is not the number — it is the instrument used to collect it.

A figure produced from recollection, at the end of a month in which the habit has not yet formed, is a report about how the month felt. Recorded workflow by workflow against a baseline captured before anything was built, the very same quantity becomes a record rather than an impression — still an estimate, but a specific, dated and traceable one.

Why Does One Average Hide the Only Finding Worth Having?

Because averages are built to absorb exactly the variation a firm needs to act on. The third anchor is the positive one, and it is positive on properly measured evidence.

In Generative AI at Work, published in the Quarterly Journal of Economics in 2025, Erik Brynjolfsson, Danielle Li and Lindsey Raymond studied 5,172 customer-support agents while an AI assistant was introduced in stages across the workforce, so that at any moment some had access and others did not, with productivity measured from system records as issues resolved per hour rather than from what anyone remembered. Access raised that measure by 15 per cent on average, with substantial heterogeneity across workers. The full reference, for anyone who wants it: volume 140, issue 2, pages 889 to 942.

The average is the least interesting number in the study. Less experienced and lower-skilled workers improved both the speed and the quality of their output. The most experienced and highest-skilled workers saw small gains in speed and small declines in quality.

Note what the headline measure actually is: issues resolved per hour — a count of problems genuinely closed, not of time spent looking busy. Which is why the average can look uneventful while 2 different things are happening underneath it.

The earlier working paper version of the same study put the gain for novice and low-skilled workers at 34 per cent, with minimal impact on the experienced and highly skilled, and reported suggestive evidence that the assistant was spreading the practices of more able workers and helping newer people move down the experience curve faster.

Sit with the shape of that for a moment. One reported average of 15 per cent sits above a group getting faster and better, and a group getting slightly faster on a quality reading that slipped slightly.

A firm that reports only the average has measured something true and learned little it can act on. It cannot say where the system is working, who it is working for, or what to do on Monday. And the movement that would most change what I would do next — a quality dip among the most experienced people — is the kind a single headline figure is least likely to surface.

One number can hide the only finding worth having. Which is why the unit of analysis is worth deciding before the measuring starts, rather than after the number disappoints.

The first 30 days are not where a return is proved. They are where the ability to prove one is either built or lost.

What Should You Measure in the First 30 Days Instead?

Delivery and adoption, which are observable now, plus a baseline captured before anything is built. The problem underneath all 3 difficulties above is the same: a single return figure collapses 3 different layers into 1, and then attributes the result entirely to the technology.

  • Delivery is the builder's responsibility. Is the thing built, tested and documented, and are the routines actually running?
  • Adoption is shared. Which workflows are genuinely in use, by whom, how often?
  • Outcome belongs to the team and sits downstream of both. Has time actually moved out of the week?

When one blended number disappoints, the diagnosis is unavailable. Nobody can say whether the system was not built, not adopted, or built and adopted but pointed at the wrong work. Those are 3 different problems with 3 different owners, and a firm that separates them can investigate a disappointing month instead of arguing about it.

So anchor the first 30 days on delivery and adoption, which are observable, controllable and leading. Track outcome, but do not headline it. This is the measurement half of the operating model itself, where governing the boundary and refusing to claim a return before it is real are the same discipline seen from two sides.

It is also where measurement meets governance, because separating the layers only helps if each layer has a named owner. A firm that has already written down what it automates, what stays assisted and who approves the sensitive category has, to my mind, done most of that work already. A firm that has not may find the measurement argument turning into an ownership argument at exactly the wrong moment.

The 5-Field Baseline

That leaves one thing to capture, and it belongs before anything is built, because it is the only part that stops being observable the moment the system is in place. For each priority workflow, record 5 fields.

  • How often it runs. Daily, weekly, monthly, or on an event.
  • Roughly how much time it takes now. An estimate, labelled as one.
  • Whether friction is high, medium or low. Three values, not a score out of 10.
  • Who owns it. A person, not a function.
  • The bar the output must clear before anyone would rely on it.

Estimates should be produced with the people doing the work rather than handed to them as a benchmark, and labelled as estimates. An estimate the team produced is still an estimate, but it is grounded in the work rather than in a target — and in my experience far harder to argue with later.

It is a short exercise: one conversation per workflow, not a project. Ongoing tracking is then monthly and bottom-up — used or not, how often, what it produced, roughly what it takes now, and whether the quality held. That last pair is what eventually gives you a reading on realised recovery: the time the workflow takes today, recorded the same way the baseline recorded what it took before. Be honest about what that produces — an estimate of recovery, built from estimates. It is more useful than a remembered total because it is specific, dated and workflow by workflow, and it is still weaker evidence than an observed duration or an operational record. Where those exist, use them. The total emerges from the parts instead of being a top-down figure someone has to defend.

Working on this inside your firm?

GustoMind works with expert-led firms on exactly this — from a readiness diagnostic to a full AI operating model. No pitch, just a conversation about where you are.

What Is the Difference Between Validated Capacity and Realised Recovery?

Validated capacity is the credible saving available from workflows that are built, tested and ready but not yet habitual. Realised recovery is time that has actually left the week. Both are worth reading, and reading only one of them is how a working installation can get written off.

Capacity can be present well before recovery shows up, because a workflow can be ready long before it is habitual.

Read the two side by side, on the same set of workflows, and the distance between them becomes a question rather than a verdict. Habit, documentation and coaching are the first places I would look, before looking at the technology. Measure only realised recovery at day 30 and a working system can look like a failed one.

There is a second reason to hold both numbers. A gap that closes slowly is worth investigating as a training and habit question. A gap that does not close at all, over months, is worth investigating as a design question — whether the workflows chosen were the workflows that mattered. Both are hypotheses to test rather than diagnoses the numbers hand you, and neither question is even available from a single blended figure.

The Verdict

The first 30 days are not where a return is proved. They are where the ability to prove one is either built or lost.

Spend them buying a baseline and an honest adoption signal, and the conversation three months later has something to stand on. Skip them, and that conversation compares against whatever records happen to exist and, where they do not, a vague memory of how busy things felt before — which is not evidence, and cannot be argued with either.

Do that, and when a number does arrive you will at least be able to say what it was measured against.

If your firm is about to start measuring an AI installation and you would rather not spend the first month generating a figure you cannot defend, see how the engagements work and send me a note — happy to share the baseline sheet I work from. And if this way of thinking about measurement is useful, it continues fortnightly in The AI Operating System, the LinkedIn newsletter this article grew from.

Frequently Asked Questions

How soon can an AI installation honestly show a return?

It depends entirely on which layer you mean. Delivery can be shown within days, because it is observable. Adoption becomes readable within weeks, once there is something to adopt.

Outcome is the slowest of the 3 to become readable, in my experience, because a saving tends to show up only once the new way of working is genuinely being used, and then only once enough time has passed for it to be more than noise. In my own practice a read at around 90 days is the first one I would put in front of anyone, and only if a baseline exists to read it against.

Is hours saved a bad measure?

No. Hours saved is a perfectly good measure and a poor survey question. The trouble is not the quantity, it is the instrument: asked at day 30 and answered from recall, it produces a memory rather than a measurement.

Recorded bottom-up against a baseline captured before the build, workflow by workflow, the same quantity becomes a record rather than an impression — an estimate still, but one you can point at. The fix is in how it is collected, not in abandoning it.

We have already built the system and never took a baseline. What now?

Reconstruct one now, and label it honestly as a reconstruction. Start with whatever records already exist — calendars, job files, ticket counts, billing narratives — and only then ask the people doing the work what each workflow used to cost them. Record their estimates as estimates, and note which ones they are confident about.

What it is worth depends on what you are reconstructing from, not on the date at the top of the document. Built on good records it can be stronger than a baseline of rough recollections written down in advance; built on recollection alone it is weaker. Either way it is a great deal better than nothing, because it at least forces the conversation to be specific about workflows rather than general about feelings. Then start the monthly bottom-up tracking immediately, so that from this point forward the record is written down as you go rather than assembled from memory at the end.

Who should own the measurement, the builder or the team?

Split it the way the layers split. Delivery is the builder's to evidence, because the builder controls it. Outcome is largely the team's, because the team controls whether the time released gets used or reabsorbed. Adoption is genuinely shared, which is why it is the most useful number in the first month and the one most worth reviewing together.

A measurement framework that makes the builder solely responsible for the team's adoption is measuring the wrong party — which is the specific mistake that got my own first version rewritten.

What is a reasonable adoption signal at day 30?

Something plain enough to be counted without a dashboard. How many of the priority workflows have been used at all, by how many of the people they were built for, and how recently.

A handful genuinely in weekly use is a healthier day-30 position than a long list built and barely touched. If adoption is low, ask which of several possible causes is operating before concluding anything: the workflow was not the bottleneck, the habit has not formed, nobody was ever told it was theirs — or the thing genuinely is not good enough yet. The adoption number tells you to ask. It does not tell you the answer.

Dr Bruno Oliveira — PhD · Associate Professor, University of Bath. Founder of GustoMind.ai. Builds and installs AI operating systems for expert-led firms, running the same system daily in his own work.

✅ The Baseline You Can Take Before the Build Starts

This is not the whole framework. It is the one part that expires — and it takes a single sitting.

Name the workflows you would actually want a system to touch first. The ones that would matter, not a catalogue of everything the firm does.

Then sit with the person who runs each one and fill in 5 fields: how often, how long now, friction high or medium or low, who owns it, and what the output has to be good enough for. Write their estimate, in their words, and label it an estimate.

One short conversation each. No tooling, no dashboard, no scoring matrix.

Do it before anything is built, because afterwards you are reconstructing it — from whatever records exist, and from memory where they do not. Three months later this page is what gives a number something to be compared against, instead of a conversation about whether things feel busier.