AI Operating Model

The Handover Test: Can Your Team Run the System Without You?

The demonstration is never the test. The test is the first ordinary Tuesday after the person who built the system stops answering — and 3 checks decide whether the team passes it.
By Bruno Oliveira 18 min read August 26, 2026

What the Evidence Says About Adoption and Key-Person Risk

65% of studied projects had a truck factor of 2 or fewerAvelino et al., arXiv 2016 (open-source projects)
133 popular open-source projects in that truck-factor corpusAvelino et al., arXiv 2016
53% of answers were positive or partially positive on the estimatesAvelino et al., arXiv 2016 (the authors’ own hedge)
≈24% more pull requests merged by adopters — an output proxy, not valueMurphy-Hill et al., arXiv 2026 (Microsoft rollout)
4 months window across which that adoption lift persistedMurphy-Hill et al., arXiv 2026

The demonstration is never the test. The test is the first ordinary Tuesday after the person who built the system stops answering.

An AI installation is not finished when it works. It is finished when the team can run it without the person who built it. That is not a sentiment about empowerment. It is a property of the system, it can be examined deliberately, and it is the difference between a capability a firm owns and a dependency it has acquired.

I have argued elsewhere that the layer a firm genuinely owns is its brain — the decisions, the curated knowledge and the encoded judgement no competitor can rent. This article asks the question that follows immediately, and that decides whether any of it is real: can your team run it without the person who built it?

The argument in 60 seconds

  • An installation is not finished when it works. It is finished when the team can run it without the person who built it. That is a property you can test, not a feeling you arrive at.
  • 3 checks decide it. Documentation: could a new joiner run the workflow from the written page alone? Exception: when the system meets a case it cannot handle, is there a named owner and a visible queue? Change: when the work changes, who updates the system, and how long does that take?
  • The change check has a hard form, and it is the one that settles ownership: whose account does the routine actually run from? A team that cannot edit a routine cannot change it, however capable it is and however good the manual is.
  • The Handover Test is not administered by the consultant. It is submitted to by the consultant — and the first artefact to point it at is the handover pack itself.
  • Ownership is not a feeling. It is an account. While a scheduled routine still runs from the installer's login, the team does not own it, however good the manual is.
  • What makes a system spread may be visible peer use rather than the manual. In Microsoft's early-2026 rollout to tens of thousands of engineers, first use spread primarily through social networks. That is coding agents rather than expert firms, so I take the mechanism as a hypothesis and leave the rest.
  • Key-person dependency has a name: the truck factor. A system only its builder can run is not a capability. It is a dependency wearing a capability's clothes.
📋

Get the AI Opportunity Spotter™

The 1-page framework I use to identify highest-impact AI use cases in any business

In this article:

  • Generating table of contents...

What Is the Handover Test?

The Handover Test is 3 questions that establish whether a team can run a system without the person who built it: a documentation check, an exception check and a change check. It is designed to take an afternoon rather than a project, and to be failed early rather than discovered late.

The documentation check. Could a new joiner run the workflow from the written page alone — not from memory of a demonstration, and not by asking the person who built it? The test is deliberately unkind about the source of the knowledge, because a team that can run the system only because it watched somebody run it has learned a performance rather than a procedure. Performances decay silently. A written procedure decays too, but visibly — and a gap somebody can see on a page is a gap somebody can correct.

The exception check. When the system meets a case it cannot handle, is there a named owner and a visible queue — or does it fail quietly? Quiet failure is the expensive kind, because it announces itself least. A system that stops obviously tends to get noticed and fixed quickly. A system that keeps running while producing subtly wrong output can travel a long way into client-facing work before anybody notices.

The change check. When the work changes — and it will — who updates the system, and how long does that take? A regulation moves, a client asks for a different format, a workflow acquires a new step. If the answer is that the installer updates it, the system carries an external dependency that will not expire on its own, and every change the firm cannot make itself costs a conversation with somebody outside it.

This third check has a soft form and a hard form. The soft form asks who knows how to update the system. The hard form asks whose account it runs from — and that is the one that settles ownership, which is why a later section is given to it in full.

Notice what none of the 3 checks asks: whether the system is impressive. Impressiveness is the property most easily demonstrated and least reliably transferred.

The Distinction That Matters
Working is not the finish line. Transferable is.

Why Does an Impressive AI System Still Fail After the Installer Leaves?

Because the properties that make a demonstration impressive and the properties that make a system transferable are different properties, and only one of them is visible in the room.

A demonstration selects for the happy path by design. The person running it chooses the example, avoids the awkward case without meaning to, and carries in their head the undocumented context that makes the whole thing work — which file to start from, which phrasing the system responds to, what to do when the output looks wrong. None of that is dishonesty. It is the residue of having built the thing, and it is invisible precisely because the builder no longer notices they are supplying it.

3 failure signatures follow, and each one is easy to mistake for success:

  • The workflow that works only when the right person starts it. Nothing is broken. The system simply has an undeclared prerequisite with a name and a diary.
  • The exception with no home. Everybody agrees the system cannot handle a certain category. Nobody has agreed who deals with that category instead.
  • The change nobody can make. The system is not broken; it is frozen at the shape the work had on installation day, and drifting further from it every month.

All 3 pass a demonstration. All 3 fail the Handover Test.

Who Is the Handover Test Actually For?

The consultant, first. The checks look like something a firm applies to its consultant's work, and that is the wrong way round. The Handover Test is not administered by the consultant. It is submitted to by the consultant — and the honest place to start is the handover documentation itself.

A handover pack for one installation — a small expert-led consultancy — went through a formal release-readiness audit with an explicit verdict. The verdict was hold: do not send the current bundle.

The underlying work was judged strong, coherent and professional. It was held anyway, because the audit's framing was exact: this is not a rewrite problem, it is a release-control problem. 2 failures generalise cleanly.

First, the documents described the target state as though it were already live. Certain controls and routines were presented as running when some were not yet cut over, and others existed as specifications rather than live jobs. The required correction was straightforward: finish the work, or label future-state items plainly as pending.

Second, the navigation announced its own incompleteness. Both documents carried contents pages stating that page numbers were placeholders still to be verified, and many of the numbers were simply wrong. Handover artefacts are judged as finished objects, and visible provisionality undermines the exact claim a handover is making — that the client can now proceed alone.

Neither failure was a failure of thinking. Both were failures of release control, which is a different discipline and an easier one to skip. The audit required that the pack not go out in that state. That gate is the product.

💡 A Manual That Describes the Target State as Though It Were Already Live Is Not Documentation. It Is a Promise.

The distinction sounds pedantic until somebody relies on it. A team that discovers the difference on an ordinary Tuesday has been handed a dependency, not a capability — and it discovers it at the worst possible moment, which is while trying to do the work.

The correction is cheap and unglamorous: finish the item, or label it plainly as pending. What is not acceptable is a document that quietly reads as though the future had already arrived.

What Does a System That Passes Actually Look Like?

I design handovers around 4 properties, and they are the ones I look for when I examine somebody else's install.

The instructions are readable prose, not code. The rules the system follows live in plain-language files the client can open and read. Changing behaviour is a sentence rather than a ticket — file drafts in a drafts sub-folder from now on — and the system records what changed and on whose instruction. Documentation a team cannot edit is a manual for somebody else's system.

Every scheduled routine can be rebuilt from a saved specification. A written specification exists for each one, so that, given the same access, credentials and dependencies, a routine can be rebuilt rather than permanently lost and the install is reproducible by someone who was not there when it was built. That is the difference between a system and a performance. It also quietly changes the risk conversation: the question stops being what happens if this breaks and becomes how long does it take to rebuild.

Changes are logged, dated, newest first. A short monthly read of those logs is a cheap governance habit, because it makes when did that change? an answerable question rather than a matter of trust. In my experience firms skip it not out of carelessness, but because an audit trail rarely feels worth the effort until the first time somebody needs one.

The architecture absorbs new work, not the consultant. A new area of work starts by copying a starter template, and existing routines pick it up automatically because they read the folder structure rather than a hidden list. New folders arrive with a default rule that treats the material as maximally confidential until real terms are captured, so the system is conservative when it does not yet know the rules. The design intent is that growth does not require the installer — because growth is a common reason a firm ends up calling one back.

These properties are what the third verb of an operating model — govern — looks like once it stops being a principle and becomes a set of files somebody can open.

What Does It Actually Mean to Own an AI System?

This is the change check pushed to its hard form, and, in my experience, it is where many apparently complete handovers are quietly still incomplete. Ownership has an exact factual form, and it is not a feeling about the manual. The question is: whose login does it run from?

In the stack I build on, a scheduled routine belongs to the account that created it, and it can only be edited from a session signed in as that account. The specifics vary between platforms. The question does not. While a routine runs from my account, the client does not own it, however good the documentation is, because every change has to be relayed through a third person.

Handover completes when the account a routine runs from is an account the firm controls — whether that is the person who uses it day to day or a firm account somebody inside the firm administers. What is not ownership is a routine still running from mine. Moving a routine into the firm's own control is not an administrative detail to be tidied up at the end of an engagement. It is the moment ownership becomes true, and it is worth scheduling deliberately rather than letting it drift into the category of things everybody assumes somebody else has done.

In my experience this is also the question that most often exposes a handover that looked complete. The manual can be excellent, the training can have gone well, the team can be genuinely capable — and the whole thing can still be sitting inside an account the firm does not control. Ask the question early enough that the answer is still cheap to change. Naming who owns what is also the third section of a firm's first AI doctrine, which is where most firms write the answer down for the first time.

Handover completes when the account a routine runs from is an account the firm controls.

What Makes a Handed-Over System Actually Get Used?

Visible peer use may do more work than a manual. That is an impression from the installs I have run rather than a measured result, and the study below offers some support from a different setting. It should unsettle anyone who writes good documentation, including me.

In a study of Microsoft's early-2026 rollout of command-line AI coding agents across tens of thousands of engineers, researchers found that first use spread primarily through social networks, and the authors suggest that organisations should treat visible peer use as central to rollout strategy. The same paper reports that adopters merged roughly 24% more pull requests than they otherwise would have — a figure its authors immediately qualify by noting that a merged pull request is not the same as the value it delivers, and one that persisted across their 4-month window.

That is software engineers adopting coding agents, not expert firms installing an operating model, so I take the diffusion mechanism and leave the rest — as a hypothesis worth testing in expert firms rather than a result imported into them. Stated as a hypothesis: a system spreads because colleagues can see colleagues using it, and the manual is necessary without being sufficient.

If that holds outside software, there is a practical instruction hiding in it. Where handover is a social act as much as a documentary one, the sequence matters: the first routine to move into someone else's hands should be one whose output other people see. A routine that quietly tidies a folder is a poor first handover, however useful. A routine that produces something a colleague reads on a Monday morning is a good one, because the handover advertises itself.

What Risk Is the Handover Test Actually Managing?

Key-person dependency — a risk software engineering already named and tried to measure. The truck factor is the minimal number of people who have to be hit by a truck, or simply quit, before a project is incapacitated. In a study of 133 popular open-source projects, researchers estimating it automatically found that 65% had a truck factor of 2 or fewer.

That is open-source software, studied in 2016, and not professional-services firms installing AI. The authors are also candid about their own method, which they proposed precisely because there was no consensus on how to calculate the metric: they surveyed developers from 67 of those systems and report a positive or partially positive answer in 53% of cases, alongside 84% of valid answers agreeing or partially agreeing that the identified authors were the main authors of their systems.

The concept remains useful within those caveats, and that is what I am borrowing: not a benchmark, but the fact that researchers thought key-person dependency worth estimating systematically at all. It is a property of a system rather than a mood about one. Used as an analogy, it points at a practical design objective. An installation should end with more than 1 person able to run the thing; an installation that concentrates capability in whoever understands the machinery ends with only 1 person able to run it. AI systems concentrate capability quietly, because much of what makes one work is configuration, convention and accumulated context that is rarely visible to colleagues who did not set it up.

Working on this inside your firm?

GustoMind works with expert-led firms on exactly this — from a readiness diagnostic to a full AI operating model. No pitch, just a conversation about where you are.

How Do You Run the Handover Test on Your Own Firm?

Pick one workflow that matters and run the 3 checks against that workflow rather than against the system as a whole. A whole-system audit tends to produce a report. A single workflow tends to produce a decision. Choose the workflow whose absence would be noticed first, because that is where a dependency costs most.

Then ground each check in observable evidence rather than in assurances. Give the written page to somebody who did not build it and watch, without helping. Find the most recent case the system could not handle and trace where it actually went — if the trail ends in somebody's inbox, the queue is not visible. Ask who updated the system for the last real change in the work, and when; if the answer is that nobody did, the useful question is how long it has been frozen without anyone noticing. Then push the change check to its hard form and ask whose account the routine runs from, because it takes very little time and, in my experience, produces the most uncomfortable answer of the lot.

A failure on any of these is not an indictment. The fix may well be smaller than the discovery suggests — a page rewritten, an owner named, a routine moved to the person who uses it. The cost is in finding out late, not in finding out. The same logic scales down: for a firm that has not yet built anything, the architecture question comes first, and the Handover Test is what makes the answer stick.

The Verdict

A good installation is designed from the first day to make the installer unnecessary. The measure of the work is not what the system does while I am there. It is what the team can do with it after I have gone — run it, fix it, change it, and extend it into work I never saw. That is why the Handover Test is not a checklist handed over at the end, but the standard the work is built to meet from the beginning.

If your firm has an AI system that only 1 person can genuinely run, the 3 checks will show you where the dependency sits. The Diagnose pathway is usually where that conversation begins: see how the engagements work, or send me a note and I will share how I run the checks. And if you would rather watch the thinking develop first, The AI Operating System — the fortnightly LinkedIn newsletter — is where this test was worked out in public, one edition at a time.

Frequently Asked Questions

How long does the Handover Test take?

It is designed to fit an afternoon, if it is run against one workflow rather than an entire system. The documentation check takes as long as it takes somebody unfamiliar to attempt the workflow from the written page. The exception check traces a past case. The change check combines tracing a past change with establishing which account actually controls the routine, and its hard form takes very little time. Running it across the whole system at once can expand the exercise enough that it never gets completed at all.

What is a truck factor, and why does it matter for an AI system?

The truck factor is the minimal number of people who would have to leave before a project is incapacitated. It comes from software engineering, where researchers have estimated it across real project corpora. It matters for AI systems because those systems concentrate capability quietly: much of what makes one work is configuration, convention and accumulated context that is rarely visible to colleagues who did not set it up. By the same analogy, a system only its builder can run resembles a project with a truck factor of 1 — and the underlying fact is about the firm rather than about the builder.

Is the Handover Test something we run on a supplier, or on ourselves?

Both, and the order matters. Run it on yourself first, on a system somebody inside the firm built, because internal systems accumulate the same kinds of dependencies and nobody is invoicing for them. Then run it on any external installation, before the engagement ends rather than after. A supplier who welcomes the test is telling you something useful. So is one who does not.

Our system works well, but only 1 person really understands it. Where do we start?

With the documentation check on the single workflow that would be missed first, and with the account question on the routines that run unattended. Those 2 moves are designed to fit an afternoon, and they can convert a vague worry into a specific, short list. Resist the instinct to start by writing a comprehensive manual: a manual written before the checks tends to document the system the builder believes exists, which is precisely the failure a release-readiness audit is built to catch.

Does handover mean the consultant disappears?

It means the consultant stops being load-bearing. There is a real difference between a firm that calls for a considered opinion on a new problem and a firm that cannot change a file path without booking a call. The first is a professional relationship. The second is a dependency, and it should not be sold as a service.

Dr Bruno Oliveira — PhD · Associate Professor, University of Bath. Founder of GustoMind.ai. Builds and installs AI operating systems for expert-led firms, running the same system daily in his own work.

✅ The 3 Checks, in the Form You Can Use on Monday
  1. The documentation check. Give the written page for one workflow to somebody who did not build it, and watch them attempt it without help. The gap between what the page says and what they need is the finding.
  2. The exception check. Find the most recent case the system could not handle and trace where it went. If the trail ends in somebody's inbox, there is no visible queue. If nobody can find the case, that is the finding.
  3. The change check, in both forms. Name the last real change in the work, then ask who updated the system for it and when. If nobody did, the system is frozen — and the question becomes how long it has been frozen unnoticed. Then push the same check to its hard form: whose account does the routine actually run from?

The hard form of the third check takes the least time and, in my experience, produces the most uncomfortable answer.