AI Operating Model

The AI Operating Model in Practice: Reorganise, Codify, Govern

Personal AI can make one person remarkably faster. By itself, it does not make the firm any smarter. Three kinds of work close that gap — reorganise the work, codify the judgement, govern the boundary.
By Bruno Oliveira 16 min read August 03, 2026

What the Evidence Says About AI Inside a Team

+15% more issues resolved per hour with an AI assistantBrynjolfsson, Li & Raymond, QJE 2025
34% productivity gain for novice and low-skilled workersBrynjolfsson, Li & Raymond, QJE 2025
5,179 customer-support agents in the field studyBrynjolfsson, Li & Raymond, QJE 2025
40%+ higher quality on tasks inside AI’s tested capabilityDell'Acqua et al. 2023
19pp less likely to reach a correct answer outside that frontierDell'Acqua et al. 2023 (GPT-4)

Personal AI can make one person remarkably faster. By itself, it does not make the firm any smarter — and that gap is, in my view, one of the largest pools of unclaimed value in professional work right now.

Almost every conversation about AI in business is still a conversation about tools: which model, which subscription, which feature shipped this week. The firms that pull ahead are answering a different question. Not what should we use, but what should we run. That is an operating-model question, not a procurement one, and it has a different shape entirely.

This article is the permanent home for the framework I use to answer it: Reorganise, Codify, Govern. Three kinds of work, in that order.

The argument in 60 seconds

  • An operating model is what the firm runs, not what its people use. The distinction sounds academic until you notice that one of them compounds and the other does not.
  • AI use that stays unshared and uncodified has an organisational ceiling. The gain sits with the individual; unless the context and the judgement are captured, the firm keeps little it can reuse.
  • Reorganise. The workflows change shape: shared context, scheduled routines and reusable skills, built so the repeated set-up no longer needs a person.
  • Codify. The firm's judgement is written down as plain text a system can carry — and because it is text, it stays readable and portable even when the wiring is rebuilt.
  • Govern. Named owners set the boundary of what runs without a human: the sign-off, the gate for sensitive material, and measurement before any return is claimed.
  • The model was never the whole story. My bet is that the firms which pull ahead are the ones that turn private AI use into an asset the firm owns.
📋

Get the AI Opportunity Spotter™

The 1-page framework I use to identify highest-impact AI use cases in any business

In this article:

  • Generating table of contents...

What Is an AI Operating Model?

An AI operating model is the system a firm runs rather than the tools its people use: shared context and routines that survive any individual, the firm's judgement written down in a form a system can carry, and named owners who decide what runs without a human. Three kinds of work build it — reorganise, codify, govern.

Each verb answers a different question, and the order matters.

  • Reorganise answers what shape does the work take now? It is the redesign of workflows around shared context, scheduled routines and reusable tools, so that the repeated set-up stops being a person's job. Which workflows to redesign first, and how much setup each one earns, is a question of zones.
  • Codify answers what does the firm know, and where does it live? It is the act of writing the firm's method down — voice, standards, frameworks, tested prompts, the rules for how a document gets made — in plain, human-readable text a system can carry.
  • Govern answers what may run without a human, and who decides? It is the explicit boundary: named decision rights, a gate for sensitive material, and measurement before any return is claimed.

None of the three is a technology choice. This is the part most firms get backwards. The architecture underneath — the move from a browser tab to a system you run on your own machine — is a separate and prior question, and I have mapped it in detail in From the Web to the Workshop. This article is about what changes organisationally once that architecture exists.

The Distinction That Matters
An operating model is what the firm runs, not what its people use. One of those compounds. The other does not.

Why Does Personal AI Use Hit an Organisational Ceiling?

Because the value stays where it was created. When AI use is unshared and uncodified, the context lives in one account, the method lives in one person's habits, and the firm retains little context or judgement it can reuse. The assistant can be genuinely good. The gain simply stays personal.

It is worth naming this stage plainly, and not as a failing. Firms reach it precisely because capable people went and solved their own problems. That is a healthy instinct, and it produces real individual gains. The ceiling is structural rather than a criticism: nothing about private use creates a shared asset, so nothing about it accumulates.

Recently I watched a small expert-led consultancy cross that ceiling. This was not a firm that was behind. It was already working with several AI tools, separately, with no shared layer beneath them. The interesting part was how cleanly the crossing sorted into three kinds of work — which is where the three verbs come from. The framework emerged from this install and from my own system; its generality is something I propose, not something I have observed at scale.

What Does It Mean to Reorganise the Work?

Reorganising means changing the shape of the workflow so that the set-up is no longer a person's job. In practice this is three moves: a shared context layer the system reads before any matter begins, scheduled routines that prepare recurring work without being asked, and a small set of the firm's own reusable tools that turn a repeated job into one instruction.

A shared context layer holds who the firm is, the standards it holds, and a file for each engagement, so the system starts every matter with the relevant standing instructions and history already to hand rather than waiting to be briefed. This is the single highest-leverage move, because everything downstream depends on it: routines have something to read, and tools have something to be consistent with.

A set of scheduled routines then takes over the recurring preparation, presenting prepared material before anyone asks. The shift here is subtle and easy to underrate. A prompt is something a person triggers; a routine is a job the system owns. I have written separately about what changes when work runs on a schedule rather than on a request, in Beyond Prompts.

Finally, a small set of the firm's own reusable tools — skills, in the actual stack — turns jobs done again and again into a single instruction. A report skill. A research skill. A document skill. The toolkit grows out of the work rather than being specified in advance.

The design intent behind all of this was not the same work done faster. It was to remove the repeated set-up, so that what is left for a person is the judgement, the relationship, the call that cannot be delegated. That is the first kind of work: the firm reorganised around the system.

💡 The Durable Asset Is Not the Wiring. It Is the Captured Judgement.

Captured judgement travels, and the wiring can be rebuilt. Being all text, a firm's written-down method stays readable and portable even when the platform beneath it changes.

Every firm that has ever migrated a CRM understands this instinctively — the records survive, the integrations do not. Codified judgement is the record layer of an AI operating model.

What Does It Mean to Codify the Firm's Judgement?

Codifying means writing the firm's method down as plain text a system can read and apply — not the confidential content, but the voice and tone, the frameworks, the tested prompts, and the rules for how a piece of work gets made. A prompt playbook. An operating manual. Per-engagement instructions.

This is the quieter kind of work and it is worth more, for a reason that only becomes obvious later: the durable asset is not the wiring but the captured judgement, and captured judgement is portable in a way that wiring never is.

It is also where the team gain shows up in the evidence. In a 2025 field study in the Quarterly Journal of Economics, Erik Brynjolfsson, Danielle Li and Lindsey Raymond studied 5,179 customer-support agents using an AI assistant. Issues resolved per hour rose by about 15 per cent on average — and by around 34 per cent for novice and low-skilled workers, with little change for the most experienced. The authors reported suggestive evidence that the assistant disseminated more able workers' best practices and helped newer workers move down the experience curve.

Read that mechanism carefully, because it is the argument for codification stated in someone else's data. The gain was not distributed evenly. It concentrated where the firm's accumulated method was least present, and the proposed mechanism was the spreading of practice that already existed somewhere in the organisation. That is the second kind of work: the firm codified its judgement.

Who Owns the Boundary, and What Does It Mean to Govern It?

Governing means deciding deliberately what may run without a human, and giving that decision an owner. In the install I watched it took three forms.

Explicit decision rights. Operating the system, approving how sensitive data is handled, and making specialist changes were each a named responsibility, not a single all-purpose owner. This matters more than it sounds. A single owner is a bottleneck at best and a key-person risk at worst; distributing the rights is what makes the model a property of the firm rather than of a person. Whether it has actually become the firm's property is a testable question, and I set out the 3 checks that decide it separately.

A local privacy gate. Sensitive material is processed locally before any cloud AI tool is used, with an audit record kept and human review retained for the outputs that matter. Note the shape of this: it is a designed boundary, not a blanket prohibition. A blanket ban is easy to write and it makes the whole system useless, which is why it never survives contact with the work.

A discipline about measurement. The engagement separated what was built, what the firm adopts, and any time recovered — refusing to claim hours saved before they were real. Most published AI return-on-investment numbers collapse the three, and that is precisely why so few of them survive scrutiny.

There is a hard-edged reason to draw the boundary carefully. In a field experiment run with BCG, Fabrizio Dell'Acqua, Ethan Mollick and co-authors gave 758 consultants tasks to do with GPT-4. On tasks inside the model's capability, the assisted group completed about 12 per cent more, worked around 25 per cent faster, and produced work rated over 40 per cent higher in quality. On one task deliberately placed outside that capability, they were 19 percentage points less likely to reach the correct answer. In that experiment, assistance lifted the inside-frontier work and, on the outside task, quietly made it worse.

The word doing the work in that finding is quietly. The consultants did not know which side of the line they were on. Neither will your team. Knowing where the line sits, and who owns it, is the third kind of work: the firm governed the boundary.

You do not install an operating model. You reorganise the work, you codify the judgement, and you govern the boundary.

Does the Model Hold at a Larger Scale?

The three verbs hold, but each one changes character. At leadership-team scale I would expect reorganising to move from shared folders to the redesign of what whole functions do; codifying to become institutional memory under change control; and governing to become formal governance with decision rights, risk boundaries and human oversight.

Reorganise stops being one shared context layer and becomes the redesign of what whole functions do and who holds which decision rights — because, as an operating principle, you cannot drop a new technology into an old work system and optimise the two separately. Codify becomes institutional memory under change control: the reasoning behind a pricing call or a hard exception, captured so it outlasts the people who made it. Govern becomes formal governance — decision rights, risk boundaries, human oversight, and measurement before any return is claimed.

Through all of it, more human effort moves towards judgement, exception-handling and governance, and away from doing the work by hand. That is the through-line at every size. It is also why the one-page doctrine — what goes where, what we automate, who owns it — is exactly this framework in its smallest legitimate form. The page becomes a running system. Nothing about the model requires scale to begin; it requires only that the three kinds of work are somebody's job.

What Happens When You Run the Model on Yourself?

You discover the difference between having components and having a system. I proved this pattern on myself before proposing it to anyone else. My own system runs through Claude Code inside VS Code, with Dropbox as its backbone and dozens of skills and scheduled routines — and an audit of it in July returned a humbling verdict: a fine set of organs, not yet an organism.

That verdict is the most useful thing the audit produced, and I include it here because it is the honest counterweight to every tidy framework diagram. More automation is not an operating model, even for one person. The compounding starts, on my reckoning, only once the judgement, routines and ownership belong to the firm rather than to whoever is at the keyboard. A person with forty automations and no codified judgement has a collection. A person with ten automations, written-down method and a clear boundary has a system.

Satya Nadella put the strategic version well in June. The real prize, he argued, is not picking the best model but building a learning loop on top of them — a loop that, in his words, "becomes the new IP of the firm" and that compounds rather than depreciates. My own test for whether a firm owns that loop: it can swap a general-purpose model and keep the company-veteran expertise its system has built. That is exactly what writing the judgement down as plain text buys you.

AI use at work is already widespread. Individual adoption is not the same as organisational capability, and the distance between the two is the whole subject of this article.

Working on this inside your firm?

GustoMind works with expert-led firms on exactly this — from a readiness diagnostic to a full AI operating model. No pitch, just a conversation about where you are.

Where Should a Firm Start?

With whichever verb is least true today, and almost always with context before automation. Reorganising without a shared context layer produces fast work built on nothing; codifying without routines produces a document nobody reads; governing without either produces a policy about a system that does not exist. The three questions in the box below locate a firm quickly — they are deliberately blunt, because the useful answer is usually the uncomfortable one.

Two practical notes. First, none of this requires a transformation programme or a headcount decision — the layers are useful individually, which is why the build is incremental by design and why small firms often move faster than large ones. I have written about that inversion for small businesses in Grow First, Hire Later. Second, the questions leaders actually ask when they meet this material are remarkably consistent, and I have set out the five most common of them in Five Questions Business Leaders Actually Ask.

The Verdict

Personal AI can make a person faster. A team system makes the firm something it was not before — the version designed to compound. You do not install an operating model. You reorganise the work, you codify the judgement, and you govern the boundary — and the firm that does all three stops visiting AI and starts running it. As I have argued since the first edition of the newsletter, the operating system is not something you buy. It is something you become.

If your firm is at that inflection point — personal AI working well, but none of it yet compounding into something the firm owns — the Diagnose pathway is where that conversation usually begins: see how the engagements work, or send me a note and I will share the diagnostic I use. And if you would rather watch the thinking develop first, The AI Operating System — the fortnightly LinkedIn newsletter — is where this framework was worked out in public, one edition at a time.

Frequently Asked Questions

What is the difference between an AI operating model and an AI strategy?

An AI strategy states intent — where the firm expects value, which risks it will accept, what it will not do. An operating model is the machinery that makes the intent real: the workflows, the written-down judgement, and the decision rights that determine what runs without a human. Strategy answers why and where; the operating model answers what runs, on what, owned by whom. A strategy with no operating model beneath it produces a document. An operating model with no strategy above it produces motion in an unexamined direction.

Do we need an AI operating model if we are only a handful of people?

Yes, and the small version is genuinely small. At a few people, reorganising may be one shared context layer and two scheduled routines; codifying may be a single operating manual and a page of tested prompts; governing may be a one-page doctrine naming what is automated, what is assisted and who approves the sensitive category. The three kinds of work do not require scale. They require that each one is somebody's job rather than nobody's.

Which comes first — reorganise, codify or govern?

In practice, context before automation and boundaries before scale. Reorganising usually leads, because the shared context layer is what everything else reads. Codifying follows quickly and often runs alongside it, because the act of writing the method down is what makes routines and tools consistent. Governing must be explicit before anything runs unattended — a boundary decided after an incident is not a boundary, it is a reaction.

How do we know whether it is working?

Separate three things that are usually collapsed into one number: what was built, what the firm actually adopts, and any time recovered. Delivery is easy to demonstrate and means little on its own. Adoption is the first honest signal. Time recovered is the last of the three to become measurable, and claiming it early is the fastest way to lose the room. Measurement before any return is claimed is not caution; it is what makes the eventual claim believable.

Does the operating model depend on a particular AI vendor?

It should not, and that is a design goal rather than an accident. Because the judgement is captured as plain text, it stays readable and portable even when the wiring beneath it is rebuilt. My own test for whether a firm owns its capability is whether it can swap a general-purpose model and keep the expertise its system has accumulated. If the answer is no, what the firm owns is a subscription.

Dr Bruno Oliveira — PhD · Associate Professor, University of Bath. Founder of GustoMind.ai. Builds and installs AI operating systems for expert-led firms, running the same system daily in his own work.

✅ Three Questions That Test Whether You Have an Operating Model
  1. If your best person left tomorrow, what would the firm still know? If the answer is what is in their head and their chat history, the judgement has not been codified. Written-down method is the version that stays.
  2. What ran this morning before anyone opened a laptop? If the answer is nothing, the work has not been reorganised — every output still needs a person to trigger it, which means capacity is still capped by attention.
  3. Who decides what is allowed to run unattended? If the answer is that it has not been discussed, the boundary is not governed. An undecided boundary is not an absence of risk; it is an unowned one.

One uncomfortable answer is the place to start. Three of them is the case for building the model deliberately rather than letting it accrete.