AI Avatar

How I Built an AI Avatar That Presents My Work — and Why I Tell People It Is Not Me

The presenter in my first film is an AI avatar, built from my own footage and speaking in a clone of my own voice. This is the receipt: the four tools, what broke, the checks around them, and why the film opens by saying it is not me.
By Bruno Oliveira • 15 min read • September 23, 2026

The Build, in Numbers

15 Minutes of Real Recording, on 3 CamerasThe recording session of 9 September 2026 — a Sony, a Logitech and an iPhone
3 HeyGen Takes to Get the Hands Right130 HeyGen credits across the 3 takes, as published with the film on 16 September 2026
56 Seconds of Finished FilmThe published cut, labelled as an AI avatar and cloned voice within its first second
80% time savings on AI-assisted tasksAnthropic
6% of orgs are AI high performersMcKinsey

This is not me. The presenter in the film below is an AI avatar, built from my own footage and speaking in a clone of my own voice, in a room that does not exist. The real me appears for a few seconds, on stage, so you can compare the two.

Film 1 of the series, first published on LinkedIn on 16 September 2026. 56 seconds, with sound.

A note on what you are watching. The presenter is an AI-generated likeness of me, made from my own recordings and my own voice. What it says is what I think, and I signed off every word.

A 56-second film has room for one sentence per idea. This page is the paragraph each sentence stood for: why I built it, how it was built tool by tool, what broke along the way, and why the first thing the avatar says is that it is not me.

The argument in 60 seconds

  • An AI avatar should say what it is before it says anything else. Said in the first three seconds, the disclosure is meant to change the question a viewer brings to the film, from whether it is really me to how it was made — and the second question is the one worth answering.
  • I build before I recommend. I test everything before it goes into a client's business, and a build of my own shows what a demonstration is not designed to show: what the tool does on an ordinary day.
  • The build was 4 tools and 1 real recording. 15 minutes on 3 cameras, HeyGen for the avatar, ElevenLabs for the voice, and Claude Code for the edit. DaVinci Resolve cut the pipeline's first film; when Resolve fell over, the edit moved into code, and that is how this one was assembled.
  • What broke is the useful part. The hands, the editing software and the moment after the avatar's last word each let the build down, and none of those failures reached the published film.
  • It is good enough to stand behind, and I expect that within a year or two most of us will hand a script to a version of ourselves and watch it deliver.
  • The discipline is the same for any AI system in a business: real inputs, a working build, a check on the output, and a person responsible. For this film, that person is me.
📋

Get the AI Opportunity Spotter™

The 1-page framework I use to identify highest-impact AI use cases in any business

In this article:

  • Generating table of contents...

Why Does the Film Open by Saying It Is Not Me?

Because I would rather tell people than risk them finding out, and in my judgement being found out would cost more trust than the film could earn. Said in the first three seconds, the disclosure is designed to change the question a viewer brings to everything that follows: not whether it is really me, but how it was made.

The disclosure is not only spoken. The frame carries a label, "AI avatar · cloned voice", from its first second; the two short passages of real footage — me on stage, and the recording session the avatar was built from — are labelled as real; and the end card names the avatar, the cloned voice and the tools behind them. Someone watching with the sound off still knows what they are looking at.

I did it this way for a reason that has nothing to do with the technology. My work is trusted judgement, and my face and my voice are part of how that trust is carried. My judgement is that a likeness which has to be explained after someone asks has already spent some of that trust, and that one which introduces itself as a likeness protects it. The design is meant to move the conversation from "is that really him?" to questions about the craft — which is the conversation a builder wants to have.

There is a practical reason as well. My concern was that if the film tried to pass the avatar off as me, the question of whether it was real would drown out everything it said.

Opening frame of the film: an AI avatar of Bruno Oliveira in a generated room, captioned This is not me and labelled AI avatar, cloned voice

Why Would I Build an Avatar of Myself?

Because I test everything before it goes into a client's business. A vendor's demonstration shows a tool on its best day. Building one myself, from my own footage and on my own machine, showed me what it does on an ordinary day — and what it takes to make the result good enough to put my name to.

The film puts it more simply: I built it to see, first hand, what is now possible for any of us. The point was never a novelty for my own channels. It was to find out what one person with a recording, a computer and a clear argument can now produce, and what it costs them in attention to do it responsibly.

There is a second reason, and it is the one that matters for a business. A capability I have built and broken myself is one I can advise on without borrowing someone else's confidence. When a firm asks me whether AI video is ready to put in front of its clients, I can now answer from a build rather than from a brochure.

How Was the Avatar Actually Built?

From one real recording and four tools. On 9 September 2026 I recorded 15 minutes on 3 cameras. HeyGen Avatar V built the avatar from that footage and a set of real photographs of me, ElevenLabs cloned my voice from earlier recordings, and Claude Code did the edit — in code, after DaVinci Resolve, which had cut the pipeline's first film, fell over on a large import.

Tool by tool, this is the receipt.

  • The recording. 15 minutes of me talking to camera, captured on a Sony, a Logitech and an iPhone at the same time. The avatar's likeness, and some of its habits, trace back to that session and to the photographs, as the next section shows. The 15 minutes is the recording, not the build: the film went live a week later, on 16 September.
  • ElevenLabs cloned my voice from earlier recordings of me speaking, read the narration in that voice, and generated the original music underneath it.
  • HeyGen Avatar V turned the footage and photographs into a presenter that can speak a new script, and generated each take of the film to the narration the voice clone had read.
  • Claude Code, Anthropic's AI coding agent, which I run on my own Mac, did the edit: the room, the captions, the cuts. On the first film this pipeline produced, it drove DaVinci Resolve Studio through Resolve's own scripting interface for the camera moves and the assembly, with the captions and graphics made in code. When Resolve fell over on a large import, the assembly moved into code as well, and that is how this film was put together: a motion engine written in HTML renders every caption and graphic frame by frame, and short Python scripts drive ffmpeg to place the avatar in the room, move the camera, mix the sound and assemble the cut.
  • The room is a generated image. It does not exist.
  • The count: 3 HeyGen takes to get the hands right, 130 HeyGen credits across them, and 56 seconds of finished film.

Frame titled Four tools, one real recording: HeyGen made the avatar, ElevenLabs cloned the voice, Claude Code and DaVinci Resolve did the edit

The tools are the least durable part of the receipt. Any one of them could be replaced by a competitor next year without changing the argument. What made the film usable was the checking around them, and that is the part that transfers.

💡 Disclosure Is Part of the Build, Not a Caveat

The label, the opening sentence and the end card were in the film's plan before its first take was generated.

They were not added after a review, and they are not a legal footnote. My judgement is that an avatar which explains itself only when asked has already spent trust it will struggle to earn back, and that one which introduces itself in the first three seconds gives the audience the chance to judge the work on its merits — which is the only judgement worth having.

What Broke, and What Caught It?

Three things broke, and none of those failures reached the published film. The early takes kept joining the hands, the editing software fell over mid-build, and the avatar did something odd once its last word had been spoken. Each was dealt with before anything went out.

The hands. The first two takes of this film kept bringing the palms together, a pose that was already in the footage the avatar's look had been cut from. Earlier in the same build, a take generated with a motion prompt asking for the hands to rest apart did not remove the habit either: a first review by eye judged that the prompt reduced the habit, and a later frame count judged that the prompt made no difference. Changing the look is what worked: a waist-up still from the same recording, with HeyGen's expressive motion switched on, produced the third take — the one in the film. A check of one frame every 2 seconds found the hands apart in all 28 frames it sampled. The lesson I took from it is small and transferable: when a generated output keeps repeating a habit, look at what it was built from before you rewrite the instruction.

The software. DaVinci Resolve crashed while importing a large composite file, and its scripting connection hung after the relaunch. Rather than wait on it, the edit moved to a route that does not depend on Resolve at all: the same cuts, camera moves and layers, rebuilt as code. A pipeline worth relying on needs a second route for the step most likely to fail, and this one found its weakest step early.

The ending. It took three cuts. A frozen final frame was tried first and rejected. The next cut let the avatar settle after its last word, and in that settle its mouth and hands did something I would not put my name to. So in the published film the avatar's part ends as the last word lands: the end card comes in as "me" finishes and covers everything the avatar does afterwards.

Behind all three sits a routine of checks at two stages. Each take is checked before it is cut in, with one frame sampled every 2 seconds for the hands. The finished film is then checked as a whole: its opening and closing frames, the loudness of the mix, the timing of the voice against the narration, and a full decode of every frame, so that a file which will not play to the end is caught before a viewer finds it. Then I watch it, and nothing is scheduled until I have signed it off.

Good enough to stand behind is a judgement about the output, and that judgement stays with a person.

Is It Good Enough to Stand Behind?

Yes, for what it is: a presenter for a short, scripted argument of my own, checked and signed off before anyone sees it. It is not a replacement for me in a room. It cannot take a question, read an audience or change its mind, and it is labelled everywhere it appears.

"Good enough to stand behind" was a deliberate choice of words. It is not a claim that the avatar is indistinguishable from me. That is not the standard, and a film built to be indistinguishable would be solving the wrong problem. The standard is whether I am content for it to carry my argument, under my name, to people whose trust I care about. Likeness and lip-sync remain judgements I make on each take, every time, rather than properties I can certify once.

And where does it go next? I expect that within a year or two most of us will hand a script to a version of ourselves and watch it deliver. That is an expectation, not a forecast I can prove. My reason for it is what this build suggested to me: the expensive part is no longer the camera. It is the judgement about whether the result is good enough — and that part does not move to the machine, however good the machine becomes.

What Does This Mean for AI in a Business?

The discipline that made this film usable is the one I bring to any AI system in a business: real inputs, a working build, a check on the output, and a person responsible. The tools change from one job to the next. The four parts do not, and a system that is missing any one of them is not ready to put in front of a client.

Real inputs. The avatar's habits came from the recording its look was cut from, not from the prompt. A firm's AI system is the same: its output is bounded by the documents, templates, standards and past work it is given to reason from. Generic inputs produce generic output, however capable the model.

A working build. Not a demonstration on a good day, but something that runs end to end, can be run again tomorrow with a new script, and survives one of its tools falling over. The demonstration is never the test; the test is whether the system still works on an ordinary day, in other hands.

A check on the output. Every take of this film was checked before it was cut in, and the finished film again as a whole. Every output an AI system produces for a client deserves the equivalent: a defined check, done every time, by someone who knows what good looks like. In the operating model I use with firms, this sits in the Govern work — named owners, the sign-off, and a gate for anything sensitive.

A person responsible. An avatar cannot be responsible for what it says, and neither can a model, a routine or an agent. Someone has to sign the output off and answer for it afterwards. In a firm, that is the named owner of the system — a decision a firm's first AI doctrine should write down. For this film, it is me.

Frame listing the four parts of the discipline: real inputs, a working build, a check on the output, and a person responsible

The film ends on four words: that person is me. The rest of this page is the working behind them.

Working on this inside your firm?

GustoMind works with expert-led firms on exactly this — from a readiness diagnostic to a full AI operating model. No pitch, just a conversation about where you are.

What Comes After This Film?

More films from the same system. The day after this film appeared on LinkedIn, I published the 50-second film the same system made for my website's homepage. The next film uses the avatar to present the three-stage map I use with clients, from AI in a browser tab, through configured workspaces, to an operating system the firm owns.

If your firm is weighing AI video, or any AI system that will speak or write in your name, and you would like it built with those four parts in place from the first day, see how the engagements work and send me a note. And if this way of building is useful, it continues fortnightly in The AI Operating System, my LinkedIn newsletter.

Frequently Asked Questions

Is the presenter in the film really you?

No. It is an AI avatar that HeyGen built from my own footage and real photographs of me, speaking in a clone of my voice that ElevenLabs made from earlier recordings. The film contains two short passages of real footage — me on stage, and the recording session the avatar came from — and both are labelled as real. What the avatar says is what I think, and I signed off every word.

Which tools did you use?

HeyGen Avatar V for the avatar, ElevenLabs for the voice and the music, and Anthropic's Claude Code for the edit. Claude Code drove DaVinci Resolve Studio for the first film the pipeline produced; after Resolve fell over, the edit moved into code, which is how this film was made: an HTML motion engine for the captions and graphics, rendered frame by frame with Playwright, and Python scripts driving ffmpeg for everything else.

How long did it take?

The recording took 15 minutes; the film went live a week later. The recording session was on 9 September 2026 and the film was published on LinkedIn on 16 September. The time in between went on the edit, the takes and the checks, not on the camera.

Why tell people it is an avatar at all?

Because the alternative is hoping nobody notices. In my judgement, a likeness revealed after the fact reads as a trick, however good the work. Said in the first three seconds, the disclosure gives the audience the chance to judge the argument and the craft on their merits, and it is how I intend to keep the trust my face and my voice carry where it belongs — with me.

Should my firm build an AI avatar?

Only if it answers a question your firm actually has. The lesson that transfers is not the avatar but the discipline around it — real inputs, a working build, a check on the output and a person responsible — which applies to any AI system you put in front of a client, from a drafted report to a film.

Dr Bruno Oliveira — PhD · Associate Professor, University of Bath. Founder of GustoMind.ai. Builds and installs AI operating systems for expert-led firms, running the same system daily in his own work.

✅ Four Questions Before Any AI Output Leaves the Building

You do not need an avatar to use the discipline. Take the next thing an AI system produces for a client — a draft, a summary, a proposal section, a film — and answer four questions about it, in one sentence each.

1. What real input did it come from? Name the documents, data or recording it was built on. If the honest answer is "whatever the model already knew", it is not yet your firm's work.

2. Would the build run again tomorrow? If it only worked because someone nursed it through once, it is a demonstration, not a system.

3. Who checked it, and against what? A named person and a named standard, not "it looked fine".

4. Whose name is on it? One person who will answer for it if it is wrong.

If any answer is "nobody" or "I am not sure", the output is not ready to leave the building — however good it looks.