A voice clone: how to stop paying a narrator for every clip

A voice clone: how to stop paying a narrator for every clip

The video is assembled and there is no voice — and from there it is either a microphone and retakes, or a narrator per clip, or a synthetic voice audible within a second. Below is a fourth option, its limits, and the rule on rights that matters more than quality.

The video is assembled and there is no voice — and from there it is either a microphone and retakes, or a narrator per clip, or a synthetic voice audible within a second. Below is a fourth option, its limits, and the rule on rights that matters more than quality.

---

The clip is ready and there is no narration

A familiar stopping point. The video is assembled, the editing is done and there is no voice.

Three options follow, each flawed.

Record it yourself. You need quiet, a microphone and several takes. At ten clips a month that is a job of its own, and it stops when you have a cold or are travelling.

Hire a narrator. Paying per clip, waiting, explaining revisions. At volume, a noticeable cost line and a permanent dependence on somebody else's diary.

The default synthetic voice. Free, available and audible within a second. The flat lifeless delivery people scroll past a clip for.

There is a fourth: a trained clone of your own voice.

What that is

You record a sample of your speech, the model trains on it, and from then on it reads any text in your voice.

The trained voice lives in your profile and gets dropped into any subsequent generation. Train it once and use it permanently.

Alongside sit two more trainable assets: a character — a permanent figure with an unchanging appearance — and a style, your own visual manner.

Three together give you what ordinary generators do not: recognisability. The same person, the same voice, the same delivery from one piece of material to the next.

What that changes practically

Speed. Text → finished narration in minutes. No quiet room, no microphone, no takes.

Independence from your condition. A cold, a trip, a late evening — the voice is always available.

Revisions are free. Changed a sentence in the script — regenerate. With a live recording that is a new take; with a narrator it is a new round and new money.

Scale. Thirty clips get narrated the same way one does. Particularly if it is part of a chain with an iterator.

And what matters most for content: the voice stops being the bottleneck. Production usually stops at the narration rather than at the filming or the editing.

Where the technology falls short

Honestly, because what you take on depends on it.

Emotion. Level informational speech comes out well. Irony, doubt, anger, fine intonational play — weakly. The voice sounds like yours but speaks flatly.

Long monologues. The longer the fragment, the more noticeable the sameness. Short lines are more reliable.

Difficult words and names. Terms, surnames, brand names may get read wrongly. Those get checked separately.

Emphasis. The model does not always understand which word in a sentence is the important one. Sometimes the stress falls in the wrong place and the meaning shifts.

The practical conclusion: a clone covers regular informational content excellently. For advertising, where the delivery decides, and for emotional formats, a live recording is still better.

How to get a decent result

Four rules, all about the source.

A clean recording with no noise. A room with no echo, no background, an even level.

A sufficient sample. The more material for training, the more accurate the clone. A stingy source gives a recognisable but flat voice.

An even pace when recording the sample. If you read the sample fast and indistinctly, the clone will speak the same way.

Testing on your own text. Having trained the voice, run it on a real script rather than a test sentence. The difference usually shows up on genuine material.

And a technique that improves the result substantially: break the text into short fragments and generate in parts. That way you fix one piece rather than the whole clip, and a misplaced stress does not force you to redo five minutes.

Rights: whose voice can be cloned

The point that matters more than quality.

Your own — allowed. It is your voice, no question.

Somebody else's — not allowed. A person's voice is as protected an attribute as their face. Cloning an actor, a public figure or an acquaintance without written consent is not on.

The voice of a client's employee or customer — only with written consent, including reworking and commercial use. The client obtains the consent, and you record in the contract that they confirm holding rights in the material supplied.

And separately on substitution. Using a cloned voice so a listener believes a real person is speaking in a real situation is no longer a technical subject. The format can be staged, a fact cannot.

What else there is in sound

A voice clone is part of a wider set, and the rest is worth knowing.

Music — tracks with and without vocals, extending and remixing existing material.

Video narration — a soundtrack for a finished clip: atmosphere, effects.

Lip synchronisation — if the voice has to come from a person in frame rather than sound over the top. It works confidently on short lines in close-up.

The price range: music and sound — $0.12–1.80 per generation, lip sync $0.12–1.50. Paid from the wallet on use, the cost visible before you press, money returned for anything not delivered.

How this sells

Your own saving. At ten clips a month you stop paying for narration and stop depending on somebody else's diary.

A service to a client. An expert's trained voice is the client's asset: from then on their content gets narrated without their involvement. For a busy person who does not want to sit at a microphone, that is the decisive argument.

What gets delivered: the trained voice, a set of narrated material, instructions for use.

The market price range. By my scrape of 848 listings: a short clip — $18–45, AI video end to end — $245–550, up to $610 for a finished clip on a per-piece arrangement. The median for one-off work is $30–75.

Narration on its own is small work. It becomes money as part of a set: the script, the frames, the voice, the assembly, the platform format. That is the difference between the bottom line and the top.

And a monthly arrangement as a separate format: a content assistant $100–180 a month, a short-form video manager and an online assistant $365–490. For somebody who can cover the full cycle including narration, that is a stable way in.

What is available on the platform

The voice trainer in Gen AI — training a clone. It lives in your profile and gets dropped into any generation.

Alongside — character training and style training, the three assets of recognisability.

Music and sound — nineteen models: tracks with and without vocals, extending, remixing, video narration.

Lip Sync — thirteen models, if the voice has to come from a person in frame.

File to text — speech transcription, if you need the reverse direction. $0.15.

AI Workflow — a chain of "script → frames → narration in your voice → assembly → publishing". Build it once and use it permanently; the iterator runs it down a list. The quote is visible before the run.

Marketing Studio — if you would rather not build a chain: eight ready formats with narration inside.

Registration is free and opens three days of full Basic access. Generation is paid separately, from the wallet on use.

Where to start today

One action. Record five minutes of your speech in a quiet room — simply read any text of your own aloud at an even pace.

Train the voice and run it on a real script rather than a test sentence.

Listen not for "does it sound like me". Listen for whether you would deliver it to a client. The answer to that determines whether a clone covers your task or only your drafts so far.

---

What narration gets done with in ready pipelines

A voice clone is one approach. Assembled solutions use another, and it is worth knowing.

ElevenLabs in the "digital department" pipeline turns text into professional narration. That settles the question of long lines and emotion: studio delivery comes out immediately, with no voice training.

AssemblyAI in the conversation analytics works the other way — recognising voice messages as text in seconds.

What to choose. Your own clone is needed when recognisability matters: a series of clips from one person, a personal brand. Synthesis from a ready pipeline is for when speed and volume matter and whose voice it is does not.

In Gen AI the music and sound mode holds nineteen models at $0.12–1.80: tracks with and without vocals, extending, remixing, video narration.

---

What to read next

[Lip sync: making a character talk](/en/blog/lip-sync-talking-character) — if the voice has to come from a face in frame.

[Your own visual style](/en/blog/train-your-own-style-lora) — the third trainable asset.

[Face rights in advertising](/en/blog/face-rights-in-ai-ads) — a voice is protected the same way.

[Editing expert video](/en/blog/expert-video-editing-pipeline) — where this fits in.

[AI video generation in 2026](/en/blog/ai-video-generation-guide) — the full guide to video models.