Making a character talk: lip sync with no filming

Making a character talk: lip sync with no filming

A silent image and a talking frame are different categories of work: the first is worth tens of dollars, the second hundreds. Below is what you need as input, where the technology fails on long speech and emotion, and three legal sources of a face for commercial work.

A silent image and a talking frame are different categories of work: the first is worth tens of dollars, the second hundreds. Below is what you need as input, where the technology fails on long speech and emotion, and three legal sources of a face for commercial work.

---

A silent character is half the job

You have learned to generate people. The frame is attractive, the face holds, the light is alive.

And the client needs a clip where a person talks. A product review, an explanation of a service, a greeting on a landing page.

At which point it turns out that a silent image and a talking frame are different categories of work. The first is worth tens of dollars, the second hundreds.

Let us take apart how it works, where the line is and what is needed as input.

What lip sync is, in plain words

Synchronising lips with a voice track. You supply an image or video of a person and an audio track — the model makes the character speak that text, hitting the sound with their lips.

The key point: this is not generating video from nothing. The frame already exists. The model adds mouth movement and expression for that particular speech.

Hence the main consequence: the quality of the result runs into the quality of the source. A bad frame will not get better by talking.

What you need as input

Three things, each affecting the result more than the choice of model.

1. A frame with a face

What you need: the face large and whole in frame, front-on or slightly turned, the mouth fully visible with nothing over it.

What breaks the result: a strong head turn, a hand near the face, large glasses with a reflection, a beard hiding the lip line, a face too small in frame.

The check is simple: if you can barely see the lips on the source yourself, the model will not manage either.

2. The audio track

What you need: clean speech with no background noise, an even pace, clear articulation.

What breaks it: music underneath, room echo, speech that is too fast, overlapping voices.

The track can be recorded yourself, taken from existing narration, or generated — Gen AI has narration and voice cloning.

3. Matching length

The speech and the frame have to match in duration. Thirty seconds of audio on a five-second clip means either cutting or extending, and both have to be planned for in advance.

Where the technology fails

Honestly, because what you can take on depends on it.

Long speech. The longer the fragment, the more noticeably the desynchronisation and the sameness of expression accumulate. Short lines work more reliably than long monologues.

Emotion. The model hits the sound but not always the intonation. Anger, irony and doubt read weakly — the face stays even whatever the content of the speech.

Profile and movement. The more the head is turned or moving, the worse the result.

Fine expression. A living person moves their eyebrows, narrows their eyes and nods while talking. A synthetic one mostly moves their mouth. That gives the result away more often than the synchronisation itself.

What follows practically: lip sync is good for short lines in close-up. For a three-minute monologue with emotion, not yet.

How this assembles into a chain

A separate generation is one frame. A working process looks different.

The full chain: text → narration → a frame with the character → lip synchronisation → assembly with sound → a finished clip.

Each step separately means download-upload-download. In AI Workflow it is one chain: each step's result becomes the next one's input. Run it once and get a clip.

There are ready templates too, so nothing has to be built from scratch. "UGC advertising from a product photo" — six steps, the character holds the product and talks to camera. "From a selfie to an advert" — six steps, your face becomes the character. "A clip from idea to result" — three steps, from one text idea.

The run continues on the server even if you close the tab.

Why a talking frame costs more

Compare with the market. By my scrape of 848 job listings over 24 days:

The gap between a clip and "end to end" is more than tenfold, and it is not about generation quality. In the second case what gets sold is the whole result: the script, the narration, the synchronisation, the assembly, the format for the platform.

A talking character is what moves you from the first line to the fourth. Not because it is technically harder but because it closes the client's task: they need not a frame but a clip they can put into an advert.

Separately on the biggest area of demand: the UGC format is 40.3% of all jobs in my dataset. That is exactly where a talking character is needed.

Cost

Generation is paid from your wallet on use, the price visible before you press, money returned on a failure.

The orders of magnitude: lip sync — $0.12–1.50 per generation depending on the model, narration and music — $0.12–1.80, animating a frame — from $0.17.

So the cost of a talking clip runs to single dollars against a fee in the tens and hundreds. The main thing is to build it into the quote from the start.

Face rights: more important than the technology

One question decides whether work like this can be delivered to a client at all.

Somebody else's face cannot be used. Not a celebrity, not a stock person, not a photograph off the internet. That is not about quality but about the client receiving a complaint and you receiving a reputation.

Three legal sources:

Your own face. A selfie, and you are the character in the advert.

A trained character. A permanent hero with an unchanging appearance, stored in your profile and dropped into any generation.

Ready avatars from the gallery — ten characters for different roles. All generated by us; they are not stock people and not somebody else's faces, so no rights questions arise.

For commercial work the last point is decisive: you can hand a client a clip and not think about whose face is in it.

What is needed for this

Gen AI — lip sync with 13 models, photo animation with 84, music and sound with 19, voice cloning. All in one section, paid on use.

AI Workflow — assembling the chain from text to a finished clip. Thirteen ready templates, including UGC advertising and "from a selfie to an advert".

Marketing Studio — if you do not want to assemble a chain: you describe the offer, attach a product photo, choose a character from the ten or your own, and get a vertical clip with sound. Eight formats, from a reel to a marketplace listing card.

The video model courses come with the Plus plan, or $9 one-off, or 1 000 experience points. They are not included in the three free days of Basic.

Registration is free and opens three days of Basic — the assistants in the creative category, including the animation and video editing expert, and an encyclopaedia of more than 190 services.

Where to start today

One action. Take a selfie, record fifteen seconds of clean speech on your phone and run it through lip sync.

Fifteen seconds is not an arbitrary length: at that length the technology works confidently, and you will see the real quality straight away rather than a perfect demo.

If the result will do, you have gained a category of work worth an order of magnitude more than static images.

---

Where to get the voice and how to assemble the whole thing

Lip synchronisation is the second-to-last step. Before it you need sound; after it, assembly.

In Gen AI the lip sync mode holds 13 models at $0.12–1.50 per generation.

Voice. In the ready "digital department" pipeline the narration is done by ElevenLabs — professional delivery from text. That also settles the question of long lines: no live recording needed.

Assembly. In the same place Creatomate joins the footage, overlays subtitles and sound into one file.

The whole ready solution

Text-to-Avatar from the archive: an article becomes a clip with your digital face in a couple of minutes. For marketers, teachers and personal-brand authors who need regular video with no crew and no studio lighting.

The economics of the digital department: the cost of one finished clip is $2.50–3, including the video model and the assembly. Against a short clip selling for $18–45, the difference is obvious.

The archive — Make, N8N and Full.

---

What to read next

[Why an AI character looks plastic](/en/blog/why-ai-character-looks-plastic) — what to fix in the frame before animating.

[A voice clone instead of a narrator](/en/blog/voice-clone) — the other half of a talking clip.

[A clip from a product photo](/en/blog/ad-video-from-product-photo) — eight ready formats.

[AI video generation](/en/blog/ai-video-generation-guide) — a full breakdown of the models and prompts.