Upload a clip and the speech you want. AI lip sync redraws the mouth so each word lines up, for dubs, new lines and AI characters.
Lips matched to every word
MP4 or MOV in, MP4 out
Any voice track
Pay per second
You get an MP4 with the new speech and matching lips. Most short clips take 2–3 minutes.
/m/
AI Lip Sync Example
One of our own tests, start to finish. Switch between the silent clip and the lip-synced result, with sound on.
The line he says: “Hi there. A minute ago this clip was silent. Now every word I say matches my lips.”
How we made this clip
01Free
An AI portrait
A photo of a person who does not exist: face to camera, mouth closed, even light.
0230 credits
Animated into a silent clip
Eight seconds of blinking and small head moves, generated from the photo with Veo 3.1 Lite.
0364 credits
Lip synced to one line
The AI lip sync redrew his mouth for every word. The rest of the frame stayed as it was.
/a/
What Is AI Lip Sync?
AI lip sync is a tool that changes the mouth movements in a video so they match a new audio track. You give it footage of a person and the speech you want, and it redraws the lips, jaw and teeth frame by frame until the words and the mouth line up.
It saves a reshoot. A dubbed ad, a corrected product name or an AI character who needs a voice used to mean filming again or animating by hand. AI lip sync does it from the clip you already have.
How AI lip sync works
The model listens to the audio and splits it into phonemes, the sounds of speech. Each sound maps to a mouth shape, called a viseme: lips closed for M, wide for A, round for O. The model then repaints the lower face in every frame to hit those shapes on time, blending them into the skin and lighting around it.
What it keeps
Framing, background, lighting, head movement and the rest of the face stay as they were. Only the mouth area is redrawn.
What it changes
Lips, teeth and jaw follow the new speech, and the new audio replaces the clip's original sound.
/e/
How to Lip Sync a Video in 3 Steps
The whole AI lip sync runs in your browser. No editing software and no plugin.
01
Upload the video
Drop an MP4 or MOV of 2 to 60 seconds. One person facing the camera, with the mouth visible, gives the best result.
02
Add the speech
Upload MP3, WAV, M4A or AAC audio up to 10 MB. Record it on your phone, export it from a voice tool or use a dub.
MP4
03
Sync and download
Tick the consent box and press Sync. Most short clips are ready in two to three minutes as an MP4.
/i/
What Makes a Good Source Video
AI lip sync can only redraw a mouth it can find. These five things decide whether it finds one.
01
A face toward the camera
Front-facing or slightly turned works best. A full side profile hides half the lips.
Why it matters
The model needs both corners of the mouth to place each shape.
02
720p or sharper
Use at least 720p for close-ups. Very small or blurry faces give soft, smeared lips.
Why it matters
More pixels around the mouth means more detail to redraw.
03
One speaker in frame
Keep one clear face in the shot. With a crowd, the sync may land on the wrong person.
Why it matters
The tool drives a single face per clip.
04
Nothing over the mouth
Hands, microphones, masks and thick beards over the lips get in the way.
Why it matters
Anything covering the mouth is redrawn badly or not at all.
05
Steady light and little blur
Fast head turns and flickering light make the mouth hard to track from frame to frame.
Why it matters
Motion blur hides the shape the model is trying to change.
/o/
When the Audio and Video Lengths Differ
The result keeps your video's length, and you pay for the video's seconds. Here is what happens when the two files don't match.
Your files
What you get
Tip
Audio as long as the video
Lips move for the whole clip
The ideal case for dubbing
Audio shorter than the video
The speaker stops talking when the audio ends
Trim the video to the audio first and pay for fewer seconds
Audio longer than the video
The audio is cut where the video ends
Speed up the line or use a longer clip
The tool shows both lengths before you sync, so a mismatch is never a surprise.
/u/
Lip Sync vs Talking Photo vs Native Speech Video
There are three ways to get a video of someone saying your words. The right one depends on what you start with.
AI lip sync
You start with
A real or generated video, plus audio
Best for
Dubbing, fixing a line, keeping a specific shot or actor
Watch out for
Needs a clear face in the footage
Price on Imgveo
40 credits for a 5 s clip
Talking photo
You start with
A single photo, plus audio
Best for
Avatars, presenters and characters when you have no footage
Watch out for
Two steps: animate the photo, then lip sync it
Price on Imgveo
30 + 64 = 94 credits for 8 s
Native speech video
You start with
A text prompt with the dialogue in it
Best for
New scenes where the model writes picture and voice together
Watch out for
You don't control the exact voice or the exact face
Price on Imgveo
From 30 credits for an 8 s clip with sound
Prices are read from our live price list. The talking photo row is the exact route we used for the example above.
/f/
AI Dubbing: One Clip, Many Languages
AI lip sync is the last step of a dub. The same footage can speak Spanish, Japanese or German without a new shoot.
1
Write the translated script
Translate the line and read it aloud. Keep it close to the original length.
2
Record or generate the voice
Record a native speaker or export the line from a voice tool as MP3 or WAV.
3
Match the timing
Trim silence at the start and end so the speech fills the clip.
4
Run the AI lip sync
Upload the original video with the new audio. Repeat for each language.
Matching the pace of another language
Some languages need more syllables for the same idea. If a German line runs long, shorten the wording rather than speeding up the voice, which sounds rushed. AI lip sync follows whatever pace the audio has.
/th/
AI Lip Sync Use Cases
Anywhere a face on screen needs to say something it didn't say on the day.
01
Ads and product explainers
Change a price, a product name or a call to action without booking the presenter again.
Try this
Keep a clean take of the presenter looking at the lens for future edits.
02
Dubbing and localization
Ship one video in several languages with mouths that match each one.
Try this
Run the most important market first and check the result before the rest.
03
AI characters and avatars
Give a generated character a voice and a line of dialogue that fits your story.
Try this
Generate the character facing the camera with a closed mouth, as in our example.
04
Fixing a single line
Swap one wrong word in a talk, a course or a social clip without reshooting.
Try this
Cut out just the sentence, sync it, and splice it back into your edit.
/l/
What AI Lip Sync Can't Fix
We tested the limits on our own clips. Here is what to expect.
Our first test: the face is hidden by a sword swing, so no face was found. The job stopped and nothing was charged.
No clear face
If the face is turned away, tiny or blurred by motion, the sync stops and your credits come back.
Several people talking
One face is synced per clip. For a conversation, cut each speaker into their own clip.
Voice and emotion
AI lip sync matches the mouth, not the delivery. A flat recording still sounds flat.
/w/
Formats, Limits and Pricing
What the AI lip sync accepts, what it returns and what it costs.
Input
Video: MP4 or MOV, 2 to 60 seconds, up to 100 MB, at least 360p. Audio: MP3, WAV, M4A or AAC up to 10 MB.
Output
An MP4 at the video's own length, with the new speech as its soundtrack and the mouth matched to it.
Price
8 credits per second of video, rounded to whole seconds. Paid plans only. Failed runs are refunded.
AI lip sync can put words in someone's mouth, so we ask for permission first. Before every run you confirm that you have the right to use the face and the voice.
Use your own face, a consenting person or a character you generated.
Don't make public figures, minors or private people appear to say things they never said.
Label edited or dubbed videos where your platform asks for it.
/r/
AI Lip Sync FAQ
Want every AI model in one place?
Open Imgveo Studio to run this and every other AI video and image model from a single workspace, with one credit balance.