1

Upload a clear headshot

Front-facing, well lit, neutral expression, mouth visible, no heavy shadow across the face. The same rule applies to human faces, cartoon characters and animal subjects.

2

Choose your audio source

Either type a script and pick an AI voice from the Voice tab — natural voices across 20+ languages with a choice of gender and accent — or upload pre-recorded audio for the avatar to sync to.

Callout

No voice input means a silent video. Plan the audio before you render, not after.

3

Use Gen-2 for anything long

Lip-Motion Gen-2 costs 0 credits, renders 720p and accepts up to 2 minutes of audio. This is the workhorse for full scripts and classroom drafts.

Credits

0

Length

2 min

Resolution

720p

4

Use Gen-3 for hero close-ups

Gen-3 co-generates audio and video, so lip movement, expression and speech timing land together. It costs around 30 credits and caps at roughly 10 seconds. Trim the audio to match, or the render gets rejected.

5

Adjust delivery, then generate

Set tone, pacing and expression intensity where available, then generate. Review with sound on and watch only the mouth on the first pass — everything else is fixable with a prompt tweak.

6

Add subtitles

Auto-generate captions from the script or audio, for accessibility and for silent-autoplay feeds like Instagram and TikTok.

The skills marketplace also surfaces a refined 'Lip Gen-3.1' packaged version of this engine (Module 09). Long script? Split it into Gen-2 sections and cut them together afterwards.

Next up: Act-Motion transfer

Continue