Upload a photo and an audio clip. The AI moves the mouth and face to match the speech, so a still portrait becomes someone talking.
A talking photo is a still portrait animated to match an audio recording - the mouth shapes, jaw and expression follow the actual speech rather than looping a generic talking motion. You supply both the photo and the audio; the tool syncs them.
One face, looking toward the camera, mouth visible and unobstructed.
A recording of what they should say.
The mouth and expression follow the speech.
An MP4 with the audio already in it.
Mouth shapes track the speech rather than looping a generic talking motion.
No video of the person, no recording session.
It still has to look like them, which means the motion stays within what the photo supports.
New accounts get 30 credits. No card.
Put a face on a voiceover when getting on camera is not practical.
A photo of someone reading their own words.
Give an illustrated or generated face a voice.
This one needs a capability we have not wired up yet. Everything else here — text to video, image to video, face swap, animate old photos — works today. Leave an email and we'll tell you the day ai talking photo opens.
An audio file. This tool syncs a face to speech you provide - it does not write the script or generate the voice.
One face, front-facing, mouth visible. Profiles, sunglasses covering the eyes, hands near the mouth and very low resolution are the four things that most often break it.
Only with their permission. Putting words in a real person's mouth without consent is the one use of this we will not help with, and depending on where you are it is also illegal.