AI Avatars, Dubbing and Sound: The New Voice Tools
Build a reusable AI presenter from one photo, have it read any script, dub your videos into 18 languages and generate sound effects from a sentence.
By the WorkCrafter team · how we write these guides

Faces and voices are what make a video feel made rather than generated. WorkCrafter now covers both: presenters that speak your script, dubbing that keeps the original voice, and sound design from a sentence. Here's what each one is for.
Three ways to put a person on screen
They look similar and solve different problems — picking the right one saves a lot of wasted credits.
- **AI presenter** — a ready-made host reads your script. Nothing to prepare; best for quick explainers and product news.
- **My avatar** — build a presenter once from a single photo, then reuse it forever. This is the one for a consistent brand face.
- **Talking avatar** — a character copies a reference performance you upload, matching its expressions and timing. Use it when the delivery matters as much as the words.
Building your own presenter
Go to My avatars, upload one clear front-facing photo, name it, describe its personality and pick a voice. It processes for a few minutes and then it's yours permanently. Creating one is free — you only spend credits when it makes a video.
The personality field does real work. “Warm, upbeat product host who speaks in short sentences” produces a noticeably different delivery from “calm expert explaining a technical topic”. Write it the way you'd brief a human presenter.
A natural voice reading a single line
Dub your videos into 18 languages
Dubbing translates the speech in a clip and keeps the speaker's own voice. That's the important part: your audience in another country hears you, not a stand-in. Upload the video or audio in the studio, choose the target language, and you get a new track back.
For a creator, this is the cheapest way to multiply reach: one video, a dozen language versions, same voice throughout.
Clean up a noisy recording
Voice isolation strips background noise and music from a track, leaving the speech. It's the fix for a good take recorded in a bad room. One constraint worth knowing: the clip needs to be at least five seconds long.
Sound effects from a sentence
Describe a sound and get it as audio — no library to trawl, no licensing to check. Transitions, UI feedback, ambience, impacts: it's the fastest way to make an edit feel finished.
“A soft cinematic whoosh followed by a bright confirmation chime”
“A silent cut looks unfinished. Voice, music and one good whoosh do more for perceived quality than another render at higher resolution.”— WorkCrafter
Putting it together
The tools are designed to feed each other. A realistic pipeline for a product video:
- Write the script with the text tool.
- Have your saved avatar read it — that's your spine.
- Generate two or three b-roll clips with the video tool.
- Generate a music bed and a couple of sound effects.
- Assemble it all in the editor and export.
- Dub the finished audio into your other markets.
Instrumental bed generated for a product demo
What makes a good source photo
The avatar is built once and reused forever, so ten seconds spent choosing the photo pays for itself many times over. The engine needs to see the face clearly and infer how it moves — anything that hides or distorts the face limits every video the avatar will ever make.
- Front-facing, eyes to camera, head fully in frame with a little space above it.
- Even light on the face. Harsh side light and deep shadow both cause artefacts around the jaw.
- Neutral or lightly smiling expression — an extreme expression gets baked into every clip.
- No sunglasses, no hand near the face, no heavy motion blur.
- A plain, uncluttered background so the subject separates cleanly.
If the first result feels slightly off, it is almost always the photo rather than the engine. Re-shoot rather than regenerate.
Writing a script that sounds spoken
Text written to be read and text written to be heard are different crafts. The most common reason an avatar video sounds robotic is not the voice — it is a script full of sentences no human would say out loud.
- Short sentences. If you run out of breath reading it, so will the delivery.
- One idea per sentence — listeners cannot re-read a clause they missed.
- Contractions: "you'll", "it's", "we've". Their absence is what makes narration sound like a press release.
- Punctuation is timing. A full stop is a beat; a comma is a small one. Use them to shape the rhythm.
- Read the script aloud yourself before generating. Every stumble is a line to rewrite.
Aim for roughly 140 words per minute of finished video. That figure is a useful planning number: a 60-second explainer is about 140 words, not the 300 people usually write.
Localising without losing yourself
Dubbing keeps the speaker's voice, which changes the economics of going multi-market. Instead of re-recording or hiring voice talent per language, you publish the same performance everywhere and only the language changes.
A few things are worth knowing before you scale it. Translated speech is rarely the same length as the original — German and Spanish run longer than English, so a tightly cut edit can drift out of sync. Leave a little slack in the visuals, or cut the localised version to its own audio. And check names, product terms and numbers in the output: those are exactly the words a translation layer is most likely to reshape.
The consent rule, stated plainly
Use your own face and your own voice, or the face and voice of someone who has explicitly agreed. This is not a formality. Cloning a likeness or a voice without permission breaks the platform rules and, in a growing number of jurisdictions, the law — and a takedown after publication costs far more than the conversation beforehand.
Frequently asked questions
How many photos do I need for an avatar?
One. A clear, well-lit, front-facing photo gives the best result — the same rules as a good passport picture.
Can I use someone else's face or voice?
Only with their permission. Cloning a voice or likeness without consent is against the rules and, in many places, against the law. Use your own, or someone who has agreed.
Does dubbing change the video itself?
It produces a new audio track in the target language. Combine it with your video in the editor to publish the localised version.
Why is my isolated voice rejected?
The clip is probably shorter than five seconds — that's the minimum the tool accepts.
Start with a face
Create your presenter in My avatars and have it read a paragraph. Once you have a consistent face and voice, everything else — b-roll, music, dubbing — is just assembly.


