Skip to main content
WorkCrafter logoWorkCrafter.online
Tutorials

AI Avatars, Dubbing and Sound: The New Voice Tools

Build a reusable AI presenter from one photo, have it read any script, dub your videos into 18 languages and generate sound effects from a sentence.

5 min read

By the WorkCrafter team · how we write these guides

An abstract glowing microphone and sound waves dissolving into colourful particles
Image generated with WorkCrafter AI

Faces and voices are what make a video feel made rather than generated. WorkCrafter now covers both: presenters that speak your script, dubbing that keeps the original voice, and sound design from a sentence. Here's what each one is for.

Three ways to put a person on screen

They look similar and solve different problems — picking the right one saves a lot of wasted credits.

  • **AI presenter** — a ready-made host reads your script. Nothing to prepare; best for quick explainers and product news.
  • **My avatar** — build a presenter once from a single photo, then reuse it forever. This is the one for a consistent brand face.
  • **Talking avatar** — a character copies a reference performance you upload, matching its expressions and timing. Use it when the delivery matters as much as the words.

Building your own presenter

Go to My avatars, upload one clear front-facing photo, name it, describe its personality and pick a voice. It processes for a few minutes and then it's yours permanently. Creating one is free — you only spend credits when it makes a video.

The personality field does real work. “Warm, upbeat product host who speaks in short sentences” produces a noticeably different delivery from “calm expert explaining a technical topic”. Write it the way you'd brief a human presenter.

A natural voice reading a single line

The same voice engine that drives the avatars — 29 languages, natural intonation.

Dub your videos into 18 languages

Dubbing translates the speech in a clip and keeps the speaker's own voice. That's the important part: your audience in another country hears you, not a stand-in. Upload the video or audio in the studio, choose the target language, and you get a new track back.

For a creator, this is the cheapest way to multiply reach: one video, a dozen language versions, same voice throughout.

Clean up a noisy recording

Voice isolation strips background noise and music from a track, leaving the speech. It's the fix for a good take recorded in a bad room. One constraint worth knowing: the clip needs to be at least five seconds long.

Sound effects from a sentence

Describe a sound and get it as audio — no library to trawl, no licensing to check. Transitions, UI feedback, ambience, impacts: it's the fastest way to make an edit feel finished.

“A soft cinematic whoosh followed by a bright confirmation chime”

That sentence is the entire prompt.
A silent cut looks unfinished. Voice, music and one good whoosh do more for perceived quality than another render at higher resolution.WorkCrafter

Putting it together

The tools are designed to feed each other. A realistic pipeline for a product video:

  1. Write the script with the text tool.
  2. Have your saved avatar read it — that's your spine.
  3. Generate two or three b-roll clips with the video tool.
  4. Generate a music bed and a couple of sound effects.
  5. Assemble it all in the editor and export.
  6. Dub the finished audio into your other markets.

Instrumental bed generated for a product demo

Generated from a one-line caption describing mood, instruments and use.

What makes a good source photo

The avatar is built once and reused forever, so ten seconds spent choosing the photo pays for itself many times over. The engine needs to see the face clearly and infer how it moves — anything that hides or distorts the face limits every video the avatar will ever make.

  • Front-facing, eyes to camera, head fully in frame with a little space above it.
  • Even light on the face. Harsh side light and deep shadow both cause artefacts around the jaw.
  • Neutral or lightly smiling expression — an extreme expression gets baked into every clip.
  • No sunglasses, no hand near the face, no heavy motion blur.
  • A plain, uncluttered background so the subject separates cleanly.

If the first result feels slightly off, it is almost always the photo rather than the engine. Re-shoot rather than regenerate.

Writing a script that sounds spoken

Text written to be read and text written to be heard are different crafts. The most common reason an avatar video sounds robotic is not the voice — it is a script full of sentences no human would say out loud.

  • Short sentences. If you run out of breath reading it, so will the delivery.
  • One idea per sentence — listeners cannot re-read a clause they missed.
  • Contractions: "you'll", "it's", "we've". Their absence is what makes narration sound like a press release.
  • Punctuation is timing. A full stop is a beat; a comma is a small one. Use them to shape the rhythm.
  • Read the script aloud yourself before generating. Every stumble is a line to rewrite.

Aim for roughly 140 words per minute of finished video. That figure is a useful planning number: a 60-second explainer is about 140 words, not the 300 people usually write.

Localising without losing yourself

Dubbing keeps the speaker's voice, which changes the economics of going multi-market. Instead of re-recording or hiring voice talent per language, you publish the same performance everywhere and only the language changes.

A few things are worth knowing before you scale it. Translated speech is rarely the same length as the original — German and Spanish run longer than English, so a tightly cut edit can drift out of sync. Leave a little slack in the visuals, or cut the localised version to its own audio. And check names, product terms and numbers in the output: those are exactly the words a translation layer is most likely to reshape.

Use your own face and your own voice, or the face and voice of someone who has explicitly agreed. This is not a formality. Cloning a likeness or a voice without permission breaks the platform rules and, in a growing number of jurisdictions, the law — and a takedown after publication costs far more than the conversation beforehand.

Frequently asked questions

How many photos do I need for an avatar?

One. A clear, well-lit, front-facing photo gives the best result — the same rules as a good passport picture.

Can I use someone else's face or voice?

Only with their permission. Cloning a voice or likeness without consent is against the rules and, in many places, against the law. Use your own, or someone who has agreed.

Does dubbing change the video itself?

It produces a new audio track in the target language. Combine it with your video in the editor to publish the localised version.

Why is my isolated voice rejected?

The clip is probably shorter than five seconds — that's the minimum the tool accepts.

Start with a face

Create your presenter in My avatars and have it read a paragraph. Once you have a consistent face and voice, everything else — b-roll, music, dubbing — is just assembly.

A glowing editing timeline with film strips and waveforms above a dark desk
Image generated with WorkCrafter AI
#AIavatarfromphoto#AIpresentervideo#AIdubbingmultiplelanguages#texttosoundeffect#voiceisolationAI#talkingavatargenerator

Keep reading

Get the next guide by email

New guides on prompting, generating and what it actually costs — a few times a month, never more. One click to unsubscribe, and we never share your address.