Skip to main content
WorkCrafter logoWorkCrafter.online
Tutorials

The AI Voice Tool: Text-to-Speech That Sounds Human

How to use WorkCrafter's AI voice generator — the Kokoro, Chatterbox and Qwen3 voice models, scripting tips, credit costs and best practices.

5 min read

By the WorkCrafter team · how we write these guides

Sound waves rising from a glowing microphone made of light
Image generated with WorkCrafter AI

The voice tool turns a script into natural-sounding speech for narration, videos, ads, explainers and accessibility. This guide covers the engines, how to write a script that reads well aloud, and how to avoid the robotic delivery that gives text-to-speech a bad name.

The voice models

Pick a model per generation depending on what you need from the voice:

  • Automatic (recommended): a natural, general-purpose voice. Start here.
  • Kokoro: clean, neutral delivery for straightforward narration.
  • Chatterbox: more emotion and dynamics — good for storytelling, ads and anything that should feel alive.
  • Qwen3 TTS: multilingual, for scripts in or mixing multiple languages.
  • Qwen3 Voice Clone: reproduce a specific voice from a sample (only with the speaker's consent).
  • Qwen3 Voice Design: craft a voice to a description when you want a specific character rather than a stock read.
The model sets the timbre; your script sets the delivery. Punctuation is how you direct the pauses and the pace.WorkCrafter

Write for the ear, not the eye

Text that reads fine on a page can sound stilted aloud. A few habits fix most of it:

  • Use short sentences. The voice breathes at full stops — long sentences run out of air.
  • Punctuate for rhythm: commas for short pauses, full stops and line breaks for longer ones.
  • Spell things phonetically when a name or acronym is mispronounced ("see-kwel" for SQL).
  • Write numbers the way you'd say them ("twenty twenty-six", not "2026") when the reading matters.
  • Read your script out loud first — if you stumble, the engine will too.
A script marked up with pauses and emphasis next to a waveform
Punctuation and short sentences are your direction marks for pacing.

Getting a natural delivery

If a read sounds flat, the fix is usually the script, not the model. Break dense paragraphs into shorter lines, add commas where a person would pause, and switch to Chatterbox for anything with emotion. For long pieces, generate in sections — it's easier to redo one paragraph than a ten-minute take.

A note on voice cloning

Voice clone is powerful and easy to misuse. Only clone a voice you own or have explicit permission to use. Impersonating someone without consent is both against the rules and, in many places, illegal.

Punctuation is your direction

A voice model has no idea what you meant — it only has the marks on the page. Those marks are the only way you get to direct the performance, and using them deliberately is the difference between a read that sounds human and one that sounds like a screen reader.

  • A comma is a short breath. Use them where you would naturally pause, even when a grammar checker disagrees.
  • A full stop is a longer stop and a small drop in pitch — the sound of a finished thought.
  • A line break is longer still. Use one between sections rather than hoping the voice will infer the change.
  • Three dots create hesitation. Powerful once in a script, exhausting three times.
  • A question mark genuinely lifts the intonation, so keep real questions as questions.

The quickest test costs nothing: read your script out loud at the pace you want. Wherever you naturally paused and the page has no mark, add one.

Names, numbers and acronyms

This is where most voice-overs break, and the fix is always the same — write it the way it should sound, not the way it is spelled.

  • Acronyms: write "S Q L" with spaces if you want the letters, or "sequel" if you want the word.
  • Years: "twenty twenty-six" reads correctly; "2026" is a coin toss.
  • Money and units: "nineteen dollars a month" beats "$19/mo" every time.
  • Brand names: spell them phonetically once you hear the mispronunciation, and keep that spelling in your script file.
  • URLs: "workcrafter dot online" rather than the raw address.

Long scripts: generate in sections

For anything past a couple of paragraphs, split the script and generate section by section. It is not about limits — it is about the cost of a mistake. If one sentence lands wrong in a ten-minute take, you regenerate ten minutes. If it lands wrong in a paragraph, you regenerate a paragraph and drop it into place in the editor.

Splitting also lets you change voice or pacing between sections — an introduction can be warmer than the technical middle without re-recording everything.

Choosing a voice for the job

  • Explainer or tutorial: a neutral, even voice. Personality competes with information.
  • Ad or trailer: an expressive voice — the dynamics are doing the persuading.
  • Long-form narration: test with a full paragraph, not one line. Voices that charm for five seconds can tire over five minutes.
  • Multilingual content: use one of the multilingual voices for every language so your brand sounds like one person.

Voice cloning only ever applies to a voice you own or have written permission to use. Cloning a public figure, a colleague or someone from a video you found is against the rules here and illegal in a growing number of places. This is not a formality — it is the one line in this guide with legal consequences attached.

What it costs

A voice generation costs 10 credits — about $0.20 at the $5 starter rate, and less on larger packs. Failed generations are refunded automatically, so it's cheap to try a line on two engines and keep the better read.

Frequently asked questions

Why does my narration sound rushed?

Not enough punctuation. The voice pauses at commas and stops — add them where a person would breathe, and split long sentences in two.

Which model for a YouTube voiceover?

Chatterbox usually wins for video: the extra dynamics hold attention. Kokoro is better for calm, informational reads where neutrality is the point.

Can I use the audio commercially?

Yes for the generated speech under the platform's terms. The exception is cloned voices — you need the rights to the underlying voice as well.

Start narrating

Open the voice tool, paste a short, well-punctuated script, and try it on Kokoro and Chatterbox to hear the difference. Generate long pieces section by section for full control.

A calm audio waveform stretching across a dark screen
Image generated with WorkCrafter AI
#AIvoicegeneratorguide#texttospeechAI#AIvoiceovertips#WorkCrafteraudiotool#expressiveAIvoice#multilingualtexttospeech

Keep reading

Get the next guide by email

New guides on prompting, generating and what it actually costs — a few times a month, never more. One click to unsubscribe, and we never share your address.