Picking a TTS voice that doesn't sound like a robot
Why synthetic narration usually fails on pacing rather than the voice model, and the script-level changes that fix it — most of which cost nothing.
Almost everyone approaches this backwards. They try four providers, decide the voices all sound slightly artificial, pick the most expensive one, and end up with narration that still does not hold attention.
The voice model is rarely the problem. Pacing is the problem, and pacing comes from the script.
What actually makes narration sound synthetic
Listen to a synthetic read that feels wrong and the fault is almost never the timbre. It is one of these:
Nothing is emphasised. A human reader stresses the word that carries the meaning. A model reads every word at the same weight unless the sentence structure tells it otherwise.
No breath. Human narration has micro-pauses everywhere — before a turn in the argument, after a number, around a name. Synthetic reads run sentences together at a constant rate, which is exhausting in a way listeners feel but cannot name.
Sentences that are too long. This is the biggest one. A thirty-word sentence gives the model no natural place to breathe, so it delivers the whole thing in one flat run and the listener loses the thread halfway through.
Wrong pronunciation of the things that matter. Names, places, acronyms and numbers. One mispronounced word every thirty seconds is enough to keep reminding people they are listening to software.
Fix the script before you change provider
These cost nothing and will improve a read more than upgrading tiers.
Cut every sentence over about twenty-five words. Split at the natural join. Check your longest sentence in the Word Counter — if it is over thirty words, that sentence is where your narration falls apart.
Punctuate for breath, not for grammar. A comma where you want a short pause, a full stop where you want a real one, a paragraph break where the topic turns. This is your main pacing control and it is free.
Write numbers as they should be spoken. “Two thousand and twenty-six” rather than “2026” when you want it read that way; “fifteen per cent” rather than “15%”. Models guess otherwise, and they guess differently between providers.
Spell out anything unusual phonetically. A name the model will get wrong is worth writing as it sounds. Nobody sees the script.
Read it aloud yourself first. Anywhere you run out of breath, the model will too. This one test catches most problems in a few minutes.
Then choose a provider — by video length
The length of your videos should decide this more than anything else.
Under five minutes. Use the cheapest usable option. At this length the differences in breath and emphasis barely have time to register, and the price gap across providers is roughly fifteen-fold.
Five to fifteen minutes. The middle tier earns its cost here. Listeners start noticing flatness somewhere in this range.
Over fifteen minutes. Pay for the best voice you can. Long continuous narration is where synthetic fatigue actually shows up in retention graphs, and it is the one case where the premium tier is straightforwardly worth it.
Price your own script across every provider in the Voiceover Cost Calculator — the difference on a real script is usually smaller in absolute terms than people expect.
Test properly, once
Take ninety seconds of your actual script — not a sample sentence, and not someone else’s demo — and generate it on two or three providers. Then listen on phone speakers at low volume, because that is how most of your audience will hear it. Differences that are obvious on headphones frequently vanish on a phone.
Do this once, decide, and stop. Switching providers every month costs you consistency, and a recognisable voice is worth more to a channel than a marginally better one.
Settings worth changing
Speed. Most defaults sit near 150 words per minute. Documentary pacing wants slower; energetic commentary wants faster. Getting this right does more for the feel of a video than the model choice does — and it changes your runtime, which you can check in the Script Length Calculator.
Stability or variance. Higher stability means a more consistent but flatter read; lower means more expressive but occasionally strange. For long narration, sit slightly toward stability.
One voice, not several. Some creators switch voices between sections. It rarely works. Consistency is part of how an audience recognises a channel.
The part nobody says
A synthetic voice reading a genuinely interesting script keeps people watching. A perfect voice reading a rewritten article does not. If retention is poor, changing provider will not fix it — and the money is better spent on having something worth narrating in the first place.
Questions
Which TTS provider sounds the most natural?
For long-form English narration, ElevenLabs is still the reference point, with PlayHT close behind. But the gap between providers is smaller than the gap between a well-written script and a badly written one.
Can listeners tell it is an AI voice?
Many can, and it matters less than people assume. What loses viewers is not the synthetic timbre but flat pacing — the sense that nothing is being emphasised because nobody is thinking.
Should I use the cheapest voice?
For videos under about five minutes, usually yes. Over roughly fifteen minutes of continuous narration the differences in breath and emphasis start to fatigue listeners, and the better voice earns its price.
Does punctuation really change the output?
Substantially. Full stops, commas and paragraph breaks are the main pacing controls you have, and rewriting punctuation alone often improves a read more than switching provider.