Most books have never had an audiobook. Not because nobody wanted one, but because human narration is expensive โ you are paying a skilled performer for studio time, then paying again for editing and mastering, and the arithmetic only works if the book will sell enough copies to recover it. For the long tail of published writing, it never will.
Synthetic narration changes that arithmetic completely. That is the real story, and it is worth separating from both the marketing around it and the reflexive dismissal of it.
I work on machine learning systems, so let me start with what the technology is actually doing, because the capability claims in this space are frequently ahead of the reality.
How voice synthesis works now
Modern text-to-speech does not stitch together recorded fragments, which is what made older systems sound the way they did. It generates audio directly from a learned model of speech.
Voice cloning adds a conditioning step. The model learns a compact representation of what makes a particular voice distinctive โ timbre, resonance, characteristic pitch contours โ and then generates new speech conditioned on that representation. This is why a few minutes of clean recording can be enough to produce a recognisable imitation, and why more recording, covering a wider emotional and prosodic range, produces a better one.
The output is very good at the level of the sentence. Timbre is often convincing. Individual phrases carry plausible intonation.
Where it remains weak is everything above the sentence. Prosody across a long passage โ the way a skilled narrator builds tension over a page and releases it, holds a pause a beat longer than is comfortable, lets a line land โ requires understanding what the text means and deciding how to perform it. Synthesis systems approximate this from surface features. Over ten minutes it is passable. Over ten hours the flatness accumulates, and listeners describe it as fatigue rather than as a specific fault they can point to.
The other reliable failure points: distinguishing characters in dialogue-heavy scenes, comic timing, invented proper nouns, and any word whose pronunciation depends on meaning. It cannot know whether that is a lead you follow or the metal, unless the surrounding context is unambiguous enough to disambiguate โ and sometimes it is not.
Where it is genuinely the right tool
Books that would otherwise have no audio edition at all. Technical books, niche non-fiction, regional-language writing, academic work, backlists from small presses. The comparison here is not synthetic narration versus a good human performance. It is synthetic narration versus silence, and silence wins nothing.
Accessibility. For readers who need audio to read at all, coverage matters more than performance quality. An adequate synthetic reading of a book that has no human recording is a real gain, and the accessibility community has been using synthetic speech for decades without apology.
An author's own voice on their own work, with their consent. There is something legitimately appealing about a writer whose book is read in their voice when they could not afford weeks of studio time. The consent question is trivial here, because it is their voice.
Drafts and proofing. Hearing your own manuscript read back is one of the most effective revision tools available, and cheap synthetic narration makes it available for a whole book rather than a chapter.
Where it is not
Anything that depends on performance. A full-cast production, comic novels, poetry, books whose voice is the whole point. If the narration is part of what people are buying, a machine approximating narration is a worse product being sold as an equivalent one.
Which brings up the part of this that is usually skipped.
The narrators
Audiobook narration is a real profession, and a substantial number of working performers have built careers in it. The honest position is that synthetic narration is a direct competitive threat to some of that work, and pretending otherwise because the technology is interesting is not a defensible thing to do.
The argument commonly offered โ that machine narration only serves books which would never have been recorded anyway โ is partly true and not sufficient. It is accurate that most of the new synthetic catalogue consists of titles nobody was going to pay a human to read. It is also true that once a cheap option exists, the pressure on budgets for mid-tier work goes one direction, and mid-tier work is where most narrators actually make their living. Both things are happening at once.
The consent issue is sharper still and is not about the technology. It is about contracts. A narrator who records for a publisher may find that the contract permits the resulting audio to be used as training data for a synthetic voice, which can then be used to produce narration that competes with them, without further payment. Some performers have signed such terms without understanding them. Union and industry pushback on exactly this point has been one of the more substantive developments in the field, and the reason it matters is that voice is the asset. Once a usable model of your voice exists and someone else controls it, you have permanently lost the scarcity your career was built on.
There is also the estate question, which people raise as a hypothetical but is not one: whether a deceased author or performer's voice can be reconstructed to narrate work they never read. Technically it is straightforward. Whether it should be done is a question about the dead and their families, and there is currently no settled answer.
The economics, without the fantasy
Production cost for synthetic narration is low enough that for most indie authors the meaningful expense becomes editing time rather than generation. That is a genuine change and it deserves to be stated plainly.
What it does not do is create demand. An audiobook of a book nobody is buying is an audiobook nobody buys. The revenue projections that circulate in this space โ where a modest backlist quietly generates a full-time income once it is narrated โ are stories, not data, and I would treat any specific number in them as invented until someone shows the receipts.
The distribution side is also shifting under everyone's feet. Royalty terms are the lever that determines whether any of this is worth doing, and they change: Audible has revised the royalty structure for titles distributed through ACX, and exclusivity terms remain the main trade-off for authors deciding where to publish. Anyone planning around specific percentages should check the current terms rather than trusting a figure from an article, including this one.
Several large retailers now run their own machine-narration programmes, generating audio editions directly from ebooks. That is worth knowing because it changes what the market looks like: the marginal cost of an audio edition is heading towards zero, and when the marginal cost of something goes to zero the supply of it goes up enormously and the average quality goes down.
What I would actually tell an author
If your book has no audio edition and no realistic path to one, synthetic narration is worth doing. Record enough source audio to give the model range, budget real time for fixing pronunciation of every name and invented term, listen to the entire output rather than sampling it, and be upfront in your product description that the narration is synthetic. Listeners mind less about the fact than about being misled.
If your book is dialogue-heavy, comic, or lyrical, or if narration would be a meaningful part of what readers are paying for, hire a person. The gap is still large there and it is exactly where the gap is largest.
And if you are asked to sign anything granting rights to your recorded voice, read that clause twice. It is the most valuable thing in the contract and it is usually the shortest paragraph in it.
Tags
Taresh Sharan
support@sharaninitiatives.com