An audiobook-ready manuscript is not simply an ebook sent through a voice generator. It is a listening script. Before generating audio, organize the book into sensible tracks, clarify pronunciations, rewrite passages that depend on visual formatting, and decide what should actually be spoken.
The process can be summarized as:
- Freeze the manuscript.
- Create a narration copy.
- Divide it into listener-friendly sections.
- Add pronunciation and performance notes.
- Generate a short sample.
- Listen, revise, and regenerate.
- Check the complete audiobook in sequence.
- Confirm the distributor’s current requirements.
That preparation matters whether you use Nellie, another text-to-speech service, or a human narrator.
Freeze the words before working on the voice#
Complete the substantive edit first. Changing a character’s name or rearranging chapters after narration begins can force you to regenerate multiple audio files. Keep the approved manuscript untouched and make a separate narration copy for audio-specific edits.
Remove material that slipped in from the writing process, including comments, revision notes, placeholder text, duplicate headings, and instructions such as “insert chart here.” Then identify every element designed primarily for the eye:
- Tables and figures
- Footnotes and endnotes
- URLs
- Decorative symbols
- Image captions
- Equations
- Bulleted reference lists
- Scene breaks represented only by ornaments
Decide whether to speak, adapt, summarize, or omit each one. Google’s auto-narration editor, for example, lets publishers edit or remove footnote references, tables, figure text, and page breaks. It also provides controls for pronunciation and other sound adjustments. This is a useful reminder that ebook formatting rarely transfers perfectly to speech (Google Play Books editor guidance).
Never silently remove information the listener needs. A table comparing three retirement accounts might become a short prose explanation. A decorative leaf between scenes should probably become a pause rather than the spoken words “leaf ornament.”
Build chapters around the listener’s experience#
Use one consistent heading system for parts, chapters, prologues, epilogues, and appendices. Every audio section should have a unique, descriptive file or project name, even if the spoken heading is brief.
A practical internal naming scheme might look like this:
00_Opening_Credits01_Prologue02_Chapter_One_Arrival03_Chapter_Two_The_Letter20_End_Credits
Leading zeroes keep files in order. Do not include these numbers in the narration unless they belong in the spoken book.
Long print chapters may need internal production segments so mistakes can be corrected without regenerating a large block. Keep each split at a natural boundary, such as a scene change or subsection. The listener should not hear an artificial announcement unless you intend to create a separate navigable track.
Also write the opening and closing credits as exact scripts. Resolve how the title, subtitle, series name, author name, and narrator attribution should be read. ACX’s director-notes template recommends documenting title and series phrasing, author-name pronunciation, story context, character notes, and unfamiliar names before production (ACX director’s notes). Although that guidance addresses author and narrator collaboration, the same decisions help when directing a synthetic voice.
Create a pronunciation sheet that can drive revisions#
Do not wait for the software to mispronounce a name before deciding how it should sound. Search the manuscript for proper nouns, invented terms, acronyms, foreign phrases, technical vocabulary, dates, measurements, and words with context-dependent pronunciation.
Use a table like this:
| Written form | Intended sound | Context or instruction |
|---|---|---|
| Dr. Ana Mier | AH-nah mee-AIR | Full name on first appearance |
| Aster Vale | ASS-ter vayl | Fictional location |
| 1906 | nineteen oh six | Year, not a quantity |
| SQL | sequel | Use consistently throughout |
| lead | led | Metal in this passage |
This is an invented worked example, not a universal phonetic standard. Simple respelling is often easiest for production notes, but a generator may require a different notation or an audio-text substitution.
Keep three things separate:
- Display text: what appears in the ebook or print book.
- Audio text: the wording supplied to the narrator.
- Pronunciation note: your record of the intended sound.
For example, the display text might remain “St. John,” while the audio text uses a spelling that prompts “Sin-jin” for a character whose name is pronounced that way. Test the result rather than assuming a phonetic spelling will work across voices.
Google warns that narrator accent and language affect pronunciation. Its guidance also recommends spelling out uncommon abbreviations, replacing special characters with words when they must be spoken, and checking homographs such as “read,” “minute,” and “present” in context (Google Play Books narration tips).
Convert visual prose into speakable prose#
Read difficult passages aloud before generation. A sentence can be grammatically correct and still be hard to follow by ear. Dense parenthetical remarks, long lists, repeated initials, and ambiguous punctuation often need audio-specific treatment.
Consider this invented worked example:
Print version: “The package includes A/B testing, CSV/XML import, and 24/7 support.”
Narration copy: “The package includes A–B testing, C-S-V and X-M-L import, and support available twenty-four hours a day, seven days a week.”
The best wording depends on the audience. Technical listeners may prefer familiar acronyms, while newcomers may need the terms expanded. Preserve meaning rather than mechanically spelling out every symbol.
For nonfiction, introduce lists so listeners know what to retain: “There are three warning signs.” For recipes, exercises, or procedures, consider whether the listener needs explicit step numbers. For fiction, check scene transitions and dialogue tags. A pause that looks obvious on a page may disappear in audio.
Generate a representative sample first#
Do not begin with the easiest page. Create a sample containing the risks most likely to recur:
- Dialogue between major characters
- An invented or foreign name
- A number, date, or abbreviation
- A scene transition
- An emotional passage
- A list or technical explanation
Choose the voice with the whole book in mind. Listen for pace, vocal range, accent, tone, and clarity. A dramatic voice may tire the listener during a long instructional book; a restrained voice may flatten a comic novel. Voice choice cannot repair confusing prose.
Nellie, published by Buzzle LLC, supports book and audiobook workflows. Its FAQ provides general information about book generation, exports, commercial use, and credits. Because this article is Nellie-owned commentary, treat Nellie as one available workflow rather than an independently assessed recommendation. This article does not compare current voices, supported languages, audiobook formats, costs, or distribution eligibility, all of which may change. Check current product details and your intended distributor’s rules before committing a project. Whatever tool you choose, verify factual content and review the complete manuscript and audio before distribution.
Listen in focused quality-control passes#
Reading along once is not enough. Different problems emerge under different listening conditions. Use several shorter passes:
Pass 1: Words and omissions#
Follow the narration copy line by line. Mark skipped sentences, repeated phrases, substitutions, bad pronunciations, and headings spoken incorrectly.
Pass 2: Performance and pacing#
Listen without looking at the manuscript. Note unnatural pauses, rushed lists, misplaced emphasis, abrupt transitions, and character voices that are hard to distinguish.
Pass 3: Continuity#
Play chapter endings and the beginning of the next track back to back. Check volume, voice, pace, chapter order, and repeated or missing text.
Pass 4: Listener conditions#
Try ordinary headphones and a small speaker at a comfortable volume. Sample sections while walking or doing a simple task. The goal is not laboratory testing. It is to discover whether key information becomes confusing without visual support.
Keep a correction log with the section, approximate timestamp, quoted text, problem, proposed fix, and retest status. Fix the audio text or pronunciation instruction at its source, regenerate only the affected section when possible, and listen across both edit boundaries.
Understand the tradeoffs before distribution#
AI narration can make iteration straightforward, especially when you control the underlying script. It can also mishandle ambiguous language, emotional turns, unusual names, and complex visual material. Human narration allows interpretive collaboration but still requires a stable manuscript, director’s notes, and careful listening.
Distribution rules are separate from production quality. Platforms may restrict eligible languages, content types, voices, partners, or file structures. As one example, Apple’s published digital-narration program requires eligible English ebooks to be available on Apple Books and routes production through preferred partners. Apple also says complex formatting and certain content categories may not qualify (Apple Books digital narration guidance). These conditions can change, so verify them for your title when you are ready to submit.
Your next step is small: duplicate one finished chapter, turn it into clean narration text, make a ten-item pronunciation list, and generate a demanding sample. The errors in that sample will tell you what the rest of the manuscript needs before it becomes an audiobook.



