VoxAnvil Editorial Desk · September 24, 2026
How AI audiobook narration works in VoxAnvil
A practical order of work: import the manuscript, prepare pronunciation, choose an ElevenLabs narrator, generate chapter audio, then review files against the distributor checklist.
Prepare the manuscript
AI audiobook narration in VoxAnvil starts with text, not with a microphone in a booth. You create a project and either paste the manuscript or upload a PDF, DOCX, EPUB, or TXT file. The importer extracts readable text. If it cannot find enough text, it stops and asks for a file it can read. That is intentional. The narrator can only speak words the studio actually extracted.
The next cut is the chapter list. The splitter looks for lines that begin with Chapter, Part, or Section followed by a number or a roman numeral. Those headings become chapter titles, and the body under each heading becomes the chapter. If the file does not use those headings, paragraphs are gathered until a chunk reaches about 2,500 words, then a new chapter starts. That size matches the pricing definition of a chapter credit, so a clean heading structure and the credit estimate stay in the same neighborhood.
Front matter and back matter deserve a look before you generate. A title page, a copyright line, a URL, and a table of contents are all text. A speech engine will try to read them if they sit in a chapter. Delete or isolate the lines you do not want spoken. Opening and closing credits, if you need them for a retailer, are easier to handle as their own short chapters than as leftovers at the top of chapter one.
Shape pronunciation and pacing before you generate
The expensive habit in text-to-speech is generating first and discovering the misread afterward. VoxAnvil’s preparation layer exists so the common misreads are visible while the text is still text. The preflight scan looks at the manuscript and returns suggested spoken replacements for numbers, abbreviations, contractions, names, punctuation, currency, dates, and similar trouble. A single scan reviews the first portion of a long paste, up to about 8,000 characters, so the useful habit is to scan the chapter you are about to narrate rather than assuming one pass silently rewrote the entire book.
Suggestions are not the same thing as a saved performance. Pronunciation Forge is where a reading becomes a rule for that project. You enter the phrase as it appears in the manuscript and the form that should be spoken. A rule can be a word, a name, an acronym, or a longer phrase, and it can be marked case-sensitive when the capitals matter. Starter and Pro can save those rules. When generation runs, each enabled rule is applied to the chapter. You can preview the spoken form before you commit it.
Pacing is handled with tags in the manuscript rather than with a separate audio editor at this step. A short pause tag becomes a brief break. A longer pause becomes a paragraph break. Expression tags such as emphasis or whispering remain in the text that is sent for synthesis. Sound-effect tags are stripped so they are not spoken as words. Narration QA, on Pro, will flag a bracket you forgot to close and a tag the studio does not recognize, which is the moment to fix them. An unclosed tag can be read aloud as raw punctuation, which is a worse outcome than no tag at all.
Choose a narrator
After the text is in shape, you choose the voice that will read the project. The free and starter plans list a curated set of narrator voices. The selected voice is stored on the project, and later chapters use it unless you change the choice. Consistency across chapters is the reason to pick once and stay there. Switching voices halfway through a book is possible, and it is also the thing listeners notice first.
If the book needs the author’s own speaking voice, Pro includes voice cloning. You record or upload audio you are authorized to use, confirm that authorization, and the studio creates an ElevenLabs clone. That clone can be assigned as the narrator. It is still a generated reading of the prepared text. It is not a recording of you performing each line in a booth, and it will not invent emphasis you did not leave in the manuscript.
Character casting is the optional extra for dialogue. Pro lists it on the plan. The studio can inspect an excerpt, propose characters who speak, and keep a cast with a voice and a short description of how they sound. Many authors never need that. A single narrator is a complete production path. Use the cast when the book’s dialogue is the point of the audio edition, not because the generator requires a different voice for every name.
Produce the chapter audio
Generation is per chapter. The studio prepares the chapter text, applies pronunciation rules, translates pause tags, drops sound-effect tags, and sends the result to ElevenLabs for the voice you selected. If the prepared text is longer than about 4,500 characters, it is cut on sentence boundaries into smaller requests and the audio pieces are combined into that chapter’s file. You hear one chapter when it finishes, not a pile of hidden fragments.
The pricing page treats that chapter as one credit when it is within about 2,500 words. Plans include a monthly number of those credits: 2 on Free, 50 on Starter at $14.99 a month, and 200 on Pro at $39.99 a month. Packs add more credits without changing the plan. The practical planning step is the estimator on the pricing page. Enter the word count, read the credit estimate, and choose a plan that covers the book you are actually producing. An 80,000-word manuscript is shown as about 32 credits, which sits inside the Starter allowance and well inside Pro.
When a chapter sounds wrong, fix the text or the rule and generate that chapter again. The preparation tools are there so the second pass is aimed at a cause, such as an abbreviation or a name, rather than a hope that the same text will be read differently. Listen in the project before you download. A waveform or a quick play-through will catch a repeated sentence and a missing paragraph that a file name will not.
Review the files and the distributor checklist
Finished chapters can be downloaded as MP3 from the project. The distribution workspace, listed on Pro, collects the practical questions that come after the download: separate files per chapter, opening and closing credits, a retail sample, cover art, and the loudness and format targets authors usually check for ACX and Findaway. The checklist is a reminder list inside the product. It tells you to normalize in post-production if a file is outside the loudness window. It does not master the file for you, and it does not submit it.
Use the list as a last desk check. Confirm each chapter file plays from the first word to the last. Confirm the order matches the book. Confirm you are not uploading a scratch take. Then open the retailer’s current specification and compare. Specifications for peak level, noise floor, sample rate, and credits files are theirs. If the checklist and the retailer ever disagree, follow the retailer.
Authors who also want a public listening page can use the RSS feed and embed tools included with Starter and Pro. That is distribution of your own episode or project audio on a site you control. It is a different act from sending a book to a store. Keep those two endings separate in your own notes so a website player is not mistaken for a retail submission.
What you still do by ear
The workflow above can be repeated for every chapter, and it still ends with a person. Listen for names that the rule missed, for numbers that should have been years rather than quantities, and for sentences so long that the voice runs out of shape. Narration QA will point at some of those patterns on Pro, including sentences over a few hundred characters and acronyms with no saved reading. It will not hear the scene.
A sensible first session is one chapter, not the whole book. Import the manuscript, save pronunciations for the names that appear in that chapter, generate it, and listen with the text beside you. If the reading matches the page, continue. If it does not, adjust the rule and regenerate that chapter before you spend the rest of the credit allowance. The free plan’s two chapter credits are enough for that experiment.
The manuscript intelligence page goes deeper on numbers, abbreviations, and phonetics. The comparison page explains how this sequence differs from pasting chapters into ElevenLabs yourself. Either one is the right next read once this order of work is clear.
Frequently asked questions
What file types can I import?
PDF, DOCX, EPUB, and plain text. The importer needs extractable text. If it cannot read the file, it will not invent a manuscript from an image.
How are chapters created?
Headings such as Chapter, Part, or Section become chapters. Without those headings, the text is split into chunks of about 2,500 words.
Which service speaks the chapter?
ElevenLabs. VoxAnvil prepares the text, applies your pronunciation rules, and sends the chapter to the narrator voice you selected.
Does the studio submit my book to Audible?
No. You download chapter audio and use the in-product checklist as a preparation aid. You submit to a retailer yourself and you confirm that retailer’s current rules.