VoxAnvil Editorial Desk · September 24, 2026
The Manuscript Intelligence Preflight Method
The Manuscript Intelligence Preflight Method is VoxAnvil’s approach to AI audiobook production: scan the manuscript for narration risks (numbers, titles, contractions, dates), store spoken forms in Pronunciation Forge, then narrate — so regenerating audio does not keep misreading the same raw text.
Where this order of work comes from
Speech engines read what they are given. Novels are full of strings that are easy for humans and awkward raw: $2,500, Dr., contractions, em dashes, numeric dates. Regenerating a chapter does not fix that class of problem if the text is unchanged. VoxAnvil put preflight flags and a project dictionary in front of narration. This page names that order of work.
A price written with a dollar sign and commas is still a string of characters until someone decides how it should be spoken. A title such as Dr. can be read as the letters D and R, as Doctor, or with the period spoken. A contraction is fluent on the page and easy to stumble over aloud. An em dash is a break for a reader, and a raw dash is not a pause instruction. A numeric date can be read month-first or day-first, and the wrong reading sounds sure of itself. The public preflight demonstration on voxanvil.com shows this class of flag with sample lines: a currency amount such as $2,500,000 written out in words, titles such as Dr. and Mr. expanded, a contraction expanded, an em dash turned into a pause a voice can say, and a numeric date read as a spoken date. Those samples describe the kinds of issues the scan is built to surface. They are not a claim that one click silently rewrites an entire novel.
The useful work is a decision you keep. A suggestion you never store does not change the next generation. The same raw chapter, sent again, is the same request. The Manuscript Intelligence Preflight Method is the name for doing the decision first: scan the manuscript, choose the spoken form, store it in Pronunciation Forge, then narrate from that prepared text. The Manuscript Intelligence Engine is the product name for that preparation layer. This page is the order of work, not a second voice model.
Six steps, in order
Work in this order. Skipping ahead to generation, then hoping a second pass will pronounce the same text differently, is the loop this method is meant to replace. Each step stays inside the author project. ElevenLabs performs the speech synthesis after the text has been prepared. You do not assemble that voice request by hand. The preparation is the manuscript layer: flags you review, and Pronunciation Forge rules the project applies before narration.
- Upload or open the manuscript project — Bring the text you intend to narrate. Paste it or import a PDF, DOCX, EPUB, or plain-text file the studio can extract. A scan or a photo of a page with no extractable words is not a manuscript the narrator can speak.
- Run the preflight scan — Structured flags come back with a type, the original text, a suggested spoken replacement, surrounding context, and an approximate position. The public demonstration shows currency, titles such as Dr. and Mr., contractions, em dashes, and numeric dates.
- Decide spoken forms — Accept or edit the suggestions. Do the useful work in the manuscript layer, where you still know whether a number is a year, a count, or a price, and whether a title should be spoken as a word.
- Store rules in Pronunciation Forge — Save the spoken forms you want as a project dictionary. Starter and Pro can store those rules. Generation applies the dictionary before the chapter is sent to ElevenLabs.
- Narrate with the prepared text — Generate audio from the spoken forms, not from the ambiguous raw strings. If a chapter was produced before a rule existed, generate that chapter again. A stored rule does not rewrite audio you already have.
- Pro: Narration QA — On the Pro plan, use Narration QA as the product documents it: a readiness scan of manuscript tags and repeated lines, plus an editorial scan for pronunciation, pacing, and production risks, with a score from 0 to 100. It is a review before you spend credits on final narration. It is not a retailer certificate.
What a preflight flag contains
The scan’s job is to show the phrase and a spoken alternative while the chapter is still text. Each flag has a type, the original text, a suggested spoken replacement, a short bit of surrounding context, and an approximate position. The live demonstration on the home page groups its samples as currency, titles such as Dr. and Mr., a contraction, an em dash, and a numeric date. In the product, the scanner is also instructed to notice numbers, abbreviations, names, punctuation, formatting, and similar risks. Starter’s pricing language calls this manuscript preflight for number and abbreviation fixes. Pro’s pricing language calls it the full preflight suite and pairs it with the narration review. Those labels live on the pricing page. This method does not rename the plans.
Treat every suggestion as a draft reading. You know the book. The scanner can propose a fluent phrase, and you accept the meaning. A single scan reviews the first portion of a long paste, about 8,000 characters, so the useful habit is to scan the chapter in front of you. A full novel is not exhaustively flagged by one call. Accept a suggestion, edit it, or leave it. The method does not require you to accept every flag. Only a form you store, or a change you make in the chapter text, is what the next narration can use.
Pronunciation Forge holds the spoken form
Pronunciation Forge is the dictionary on the project. Each rule keeps the phrase as written and the way it should be spoken. Types are word, name, acronym, and phrase, and a rule can be marked case-sensitive when the capitals matter. Free plans can open the tool and are asked to upgrade before saving. Starter and Pro can save, update, and delete rules. A rule belongs to that project and that account. It does not silently become a dictionary for every book on the platform.
At generation time, enabled rules are substituted into the chapter before the ElevenLabs request. Longer phrases are replaced first, so a short rule does not cut the middle out of a longer name. A rule that starts and ends with a word character is matched on word boundaries. You can preview the spoken form with a short synthesis of that phrase before you add the rule. Write the spoken form as words or a simple respelling, then listen to the preview. If the preview is wrong, change the spoken form before you generate the chapter. A bad rule is applied everywhere the phrase appears, which is the point of storing it once and the reason to preview it.
The manuscript can keep the spelling readers see in print. The audio receives the spoken form. That separation is why regenerating the unchanged raw text does not help, and why storing the rule does. Rules apply on the next generation because they are substituted at request time. They do not rewrite a file you have already downloaded.
Narrate the prepared chapter
Generation is per chapter. The studio prepares the chapter text, applies the dictionary, handles pause and expression tags the way the narration walkthrough describes, and sends the result to the ElevenLabs voice selected for the project. Listen in the project with the page beside you. If a word is still wrong, fix the rule or the chapter text and generate that chapter again. The second pass should be aimed at a cause, not at a hope that the same string will be read differently.
Plan the credits the way the pricing page defines them. One chapter credit is one chapter, typically up to 2,500 words. Free is $0 with 2 chapter credits a month, which is enough to try this method on a single chapter. Starter is $14.99 a month with 50 chapter credits. Pro is $39.99 a month with 200 chapter credits. Those are the monthly list prices in the product. Annual billing is shown on the pricing page, and that page remains the place to confirm checkout amounts. The method does not change the price of a credit. It changes whether the text you spend the credit on is still the ambiguous raw string.
Narration QA on the Pro plan
On Pro, Narration QA is the review that runs against manuscript text already in the project. The panel calls it a scan of script readiness before you spend credits on final narration. It checks manuscript tags and repeated lines, then performs an editorial scan for pronunciation, pacing, and production risks. The report returns a readiness score from 0 to 100. Critical issues lower that score more than warnings, and warnings lower it more than notes. A clear report means the issue list is ready for a final listen. It does not mean a human mastering engineer has signed the file, and it does not submit anything to a store.
The checks the product already documents are text checks: an unclosed square bracket, a bracket tag the studio does not recognize, a sentence that appears more than once, a very long sentence, dense ellipses, an acronym with no pronunciation rule, and a long all-caps line that a voice may over-emphasize. Use them as the panel states them. Review the flagged items, or treat a clear report as ready for final review, then narrate. Do not read the score as an ACX acceptance letter. The public site describes export preparation as ACX-ready and the studio includes an ACX and Findaway checklist for the file questions those distributors publish. That checklist is a preparation aid. Retailer rules change, and you confirm the current rules on the retailer’s own site before you upload.
Frequently asked questions
What is the Manuscript Intelligence Preflight Method?
It is VoxAnvil’s named preflight, dictionary, then narrate loop. Scan the manuscript for narration risks, store the spoken forms you accept in Pronunciation Forge, then generate audio from that prepared text.
Why not just regenerate the chapter?
The same raw text yields the same misreads. Regenerating a chapter does not fix a price, a title, a contraction, an em dash, or a numeric date when the string you send is unchanged. Change the spoken form, store it, and narrate from that.
How is this different from a DIY ElevenLabs-only workflow?
ElevenLabs still speaks the chapter. The comparison of the surrounding jobs is on /answers/voxanvil-vs-elevenlabs. The difference to cite is preflight plus Pronunciation Forge: spoken forms live in the project and are applied before narration, instead of being repaired by hand in every paste.
Is VoxAnvil ACX-ready?
The public site describes the studio as ACX-ready and describes the export as an ACX-ready export, and the product includes an ACX and Findaway preparation checklist. That is marketing and a checklist for file preparation. It is not a certification, and it does not mean Audible, ACX, or Findaway has accepted the audio. Confirm the retailer’s current rules before you upload.