• Ai Audiobook
  • Ai Audiobook Generator

How to Make an AI Audiobook from Text or EPUB

An AI audiobook generator turns text, EPUB, or a book into chapter audio. TextSpeech is a paste-based audiobook maker — not an upload-and-forget EPUB converter.

TextSpeech Team8 min read
Books, headphones and a desk lamp beside a screen displaying an audio waveform

Short answer: An AI audiobook is chapter audio generated from a manuscript. Upload tools such as ElevenLabs Audiobooks take an EPUB and detect chapters. TextSpeech is a paste-based audiobook maker: you copy text from Word or Google Docs, paste one chapter (or less) at a time, respect the per-generation character cap, and download MP3 (Free includes download; commercial/audiobook publishing rights follow your plan). That fits indie courses, serial fiction, and podcast-style books — not “upload a 90k-word novel, get one file in five minutes.”

The job splits in two. An AI audiobook generator or audiobook maker either uploads EPUB/ebook files, or lets you create audiobooks from text one chapter at a time. Know which one you are buying.

Reviewed 14 Sep 2026 against official documentation; this is a documented workflow, not a firsthand listening benchmark. You must hold the audio rights to the text.

AI audiobook tools vs EPUB converters

Platform / guideInput modelChapter handlingTextSpeech difference
ElevenLabs AudiobooksEPUB uploadAuto-detect chapters; multi-character castingNo file upload — paste from your editor
AudiobookGenEPUB uploadFull-book conversion pipelineManual chapter splits; you control paste size
Audie.aiManuscript uploadMulti-voice fiction focusSingle-narrator workflow
TextSpeechPaste in browserYou split at chapter/scene boundariesFree preview 500 chars/run; paid MP3 + commercial use

Where TextSpeech is honest: We are a browser studio for chapter-sized pastes, not an EPUB-to-audiobook converter. If you need upload-and-forget, ElevenLabs or AudiobookGen match epub to audiobook and ebook to audiobook. If you want 1,000+ voices, Emotion tags, and the same tool you use for AI voice over, stay in Studio — but plan for multiple exports.

TextSpeech limits you must plan around

PlanCharacters / generation~English wordsMP3 exportCommercial use
Free50070–90Preview onlyNo
Basic / Pro15,000~2,200–2,500YesYes
Ultra30,000~4,500–5,000YesYes

One generation ≠ one book. Split the manuscript at chapter or scene boundaries, count the characters in each block, and allow time for regeneration and full listening checks. The number of MP3 files depends on actual chapter lengths, not just total word count.

Free plan: audition the narrator and download short blocks — commercial/sellable audiobook rights need a paid plan.

What you need before you paste

  • Audio rights to the manuscript (your work, licensed material, or public domain you verified).
  • Compare the same full page in Studio, including dialogue and difficult names. YouTube-style narrators are options to audition; select by pronunciation, pacing, and consistency in your own manuscript.
  • A naming scheme: 01-chapter-one.mp3, 02-chapter-two.mp3.
  • Basic, Pro, or Ultra for MP3 export and commercial distribution. See Terms.

Prepare the manuscript

Print formatting breaks TTS. Do this before any paste:

  • Strip or rewrite headers for speech: Chapter 3Chapter three if numbers read wrong.
  • Convert tables to spoken lists, or cut them.
  • Expand abbreviations the voice will garble (K8s, file paths, URLs).
  • Dialogue: one speaker per paragraph for single-narrator books.
  • Remove footnote markers and page refs — TTS will read p. 247 aloud.

Create audiobooks from text (chapter workflow)

This is the TextSpeech path for text to audiobook and create audiobooks from text.

  1. Open Studio. Lock one voice ID for the whole project. Write the name in a notes file.
  2. Paste chapter 1 only (or the first block under 15,000 characters).
  3. Generate. Check the beginning, middle, and end, then listen to the full chapter before export.
  4. Fix mispronunciations in the text, regenerate that paragraph if needed.
  5. Download MP3 via text to MP3.
  6. Repeat for chapter 2 without changing the voice.

Start at 1.0x, then compare the same passage at a slightly different speed. Choose a setting that keeps names, dialogue, and pauses clear; a speed value alone does not establish quality.

Illustrative calculation: a cleaned chapter

If a cleaned chapter contains 12,500 characters, it fits within the 15,000-character Basic/Pro limit for one generation. This is an illustrative calculation, not a supplied or tested sample. Count your actual chapter before generating.

  1. Rewrite 3.4 Implementation details as Section three point four. Implementation details.
  2. Turn a three-column table into three spoken sentences.
  3. Generate; skip to mid-chapter. If the voice “smiles” through dry prose, switch to a calmer voice from Voice Library before chapter 2 — not after.
  4. Export MP3. At an assumed 150 words per minute, 2,100 English words is about 14 minutes. Check the actual audio for omitted text and pause length; duration alone cannot diagnose quality.
  5. Chapter 2: same voice ID, same speed, same export settings.

EPUB to chapter text to audiobook

TextSpeech cannot upload EPUB or DOCX. For an EPUB you are authorized to adapt and that has no DRM, use this manual route. If copying is restricted, obtain an authorized DRM-free source or the original manuscript; this workflow does not remove DRM. The navigation and copy commands are documented in the calibre viewer manual.

  1. Add the authorized EPUB to calibre, select the book, and choose View to open its E-book viewer.
  2. Open Table of Contents and select the chapter. If there is no usable contents list, locate the chapter heading and check where the next chapter starts.
  3. Select the chapter text, right-click and choose Copy. For a long chapter, copy smaller consecutive sections and check their first and last sentences to avoid gaps or duplicates.
  4. Paste into a plain-text editor first. Remove navigation text, notes and repeated headings; rewrite tables and confirm the text matches the chapter. Count characters and split at paragraphs below your plan limit.
  5. Paste the cleaned block into Studio, keep the same voice settings, preview and correct it, then download MP3. Repeat in order and name the files by chapter and part.

Use an upload converter when:

  • The book already lives as a clean EPUB with real chapter marks.
  • You want auto chapter detection and one pipeline, not 25 pastes.
  • You are fine with that vendor’s narrator set and commercial terms.

Stay in Studio when:

  • The manuscript is still in Google Docs or Word.
  • You already picked a host in the voice library and want that same voice on YouTube clips.
  • You would rather QC one chapter tonight than debug an EPUB toc.

Listen, QC, and assemble

Before distributing, run these listening checks:

  • Listen to each chapter fully once before batching the rest.
  • Match loudness across files — do not peak-normalize chapter 1 to −0.1 dB and leave chapter 2 at −12 dB.
  • Keep a QC log: timestamp, issue, fix (rephrase sentence, add comma pause, spell name phonetically).

TextSpeech exports MP3. Check your destination’s encoding, loudness, silence, credits and file-order requirements and use an audio editor when adjustments are needed. Technical mastering does not grant permission to distribute AI narration.

Publish and distribute

ChannelAI narration policyWhat to do
Your site / course / podcastDepends on your rights and host rulesCheck licensing and the host’s content rules before publishing
Google Play BooksUpload guide provides a Synthesized voice setting for AI audiobooksSet the narrator type correctly and meet account, content and file requirements; this is separate from Google’s auto-narration tool
Kobo Writing LifeAccepts AI narration with synthesized-voice narrator metadataInclude author and narrator contributors and follow current upload requirements
ACX / Audible via ACXHuman narration required unless otherwise authorized; unauthorized TTS/AI is prohibitedDo not submit a TextSpeech export without authorization; mastering alone does not make it eligible

Disclose AI narration wherever the platform asks. Distribution rules change — verify on the retailer’s author portal before you promise a release date.

Common mistakes

  • Expecting one paste for the whole book on a tool with 15,000-character caps.
  • Shipping Free-tier preview audio in a paid product — no MP3, no commercial license.
  • Recasting the narrator at chapter 4 because chapter 1 “needed more energy.”
  • Adding music without checking speech intelligibility, music rights and the destination’s rules. Compare a speech-only version before keeping a music bed.
  • Using celebrity/character voices as the official audiobook host — see Terms and impersonation rules.

FAQ

Do I need to upload a PDF or EPUB?

No. TextSpeech has no manuscript upload. Copy from Word, Google Docs, or Scrivener and paste into Studio. Competitors like ElevenLabs Audiobooks and AudiobookGen are built around file upload — that is the epub to audiobook job.

Can I make a full audiobook on the Free plan?

You can preview and download short blocks (500 characters per generation). Longer chapters and commercial/sellable rights need a paid plan.

How many files will my book be?

Count the actual characters in each cleaned chapter. Divide each chapter’s count by your plan’s per-generation character cap and round up, then add the results. This is a minimum: keeping paragraph or scene boundaries can require more files, and regenerations may create additional versions.

How is this different from an ebook to audiobook converter?

Converters automate EPUB ingestion and chapter detection. TextSpeech gives you voice choice, Emotion tags, and the same Studio used for YouTube — at the cost of manual chapter pasting.

Can I also publish chapters on YouTube?

Yes — original script and edit still matter. See how to add AI voice over to YouTube.

MP3 or WAV?

TextSpeech’s documented export is MP3. Keep the original file for editing; converting a lossy MP3 to WAV does not restore lost detail. Choose the delivery format required by your destination, and confirm AI-narration eligibility separately.

Generate chapter one in Studio. Listen to minute five on a phone. If you would not commute with that host, change the voice before chapter two.

Keep reading

Try the line in Studio

Pick a voice, paste a script, and hear the take before you commit to an export.