Lex

Browse

GenresShelvesPremiumBlog

Company

AboutJobsPartnersSell on LexAffiliates

Audio

Lex StudioVoicesFirst chapter free

Resources

DocsInvite FriendsFAQ

Legal

Terms of ServicePrivacy Policygeneral@lex-books.com(215) 703-8277

© 2026 LexBooks, Inc. All rights reserved.

Blog
July 22, 2026·9 min read

How to Record a Voice Sample for Your AI Audiobook Clone (Under 2 Hours)

Gear, room, levels, and script tips so Lex can clone your voice from under two hours of clean narration. From real author sessions.

ai audiobooksvoice cloningindie authorsaudiobook narration
How to Record a Voice Sample for Your AI Audiobook Clone (Under 2 Hours)

In this article

  1. How much audio do you actually need?
  2. What microphone and interface should you use?
  3. What room and setup work at home?
  4. What recording settings should you use?
  5. What should you read into the mic?
  6. What ruins a voice clone?
  7. How should you run the session?
  8. How do you deliver the sample to Lex?
  9. What happens after we have the sample?

You do not need sixty hours in a booth to hear your book in your own voice. At Lex, authors and partner narrators record up to two hours of clean speech and we extend that voice across the full title. The clone is only as good as the sample. This is the recording brief we send before every author-voice session.

Key numbers: aim for up to 120 minutes of usable speech (hard floor ~30 minutes; target range 1–2 hours). Record WAV at 44.1 or 48 kHz, 24-bit. Keep average level around −23 to −18 dB RMS with true peaks under −3 dB. One mic, one room, one narration style.

How much audio do you actually need?

Total spoken runtime matters more than the number of files. Under half an hour, the clone gets thin on long chapters. Our ask for author and partner sessions is up to 120 minutes of usable speech (the 1–2 hour window). That is enough for a production-grade clone without a full-book booth day. Past two hours of the same clean take, returns flatten and you mostly add fatigue, so stop there.

That is the same model we use with partner narrators like Bob Neufeld: a 1–2 hour sample per book, then Lex extends the voice across the rest. Authors get the same deal for memoir, essays, and any book that should sound like you.

What microphone and interface should you use?

Pick by room first, then by budget. A clear mic in a quiet closet beats a famous mic in a noisy kitchen. Prefer an XLR mic into a dedicated interface over a laptop mic or Bluetooth headset. Amazon prices move; figures below are approximate totals as of mid-2026.

Rule: quiet closet or treated corner → condenser. Untreated bedroom → dynamic (rejects the room).

Desperate cheaper (~$30): only if the room is dead

FIFINE USB condenser (~$30) or similar Maono/Tonor-class plug-and-play mics. Possible for a usable clone only in a quiet closet with every auto-gain, RGB gimmick, and “AI noise reduction” feature off, recording flat WAV. Expect more rejects: hiss, harsh highs, USB bus noise, and processing that bakes into the voice. If you can stretch to ~$80, skip this tier.

Budget (~$80): cheapest we reliably accept

Samson Q2U (~$80). Dynamic USB-C/XLR handheld. Plugs straight into a computer, no interface required. Rejects a lively room better than most cheap USB condensers. Includes a stand and windscreen. This is the lowest kit we recommend for a production clone sample.

Recommended (~$250–$360): what most authors should buy

Quiet room / closet: Audio-Technica AT2020 (~$120) or the RØDE NT1 Signature Series kit (~$140, includes shock mount, pop filter, and cable), plus a Focusrite Scarlett Solo (4th Gen) (~$160). Condensers capture more detail; they also hear your fridge, so only go this route if the room is dead.

Untreated bedroom: RØDE PodMic (~$80) or RØDE Procaster (~$200), plus the same Scarlett Solo. Both are broadcast dynamics built for speech. PodMic covers most of the job at the lower price; Procaster is the mid-high step up (fuller presence, still without SM7B money or gain headache). The classic “podcaster mic” people mean when they say SM7B is fine too, but it belongs in the high-end tier below, not here.

High end (~$700–$900): classic broadcast chain

Shure SM7B (~$440) + Cloudlifter CL-1 (~$100) + Scarlett Solo, or skip the Cloudlifter and buy the Shure SM7dB (~$550, built-in preamp) + Solo. Excellent room rejection and long-session comfort. Not required for a good Lex clone; buy it if you already want a permanent booth mic.

USB condenser option if you refuse XLR: Audio-Technica AT2020USB+ (~$200). Note it costs more than the Q2U and more than the AT2020 body alone.

Skip: AirPods, phone voice-memo defaults, Blue Yeti (stereo modes and processing bake into the clone), conference headsets, and anything with aggressive noise suppression or auto-gain.

Sit about two fists from the mic. Stay on-axis. Use a pop filter on condensers (NT1 Signature kit already includes one; Q2U ships with a windscreen). Do not ride the chair closer and farther between paragraphs; distance changes are baked into the clone as tone shifts.

What room and setup work at home?

The model clones everything it hears: HVAC hum, fridge buzz, street traffic, and slap echo. Treat the room before you treat the file.

  • Smallest quiet room you have. Closets full of clothes work surprisingly well.
  • Hang a thick duvet or moving blanket behind and beside the mic to kill first reflections.
  • Turn off fans, AC, dishwashers, and notifications. Soft shoes or socks beat hard soles on wood floors.
  • Record at the same time of day for every session if you split the sample across days. Same room noise floor, same voice.

If you can book a cheap hour in a podcast booth or coworking recording pod, do it. Consistency beats luxury: one good room for the whole sample is better than mixing a studio hour with a bedroom hour.

What recording settings should you use?

Capture clean source. Do not "master" the sample like a finished audiobook.

  • Format: WAV (or AIFF). Not MP3, not M4A, not Voice Memos compressed exports.
  • Sample rate / bit depth: 44.1 kHz or 48 kHz, 24-bit.
  • Channels: mono is ideal; stereo is fine if both channels are identical (no wide "podcast" dual-mic tricks).
  • Levels: average speech around −23 to −18 dB RMS, true peaks under −3 dB. Loud enough to be clear, never clipping.
  • Processing: flat. No music bed, no reverb plugin, no heavy compression, no "studio enhancer," no noise reduction that pumps or chirps. Light high-pass (around 80 Hz) to kill rumble is fine. We would rather remove noise carefully than inherit a bad denoiser.

Free tools that work: Audacity, GarageBand (export WAV), Reaper's eval build. Record, trim long silences and throat clears, export. Stop there.

What should you read into the mic?

Read the way you want the finished audiobook to sound. The clone copies cadence, pause length, breathiness, and energy. If you whisper half the sample and announce the other half, the book will wander between those poles.

  • Best material: chapters from this book, in your natural narration voice. Include narrative prose and a few stretches of dialogue so the model hears both registers inside one consistent read.
  • One style: audiobook narration, not a YouTube explainer voice, not a stage performance, not a whispered ASMR take. Pick the delivery listeners will live with for eight hours.
  • Variety inside that style: questions, quiet interior moments, a higher-stakes paragraph. Same person, same booth energy, wider emotional range.
  • Names and invented words: say them the way you want them forever. We also run a pronunciation pass in production, but the sample is the ground truth for your mouth shapes.

Scripts that are not your book still work if the tone matches (another chapter of yours, a public-domain passage in the same register). Avoid reading marketing copy, poems (unless the book is poetry), or songs. Spoken voice only.

What ruins a voice clone?

We have thrown away sessions for the same handful of problems. Fix them in the booth, not in email.

  • Background noise or music. The clone will hum, tick, or sing along forever.
  • Room echo. Reverb becomes a permanent "wet" narrator. Dead rooms clone clean; bathrooms do not.
  • Multiple speakers. Partner reads, kids, podcast guests. One voice only.
  • Inconsistent mic or room. Mixing a USB mic day with an SM7B day teaches the model two different people.
  • Over-acted character voices. In every long-form probe we have run, clones from clean, consistent studio narration stay stable across chapters; "folksy grandpa" and noir caricature invite the model to re-improvise every take. Save character work for light color inside a steady narrator, or cast separate voices later.
  • Heavy editing artifacts. Chopped breaths, robotic denoise, brickwall limiting. The model learns the artifact as timbre.
  • Phone recordings with auto-gain. Levels pump; the clone pumps.

Quality beats quantity. Two clean hours outperform three hours of mixed phone takes every time. That matches what we see when casting catalog voices: stable studio sources hold register; noisy or theatrical sources drift.

How should you run the session?

Plan for more clock time than spoken time. A full 120-minute usable sample usually means three to four hours including setup, water, and retakes. Hitting an hour of keepers is fine if that is all you have that day; fill toward two hours when you can.

  • Warm up for five minutes (read aloud, not silent). Discard the warm-up.
  • Work in 20–25 minute blocks with short breaks. Fatigue thins the top end of the voice; the clone will sound tired in hour ten of the book if hour two of your sample already is.
  • Hydrate. Avoid dairy and heavy chips right before. Sip, do not gulp mid-sentence.
  • Mark retakes out loud ("retake") or leave a clear gap so you (or we) can cut them. Do not stitch bad takes into the keep file.
  • Split long sessions into files of roughly 20–30 minutes each. Easier to upload, easier to QA, same total runtime.

How do you deliver the sample to Lex?

Send lossless files, labeled, with a one-line note on mic and room.

  • Files: WAV, 44.1 or 48 kHz, 24-bit, mono preferred. Name them yourname_book_sample_01.wav, _02, and so on.
  • Transfer: a single cloud folder (Google Drive, Dropbox, WeTransfer). No Facebook/Messenger compressions.
  • Notes: mic + interface, room (closet / treated office / studio), and anything we should know (accent goals, words you always mis-say, chapters you want emphasized).
  • Consent: author-voice clones are consent-verified. You will confirm the voice is yours before we train. Do not send someone else's audiobook and ask us to "sound like them."

If you already have a clean podcast or prior audiobook in exactly the narration style you want, we can often train from that instead of a new session. Same rules apply: one speaker, no music, consistent chain.

What happens after we have the sample?

We train a dedicated voice model from your files, then produce the book with the same continuity stack we use on every commission: request stitching, take selection at chapter joins, level matching, and human QA for seams and pronunciation. You hear a first-chapter proof before we burn the rest of the title. If a line needs a different read, we micro-retake and splice rather than regenerating whole chapters.

Want to hear the production style first without recording anything? Send your manuscript chapter and we return finished audio on a cast voice within 24 hours, free, on the first-chapter page. When you are ready for your voice, use this brief, record under two hours once, and we take it from there.


Lex produces scored AI audiobooks for indie authors and classic literature: narration, music, ambience, and sound effects, with optional author-voice or partner-narrator clones. See packages on /audio and author terms on /authors.

More articles

How Much Does Audiobook Production Actually Cost in 2026?

How Much Does Audiobook Production Actually Cost in 2026?

Real numbers for audiobook production in 2026: per-finished-hour narrator rates, editing and mastering fees, the royalty-share trap, and what AI production changes.

ACX Royalty Share vs. Paying Upfront vs. AI Narration: The Real Math

ACX Royalty Share vs. Paying Upfront vs. AI Narration: The Real Math

A working comparison of the three ways indie authors get audiobooks made in 2026 — with the 7-year exclusivity math ACX doesn’t put on the pricing page.

Selling Ebooks at 90% Royalty: How Authors Keep More Without Leaving Amazon

Selling Ebooks at 90% Royalty: How Authors Keep More Without Leaving Amazon

KDP pays 35–70% and charges delivery fees. Lex pays 90% with no exclusivity — so you can stay on Amazon and still make more per sale everywhere else.