Fish Audio Instant Voice Clone Best Practices (and ElevenLabs IVC)
How to record 10–90 seconds that actually clone well: clean speech, exact transcripts, levels, and what Lex gates for authors.
Instant clones are not the same job as a production audiobook sample. A 90-second phone take can sound like you on a demo line; a thin, noisy take will also sound like you: thin and noisy, forever. This is what we actually do when authors clone on Fish Audio through Lex, with a short ElevenLabs instant-clone (IVC) comparison when that is the tool.
Key numbers (Fish, Lex default path): hard floor 10 seconds of usable speech, target ~30 seconds, better at ~90 seconds, cap around 2 minutes for this lane. Room noise and clipping kill more clones than “not enough minutes.” For a full-book production sample (hours of clean narration), use our longer brief: how to record a voice sample under 2 hours.
What is an “instant” voice clone?
An instant clone builds a usable voice from a short reference instead of a long training set. On Fish you can either upload samples into a saved voice, or pass reference audio (plus transcript) on a single generation. On ElevenLabs, Instant Voice Cloning is a short-sample product separate from Professional Voice Cloning (tens of minutes to hours). At Lex, author self-narrate intake for quick demos lands on Fish; longer partner / author production samples use a different, stricter pipeline.
Instant is for “hear yourself on a page of your book now.” Long samples are for “ship a full title without the clone collapsing at hour six.” Do not mix those quality bars.
How long should the Fish sample be?
Fish’s own guidance for short reference clones is roughly 10–30 seconds of clean speech on the instant path, with about a minute or two helping when you train a persistent model. Below ~10s the clone often lacks prosody; past a few clean minutes of instant material you usually stop buying fidelity and start buying inconsistency.
What we gate in product: 10s minimum, 30s target, 90s better, about 120s max for the self-narrate flow. If you only have one good take, one continuous read of the guided script beats stitching three bad phone clips.
What makes Fish clones sound like you?
Signal quality first. Then transcript. Then performance consistency.
- One speaker, dry room. No music, HVAC roar, café wash, or bathroom slap. The model treats background as part of the voice.
- Distance and level hold still. Two fists from the mic, same seat, same energy. Riding the chair in and out becomes volume and timbre drift in every later take.
- Peaks under about −3 dB. Do not slam the sample into 0 dBFS. Clipped source reads as metallic or crunchy later. Loud enough to be clearly above noise is enough.
- Exact transcript when you can send one. Fish asks for matching text on the reference, including punctuation. Lex’s guided scripts exist so the words in the file match the words we send. Hand-edited freestyle is fine if you paste a true transcript; wrong words hurt prosody more people expect.
- Natural pace, not auctioneer, not ASMR. Read the way you want the book to sound. A forced “announcer” voice clones as an announcer.
- Variety inside one register. A question, a calmer sentence, a slightly warmer line: still the same person. Do not half-whisper and half-shout inside one clip.
Format is secondary. Clean mono speech at a normal speech sample rate beats a noisy 48 kHz stereo export. WAV or high-bitrate M4A/MP3 are both fine for this lane; compression artifacts only matter when the room already sounds cheap.
Fish can enhance noisy samples before training. Useful for a slightly hissy bedroom take. If you already have a dry, studio-clean capture, skip heavy “enhance / denoise” so the model does not learn the processor.
What should you actually read?
On Lex, use the on-screen guided script unless you have a reason not to. Those lines are written for continuous speech at a calm narration pace and we attach them as the Fish texts so alignment stays honest. Reading a page of your own book works if your mic take matches a transcript you can provide word-for-word.
Skip poems unless the book is poetry, skip songs, skip co-reads. Skip “I am testing, one two three” into the only file you upload. The clone will sound like a mic check.
How is ElevenLabs IVC different?
ElevenLabs Instant Voice Cloning wants slightly more speech than Fish’s shortest path, and it is even more allergic to inconsistency.
- Length: about 1–2 minutes of clean audio is the happy band. Past roughly 3 minutes of IVC material, gains flatten and bad takes can make the clone worse.
- Consistency over mileage. One excellent minute beats five mixed sources. Same mic, same room, same delivery, same distance.
- Levels: aim average speech around −23 to −18 dB RMS, true peak under about −3 dB. That matches what we already tell authors for longer samples.
- Codec: MP3 at 128 kbps or higher is usually enough for IVC; how you recorded matters more than WAV vs MP3.
If you are logging hours for a professional / long-form ElevenLabs clone, that is a different brief (think tens of minutes minimum, often hours). Use the production sample guide, not this instant list.
What ruins an instant clone in practice?
- Noise as identity. Fridge, fan, traffic, keyboard. The clone will bring them into every sentence.
- Room reverb. Wet bathrooms and hard kitchens bake space into the speaker.
- Auto-gain and aggressive “AI denoise.” Phones and consumer apps that pump or chirp train the model on the artifact.
- Multiple mics or days mashed together. Two different rooms read as two people.
- Acting extremes in a short sample. Instant clones have little average to fall back on; a cartoon villain line in a 20s clip can dominate the whole voice.
- Lying transcript (Fish). Mismatch between spoken audio and submitted text blurs rhythm and stress.
Quick checklist before you upload
- Quiet room (closet full of clothes is fine).
- One mic or phone, fixed distance, no headset processing if you can avoid it.
- At least 10s of real speech; 30–90s if you want a fair self-narrate demo.
- Natural narration energy; no music, no second speaker, no clipping.
- Transcript matches the take (use Lex’s script when you are in our flow).
- Listen once before upload: if you hear noise or slap, re-record. Do not “fix it in software” unless you know what you are doing.
Want more than a demo line? Start with a free first-chapter sample on Lex Audio, or see full production options at lex-books.com/audio. For full-length author clones, record the longer clean sample, not a single 30-second take stretched across a novel.