Scoring an AI Audiobook: Music, Ambience, and SFX That Don't Drown the Narrator
How Lex scores AI audiobooks like films: anchored music beds, looped ambience, word-synced sound effects, and a stem mix where silence is a choice.
A good narration doesn't need help. But a scored audiobook, one with a film-style music score, environmental ambience and sound effects, can do something plain narration can't: put you inside the scene. You hear the river the characters are walking beside. A low drone arrives two seconds before the text tells you why you're tense.
Scoring is also easy to ruin. Most attempts fail the same way, with wall-to-wall music that turns a book into a podcast intro. After scoring everything from Homer to original historical fiction at Lex, this is the production approach we've settled on.
Silence is a real option
The narration is the product. The score is seasoning. In our chapters, music is present for well under half the runtime. A cue enters for a beat, a confrontation, a vision, a departure, makes its point and leaves. Long stretches run on ambience alone or on nothing. Intimate dialogue usually plays completely dry, because music under a scene like that tells the listener how to feel about words that are already doing the job.
And a cue plays once. If the scene outlasts its track, the tail is silence, not a loop. Looping a melody under twelve consecutive paragraphs is the fastest way to send a listener back to the plain version. Ambience can loop, since wind and river are texture rather than melody, but each musical moment gets its own track. Leitmotif, not playlist.
The voice always wins
Music sits roughly 20 dB under the narration. Ambience floats just above audibility. One-shot effects come in a touch louder, then get out of the way. If a moment genuinely calls for prominent, swelling music, a human signs off on that as an editorial decision. It is never a default. At these levels the score works the way film scores work: you feel it more than you hear it.
Anchor cues to words, not timecodes
Our narration pipeline emits word-level timestamps, and the entire score hangs off them. A scene's music starts where a specific phrase begins in the text, not at minute 7:05. The sound of clinking beads lands on the word where the beads clink.
This matters because AI narration gets re-rendered constantly. A better voice, a fixed line, new pacing. When the narration changes, every anchor re-resolves against the new timestamps and the score re-conforms in seconds. Hard-coded timecodes would turn every narration fix into a manual re-mix.
Mix in stems, master once
The mix is built as four full-length, time-aligned stems: voice, music, ambience, effects. Gains and fades are baked into each stem. Loudness normalization is applied only to the combined master, never per stem, so the four stems always re-sum to exactly the same mix.
The payoff is that every future edit touches one stem. A narration fix doesn't re-render the music. A music re-balance doesn't touch the voice. The stems drop straight onto four DAW tracks for human QA, and the master is bounced to lossless WAV first, with the delivery MP3 derived from it, so the editorial master never passes through a second lossy encode.
The workflow
- Narrate the chapter with word-level timestamps (and deal with the seam problem).
- Spot the chapter like a film composer: where music enters, what mood, where it yields to silence.
- Source licensed tracks, beds and effects. Log every asset with its license, and report usage to the rights holder at mix time.
- Anchor every cue to a phrase in the text.
- Mix: resolve anchors against the timestamps, build the four stems, master once.
Once the pipeline exists, the marginal cost of scoring a chapter is minutes of compute. That's the real promise of AI audiobooks. Not replacing narrators, but making the cinematic version of every book economically possible.
Hear the finished results, from Homer with a leitmotif score to gothic London under dockside piano, on the Lex audio showcase or in the free app.