Sidechain Ducking: Adjust Music Volume Under Speech
Sidechain ducking reduces one audio signal in response to the level of a separate reference signal, called the key. For example, a compressor can lower music when its dialogue reference becomes louder. Its response depends on the detector and settings; it does not recognize words.
If music covers your words, start with the music level. Then decide whether a changing level would help. A quiet bed, a timed reduction and a sidechain compressor solve different mixing needs.
Adjust music under speech in four steps
- 1Identify the voice and competing musicSelect the intended edit and name the original audio, added music and any uploaded narration separately. If music and speech are already mixed into one recording, lowering that whole recording lowers both; an added-layer duck is a different operation.
- 2Set a useful starting balanceListen to the quietest important words and lower the added music until they are clear. A steady, quiet bed can be appropriate. Decide whether you also want music to rise in gaps; that determines whether ducking adds value.
- 3Choose and inspect the ducking behaviorIdentify which signal or timed window should lower which layer. In Valmera, request music ducking on or off and smooth or step mode explicitly. Review the actual setting and preview; changing mode alone has no effect when ducking is off.
- 4Compare the difficult passagesCheck quiet speech, loud delivery, pauses, ambience, transitions and the ending. Change one relevant control at a time. Recheck after cuts or audio changes, then export the approved version in Studio and play the downloaded MP4.
For example, music chosen against a loud speaker can cover their quieter words. Lowering the bed for those quiet words also lowers it in pauses; that may be acceptable, or you may prefer controlled movement. Neither choice guarantees intelligibility.
Follow the reference signal
The reference whose level is measured.
The settings determine gain reduction.
The reduction is applied to the music.
A louder sound in the key can cause more ducking, whether it is a word, a door or a crowd. Turning down an independently routed music layer changes its level without changing that separate key. The FFmpeg sidechain compressor reference describes this two-input arrangement.
What threshold, ratio, attack and release mean
- Threshold
- The reference level around which compression begins. A soft knee makes the transition gradual; it is not always a sharp on/off boundary.
- Ratio
- How strongly level above threshold affects gain reduction. A ratio is not a fixed number of decibels removed.
- Attack
- How quickly the reduction responds as the key grows. A slower response can leave more of the initial sound un-ducked.
- Release
- How quickly gain recovers as the key subsides. A short recovery can produce repeated rises in gaps; a long one can keep music down into a pause.
These are response controls, not promises that a fade finishes exactly at the stated millisecond. Knee, detection, routing and the changing source also matter. Read Ableton's compressor reference and Apple's knee and sidechain guidance for their implementations. Listen to the actual passage instead of assuming one setting suits every voice.
Explore ducking depth with a simple model
Suppose the settled key is −18 dBFS and the threshold is −30 dBFS: the key is 12 dB above threshold. In a hard-knee model at 4:1, 12 ÷ 4 = 3 dB of that excess remains, so the music is attenuated by 12 − 3 = 9 dB. At 12:1, the same model gives 11 dB. A quieter key produces less reduction at the same settings.
An illustrative, settled hard-knee model with no makeup gain. It ignores attack, release, knee smoothing and detector behavior. These sliders explain the arithmetic; they do not change Valmera settings or predict its rendered mix.
Modeled gain reduction:
12.0 dB above threshold × (1 − 1/4) = 9.0 dB of attenuation. 35.5% of the music's pre-duck amplitude remains; this is not a perceived-loudness percentage.
Try moving the key to or below the threshold, then try a ratio of 1:1. Both produce zero reduction in this simplified model. This arithmetic excludes real detector and timing behavior; compare it with the measured example below without treating the numbers as interchangeable.
Hear the difference in a controlled comparison
The published audio comparison uses a synthesized chord and two higher test-tone windows. With smooth ducking enabled, the measured background frequency band drops by about 14 dB during those windows and recovers afterwards. The same source, gain, fades and encoding settings are used in both variants.
These are approximately 12-second examples generated locally through the production-revision renderer with supplied edit settings. They demonstrate level-triggered reduction, with no human voice, Studio agent request or account export. Measurements and reproduction inputs and instructions are available.
Choose a steady bed, a timed duck or a sidechain
| Approach | What drives it | What to check |
|---|---|---|
| Steady gain | One chosen level | Quiet speech stays clear; the bed also stays quiet in gaps. |
| Timed duck or automation | Authored or detected windows | Window placement, fade shape and whether later edits update it. |
| Sidechain compression | The reference signal's level | Quiet words, non-speech triggers and distracting level movement. |
Automatic ducking does not always mean a sidechain compressor. For example, Premiere's documented Auto Ducking workflow generates editable gain keyframes. Regenerating them overwrites manual changes. Identify the method in the editor you use before assuming how it follows revisions.
What Valmera exposes and what its defaults mean
Request the music layer's gain, ducking on or off, and smooth or step mode. Mode alone does not enable ducking. Smooth music ducking uses program audio as its reference; the legacy step mode applies a 12 dB reduction over mapped transcript speech windows. Check the current item and preview after changing either.
New music defaults depend on indexed source speech surviving in the selected edited window. At least one second selects a −18 dB bed with smooth ducking; less than one second selects −4 dB lead music with ducking off. Explicit gain and duck arguments override their respective defaults. Existing items keep their stored settings.
Uploaded narration has a separate path. Its ducking option lowers original program audio by 12 dB during the voiceover window. The narration is mixed later, so it does not automatically drive smooth music ducking. Reducing or muting the original program can also change that reference. Balance music against the voiceover explicitly.
Renderer settings and limits
The checked renderer uses one sidechain compressor per smooth-ducked music item, with threshold 0.03, ratio 12:1, attack 180 ms and release 550 ms. A linear threshold of 0.03 corresponds to approximately −30.5 dBFS using 20 × log₁₀(0.03). That is a threshold conversion, not a measured speech level or a guaranteed amount of attenuation.
The music tools do not expose individual threshold, ratio, attack or release controls. This calculator does not add those controls to Studio. Gain, mode and enablement remain separate requests; preview and final rendering determine the actual sound.
In the current edit, lower [music item] by 3 dB from its existing gain and enable smooth ducking. Show the preview around [quiet spoken passage] and [loud background sound]. Keep the track timing and repeat behavior unchanged.
This is an illustrative revision brief, not a tested agent response. Account creation and uploads are free; indexing and agent editing require a subscription and processing credits. Use Studio for the final MP4.
If the music still gets in the way
Quiet words remain masked
Check the starting music level and key strength. A quiet key can produce less reduction, so increasing the ratio may not solve the balance. Try a lower bed, an appropriate timed duck or a less competing track.
Music dips on room noise
Identify the reference. If program audio contains that noise, smooth ducking can react to it. Compare a steady level or the available step mode, and inspect the transcript windows before relying on them.
Music repeatedly rises and falls
Check whether key changes and recovery timing explain the movement. In editors with compressor controls, compare timing settings one at a time. In Valmera, compare gain and available modes; individual compressor timings are not exposed.
Changing the mode seems to do nothing
Confirm ducking is enabled on the intended item, that music and reference overlap, and that you are playing the newly rendered edit. An old downloaded file cannot reflect the revision.
Music already mixed into the original recording needs a different approach from lowering an added layer. Valmera can attempt source music/voice separation subject to service, source-length and media-readiness constraints; residue or artifacts can remain. See the audio reference.
Frequently Asked Questions
Make space for the words that matter
Describe the balance, compare the preview and review the final file. Editing requires a subscription and processing credits.
Open Studio →