← Home
GLOSSARY

By Valmera Editorial · Published · Updated

Sidechain Ducking: Adjust Music Volume Under Speech

Sidechain ducking reduces one audio signal in response to the level of a separate reference signal, called the key. For example, a compressor can lower music when its dialogue reference becomes louder. Its response depends on the detector and settings; it does not recognize words.

If music covers your words, start with the music level. Then decide whether a changing level would help. A quiet bed, a timed reduction and a sidechain compressor solve different mixing needs.

Adjust music under speech in four steps

  1. 1
    Identify the voice and competing music
    Select the intended edit and name the original audio, added music and any uploaded narration separately. If music and speech are already mixed into one recording, lowering that whole recording lowers both; an added-layer duck is a different operation.
  2. 2
    Set a useful starting balance
    Listen to the quietest important words and lower the added music until they are clear. A steady, quiet bed can be appropriate. Decide whether you also want music to rise in gaps; that determines whether ducking adds value.
  3. 3
    Choose and inspect the ducking behavior
    Identify which signal or timed window should lower which layer. In Valmera, request music ducking on or off and smooth or step mode explicitly. Review the actual setting and preview; changing mode alone has no effect when ducking is off.
  4. 4
    Compare the difficult passages
    Check quiet speech, loud delivery, pauses, ambience, transitions and the ending. Change one relevant control at a time. Recheck after cuts or audio changes, then export the approved version in Studio and play the downloaded MP4.

For example, music chosen against a loud speaker can cover their quieter words. Lowering the bed for those quiet words also lowers it in pauses; that may be acceptable, or you may prefer controlled movement. Neither choice guarantees intelligibility.

Follow the reference signal

1 · Key signal

The reference whose level is measured.

2 · Detector and response

The settings determine gain reduction.

3 · Music gain stage

The reduction is applied to the music.

Conceptual signal flow. The key controls the reduction; the music is the signal being reduced.

A louder sound in the key can cause more ducking, whether it is a word, a door or a crowd. Turning down an independently routed music layer changes its level without changing that separate key. The FFmpeg sidechain compressor reference describes this two-input arrangement.

What threshold, ratio, attack and release mean

Threshold
The reference level around which compression begins. A soft knee makes the transition gradual; it is not always a sharp on/off boundary.
Ratio
How strongly level above threshold affects gain reduction. A ratio is not a fixed number of decibels removed.
Attack
How quickly the reduction responds as the key grows. A slower response can leave more of the initial sound un-ducked.
Release
How quickly gain recovers as the key subsides. A short recovery can produce repeated rises in gaps; a long one can keep music down into a pause.

These are response controls, not promises that a fade finishes exactly at the stated millisecond. Knee, detection, routing and the changing source also matter. Read Ableton's compressor reference and Apple's knee and sidechain guidance for their implementations. Listen to the actual passage instead of assuming one setting suits every voice.

Explore ducking depth with a simple model

Suppose the settled key is −18 dBFS and the threshold is −30 dBFS: the key is 12 dB above threshold. In a hard-knee model at 4:1, 12 ÷ 4 = 3 dB of that excess remains, so the music is attenuated by 12 − 3 = 9 dB. At 12:1, the same model gives 11 dB. A quieter key produces less reduction at the same settings.

An illustrative, settled hard-knee model with no makeup gain. It ignores attack, release, knee smoothing and detector behavior. These sliders explain the arithmetic; they do not change Valmera settings or predict its rendered mix.

Explore the key, threshold and ratio

Modeled gain reduction: 9.0 dB

12.0 dB above threshold × (1 − 1/4) = 9.0 dB of attenuation. 35.5% of the music's pre-duck amplitude remains; this is not a perceived-loudness percentage.

Try moving the key to or below the threshold, then try a ratio of 1:1. Both produce zero reduction in this simplified model. This arithmetic excludes real detector and timing behavior; compare it with the measured example below without treating the numbers as interchangeable.

Hear the difference in a controlled comparison

The published audio comparison uses a synthesized chord and two higher test-tone windows. With smooth ducking enabled, the measured background frequency band drops by about 14 dB during those windows and recovers afterwards. The same source, gain, fades and encoding settings are used in both variants.

These are approximately 12-second examples generated locally through the production-revision renderer with supplied edit settings. They demonstrate level-triggered reduction, with no human voice, Studio agent request or account export. Measurements and reproduction inputs and instructions are available.

Choose a steady bed, a timed duck or a sidechain

Different ways to control music
ApproachWhat drives itWhat to check
Steady gainOne chosen levelQuiet speech stays clear; the bed also stays quiet in gaps.
Timed duck or automationAuthored or detected windowsWindow placement, fade shape and whether later edits update it.
Sidechain compressionThe reference signal's levelQuiet words, non-speech triggers and distracting level movement.

Automatic ducking does not always mean a sidechain compressor. For example, Premiere's documented Auto Ducking workflow generates editable gain keyframes. Regenerating them overwrites manual changes. Identify the method in the editor you use before assuming how it follows revisions.

What Valmera exposes and what its defaults mean

Request the music layer's gain, ducking on or off, and smooth or step mode. Mode alone does not enable ducking. Smooth music ducking uses program audio as its reference; the legacy step mode applies a 12 dB reduction over mapped transcript speech windows. Check the current item and preview after changing either.

New music defaults depend on indexed source speech surviving in the selected edited window. At least one second selects a −18 dB bed with smooth ducking; less than one second selects −4 dB lead music with ducking off. Explicit gain and duck arguments override their respective defaults. Existing items keep their stored settings.

Uploaded narration has a separate path. Its ducking option lowers original program audio by 12 dB during the voiceover window. The narration is mixed later, so it does not automatically drive smooth music ducking. Reducing or muting the original program can also change that reference. Balance music against the voiceover explicitly.

Renderer settings and limits

The checked renderer uses one sidechain compressor per smooth-ducked music item, with threshold 0.03, ratio 12:1, attack 180 ms and release 550 ms. A linear threshold of 0.03 corresponds to approximately −30.5 dBFS using 20 × log₁₀(0.03). That is a threshold conversion, not a measured speech level or a guaranteed amount of attenuation.

The music tools do not expose individual threshold, ratio, attack or release controls. This calculator does not add those controls to Studio. Gain, mode and enablement remain separate requests; preview and final rendering determine the actual sound.

In the current edit, lower [music item] by 3 dB from its existing gain and enable smooth ducking. Show the preview around [quiet spoken passage] and [loud background sound]. Keep the track timing and repeat behavior unchanged.

This is an illustrative revision brief, not a tested agent response. Account creation and uploads are free; indexing and agent editing require a subscription and processing credits. Use Studio for the final MP4.

If the music still gets in the way

Quiet words remain masked

Check the starting music level and key strength. A quiet key can produce less reduction, so increasing the ratio may not solve the balance. Try a lower bed, an appropriate timed duck or a less competing track.

Music dips on room noise

Identify the reference. If program audio contains that noise, smooth ducking can react to it. Compare a steady level or the available step mode, and inspect the transcript windows before relying on them.

Music repeatedly rises and falls

Check whether key changes and recovery timing explain the movement. In editors with compressor controls, compare timing settings one at a time. In Valmera, compare gain and available modes; individual compressor timings are not exposed.

Changing the mode seems to do nothing

Confirm ducking is enabled on the intended item, that music and reference overlap, and that you are playing the newly rendered edit. An old downloaded file cannot reflect the revision.

Music already mixed into the original recording needs a different approach from lowering an added layer. Valmera can attempt source music/voice separation subject to service, source-length and media-readiness constraints; residue or artifacts can remain. See the audio reference.

Frequently Asked Questions

Identify the music layer and start by lowering its gain while listening to the quietest important speech. If music should rise during pauses, try ducking against the appropriate signal or speech windows. Compare the actual preview and downloaded final; one music percentage does not fit every recording.
A level-driven compressor responds to its key signal, not the meaning or presence of words. Noise, laughter or existing music can trigger it if they are in that reference. Quiet speech may cause little reduction. A transcript-timed duck uses different evidence.
No. A steady, low music bed can work well when the recording and arrangement allow it. Ducking is useful when you want the balance to change with another signal or a defined window. Check whether the movement helps the video or adds distracting dips.
There is no universal depth. The useful amount depends on the source levels, track, speech, timing and intended result. The calculator illustrates how key level, threshold and ratio interact under simplified assumptions; it does not recommend a target for your recording.
The checked music-editing tools expose gain, ducking enablement and smooth or step mode, but do not expose individual compressor parameters. Smooth rendering uses threshold 0.03, ratio 12:1, attack 180 ms and release 550 ms. These settings do not establish a fixed ducking depth or exact completion time.
When gain and ducking are omitted, the add-music tool checks indexed source speech surviving in the selected edited window. With at least one second it defaults to a −18 dB smooth-ducked bed; with less than one second it defaults to −4 dB lead music with ducking off. Explicit gain and duck arguments override the respective defaults.
Uploaded voiceover is mixed after the program reference used by smooth music ducking. Its own ducking control lowers original program audio by 12 dB during its active window; it does not automatically make that narration the music sidechain. Balance music against uploaded narration explicitly.
Overall loudness treatment processes the combined mix. It does not independently choose a better music-to-voice balance. Adjust the competing layers first, then apply any required overall treatment and check the delivered result.

Make space for the words that matter

Describe the balance, compare the preview and review the final file. Editing requires a subscription and processing credits.

Open Studio →
See pricing →

Related Articles

Add background music
Track timing, mixing briefs and a reproducible audio comparison.
Audio controls
Original audio, music, voiceover and final loudness treatment.
LUFS and loudness
Understand overall level separately from layer balance.
Recorded narration
Place a voiceover and review the separate audio layers.