← Home
TOOL

By Valmera Editorial · Published · Updated

Add Background Music to Video with AI: Timing, Volume and Ducking

Bring a music track or ask Valmera to find candidates. Tell the editing agent where the music belongs, how prominent it should be and how it should enter and end. Review the mix against your footage before creating the final MP4.

Account creation and uploads are free; indexing and agent editing require a subscription and processing credits. Open Studio · See current plans

Hear what smooth ducking changes

This 12-second comparison uses a steady synthesized chord and a higher test tone at 00:02–00:04 and 00:06–00:08. In A, the chord stays at its set level. In B, the chord dips while the higher tone plays and recovers afterwards. Play one at a time using the same device volume.

A · Ducking off

Download A

B · Smooth ducking on

Download B

We generated both variants locally with the production-revision Valmera renderer and supplied edit settings. There is no human voice or customer footage in this test. It demonstrates level-triggered ducking, not speech recognition, song selection, prompt success or a Studio export. The synthetic inputs were prepared for this comparison; Valmera does not offer AI music generation.

Inputs, measurements and reproduction

Both variants use the same source, music gain, fades and encoding; only ducking enablement and mode differ. Mastering is off. In the measured 3.0–3.8 and 7.0–7.8 second windows, the background frequency band is about 14 dB lower with ducking. Before the tones and after recovery, it is within 0.01 dB of the version without ducking. This is a signal comparison, not a loudness or subjective-quality score.

Source with test tones · Synthetic chord input · Measured evidence · Reproduction instructions

The A edit, B edit and generated A/B filtergraphs are downloadable. Read the FFmpeg compressor reference for the underlying signal control.

Give the agent a track, a window and a mixing goal

Use [approved track filename] from preview 00:05 to 00:25 in the current edit. Begin at 00:12 inside the music file. Keep the original speech clear, use smooth music ducking, add a one-second entrance fade and a two-second exit fade, and show the preview. Tell me if the track will run out or repeat.

This is an illustrative brief to adapt. For a montage without speech, specify that music should be the lead sound and state whether original location audio should remain. For uploaded narration, name its file separately and request the balance you need under that voiceover.

Add and review music in five steps

  1. 1
    Choose footage and an approved track
    Open the project and intended edit version. Upload a supported audio file, use an accessible link, or ask for music candidates. Check the actual track, its source and the terms for your intended use before placing it.
  2. 2
    Specify the video window and track offset
    Name when the music should start and end using the edited preview clock. Separately state where playback should begin inside the music file. Say whether the track should repeat and how its entrance and ending should fade.
  3. 3
    Set the balance and ducking request
    Describe whether music is a quiet bed under speech or the main sound of a montage. Identify the original audio, music and any uploaded narration separately. Ask for the required levels and ducking, then inspect the settings and preview.
  4. 4
    Listen to speech, gaps and transitions
    Check the quietest spoken passage, louder moments, pauses, track entrance, repeats and ending. Listen for competing music, unwanted dips, clipping, a sudden stop or a duplicated audio file. Recheck after cuts, inserts or speed changes.
  5. 5
    Export and check the delivered mix
    Create the approved final MP4 in Studio once the required original media is ready. Listen to that downloaded file through its ending. Check any requested loudness treatment and branding; a saved setting or old preview is not the delivered result.

Separate video start time from the music-file offset

The music window uses the edited preview clock. The offset skips into the track itself. In the example brief, music starts five seconds into the video, but begins twelve seconds into the music file. These are two different controls.

The add-music tool defaults to the current program window when no start or end is given. A later structural edit can change which picture or speech that window covers. Check music timing after cuts, inserts and speed changes; an old downloaded MP4 will not update.

Looping can fill a window longer than the remaining track; without it, the tail can be silent. Listen to repeat joins and fade placement. A repeat is not a guarantee of a seamless musical edit. Beat or energy analysis can help locate a cue, but review the actual reveal and preserve meaningful speech around any suggested cut.

Track gain, ducking and final loudness do different jobs

Gain sets a layer's level. Ducking reduces a layer in response to another signal or a defined window. Mastering processes the completed mix. Turning on mastering does not fix music that already overwhelms the voice.

Linear amplitude and relative gain — illustrative conversions
AmplitudeGain
100%0 dB
50%About −6 dB
20%About −14 dB
10%−20 dB

These values are amplitude ratios, not perceived-loudness percentages or LUFS targets. Ask for an explicit gain in dB when precision matters. A track already at −18 dB set to −21 dB is reduced by 3 dB; setting it to −3 dB would make it much louder.

Defaults depend on surviving indexed source speech. A quiet bed and lead music can receive different initial levels and ducking settings. Smooth music ducking listens to program audio, so loud ambience can trigger it. Uploaded voiceover is a separate layer: its ducking control lowers original program audio during its window, and it does not automatically become the smooth music sidechain reference. Listen and balance the music against narration explicitly.

Valmera's optional social mastering targets −14 LUFS with −2 dBTP headroom and an output limiter. This is a processing preset, not a universal YouTube, TikTok or other platform requirement. Check the intended destination and measure the delivered file when a loudness specification matters. See audio controls and the normalization reference.

Check the track source and music already in the recording

Upload MP3, WAV, M4A, AAC or OGG audio up to 50 MB per track, or use a supported accessible media link. Search results and source-supplied license labels are leads to check, not publication clearance. Keep the original source, relevant terms and required attribution with the project. For Creative Commons music, read the official license FAQ, including its guidance on music synchronized with video.

An added music layer is separate from music already mixed into the original recording. Valmera can attempt source music/voice separation and rebalance speech or vocals against the remaining soundtrack, subject to service, readiness and source-length constraints. Dense mixes may retain unwanted sound or develop artifacts. Listen before relying on removal. Muting the entire original track would also remove the recorded speech.

Listen for the problem before asking for a revision

Speech is hard to understand

Name the passage and the competing music layer. Reduce that layer, inspect ducking and listen to the quietest words. Loudness mastering cannot repair the balance between layers.

Music dips during a crowd or wave sound

Smooth ducking reacts to program sound level. Consider a lower steady music level or different ducking arrangement, then compare the affected passage.

The track starts at the wrong part of the song

Check the preview start separately from the music-file offset. Confirm the intended file and identify the cue inside it.

The ending goes silent or stops abruptly

Check the item end, remaining track duration, looping and fade-out. Listen through any final card and the complete downloaded ending.

The same audio sounds doubled

Check whether the same file was added both as music and as voiceover. Remove the unintended duplicate from the selected edit and review the new preview.

Finish with the Studio export checks. A successfully saved edit or processing setting does not establish that the final mix sounds right.

Frequently Asked Questions

Open your footage and add an approved track, or ask for music candidates. Name its start and end in the edited preview, the starting point inside the track, and the desired level, fades and repeat behavior. Ask for a preview, check the complete mix and export the approved version in Studio.
Account creation and uploads are free. Indexing and agent editing require a subscription and processing credits. Music editing usage depends on the work performed and revisions; this guide does not promise a fixed price or a finished mix in one message. Check current plans.
Studio supports MP3, WAV, M4A, AAC and OGG audio uploads up to 50 MB per track. A supported accessible link can also supply media, but availability, download limits and processing still apply. Identify the imported file before mixing it and confirm you can use it in the intended publication.
No. Search can provide a source, uploader and source-supplied license information, but a downloaded file does not establish authorship or publication rights. Check the original source, applicable terms and required attribution for your use. You can instead upload a track for which you already have the required rights.
The add-music tool chooses its defaults from the amount of indexed source speech surviving in the selected window. A speech bed and lead music can therefore receive different gain and ducking defaults. Explicit settings override those defaults. Check the actual mix; smooth ducking responds to program sound level and can react to non-speech sounds too.
Do not assume it does. The voiceover ducking control lowers original program audio during its active window. Smooth music ducking uses the program audio as its reference; the separate voiceover layer is mixed later. Balance the music against uploaded narration explicitly and listen to every relevant passage.
Check the music item's preview-time end, the source track's length and its offset. If the track runs out and looping is off, the remaining part of its window can be silent. Request an appropriate repeat or a different track/window, then inspect the join and ending.
Valmera can attempt source music/voice separation and apply separate gains to speech or vocals and the remaining soundtrack. This needs suitable processed source media, available separation services and compliance with the source-length limit. Dense mixes can leave residue or artifacts. Listen before relying on complete removal; muting the whole original track would also remove its speech.
There is no single music percentage that fits every source. Balance the track against the actual speech first. Track gain changes one layer; integrated LUFS describes the overall mix. Valmera's optional social mastering targets −14 LUFS with −2 dBTP headroom, followed by output limiting. Those are processing settings, not a universal platform requirement or a guarantee of the measured final result.

Bring your footage and an approved track

Set the music window and balance, review the preview, then export in Studio. Editing requires a subscription and processing credits.

Open Studio →
See pricing →

Related Articles

Audio and music reference
Original audio, music, voiceover and final loudness treatment.
Add recorded narration
Place a voiceover and review the balance of the separate layers.
What sidechain ducking means
How the reference signal, attack and release affect gain reduction.
Export the final MP4
Check the intended edit, media readiness and delivered file.