Music, Voiceover & Audio Mixing
Add music or recorded narration, specify the balance, and review the complete mix before exporting.
Updated 2026-09-08. For track timing, mixing prompts and a synthetic ducking comparison, use the background-music guide.
How Audio Works in Valmera
Valmera mixes the audio layers of your edit — the original speech in your footage, background music (found on the web, your own uploads, or a pasted link), a voiceover track, and sound effects — and the AI agent controls the volume of each layer independently. You describe the mix in plain English ("add my track under the whole video, keep it quiet, duck it when the voiceover speaks") and the agent builds it, then renders a preview you can check before exporting.
Because every layer has its own gain, you never have to choose between "music on" and "music off." Speech can stay at full volume while music sits at 10–20% underneath it, and a voiceover can ride on top of both.
Music: Found on the Web or Your Own Tracks
Valmera finds music on the web: ask for "something chill under the whole video" and the agent searches by that vibe, fetches a real track, sets the level, and ducks it under your speech. Name a specific song and it looks for that one. Results preserve the provider, uploader and any source-supplied license metadata, and clearly marked no-derivatives results are excluded. That metadata is not proof of identity, ownership or publication rights, so verify the original source and terms before publishing to YouTube, TikTok, or a client.
Prefer your own sound? Upload tracks up to 50MB each, or paste a link to bring a song into the project. Attach the file in the chat or drop it straight onto the timeline — music appears as a draggable block you can reposition by hand (see Timeline, Transcript & Player). Upload limits for every asset type are listed in Uploading Footage.
Voiceover and Automatic Ducking
Upload a recorded voiceover and tell the agent where it belongs — over the intro, across a montage, or under a specific section. The voiceover becomes its own layer on the timeline, draggable like any other block.
These controls use different signals. With its ducking option enabled, a voiceover lowers original program audio by 12 dB during its active window. Smooth music ducking reacts to program sound level; it can react to ambience as well as speech. The separate voiceover layer is mixed later and does not automatically drive that music sidechain. Balance music against uploaded narration explicitly and listen before relying on intelligibility.
Sound Effects
Sound effects are real recordings found on the web. Describe the sound you want — a whoosh on a cut, a soft click, an impact under your strongest word, rain on a window — and the agent searches the web (Openverse, then Freesound), finds a real recorded effect (capped at about 15 seconds by default, license-labeled, no-derivatives excluded), and places it at the moment you asked for. Valmera does not synthesize audio. See AI Generation for how found and generated media is charged.
You can also bring your own: upload a sound file and the agent places it the same way. Either kind can be moved to a different moment or removed later — just say which effect and where it should go.
Loudness Mastering & Audio Analysis
After balancing the layers, optional social mastering targets −14 LUFS with −2 dBTP headroom, followed by output limiting. This processing preset does not establish a universal platform requirement or guarantee the measured final output. Check the destination specification when one applies and inspect the delivered file. Mastering changes the overall mix; it does not repair a poor music-to-voice balance.
Track energy or beat analysis can help find a cue or propose cut timing. A reliable result depends on useful source evidence. Listen to the selected point and inspect the reveal and surrounding speech before accepting a beat-aligned edit.
What You Can Ask For
Usage depends on processing and revisions. Check the actual request result and see How Credits Work for the details.
Honest Limits
- No denoise or "studio sound." Valmera does not remove background noise, hum, or echo from your recording. If the source audio is rough, it stays rough — record clean when you can.
- No per-speaker balancing inside one track. If two people were recorded into the same audio track at different volumes, Valmera cannot level them against each other. Layer gain applies to the whole speech track.
- Music separation needs review. Ask to lower the music already mixed into your original recording while keeping the speech. Subject to service availability, processed source media and the source-length limit, Valmera can separate speech/vocals from the remaining soundtrack and apply independent gains. This differs from lowering an added music layer. Dense mixes can leave residue or artifacts; listen before relying on removal.
- No AI audio generation. Valmera does not synthesize music, songs or sound effects — music and sound effects are real recordings the agent finds on the web, or files you provide.
- Check the supported tools. Confirm that the requested operation is available and audition the output. Recorded settings and a completion reply do not prove that an unsupported audio treatment was performed.
Common Questions
Where does Valmera's music come from?
Use your own approved audio, an accessible supported link, or ask the agent for music candidates. Source-supplied license information is evidence to check, not publication clearance. Studio audio uploads support MP3, WAV, M4A, AAC and OGG up to 50 MB per track.
Can Valmera automatically duck music under a voiceover?
Voiceover ducking lowers original program audio during the voiceover window. Smooth music ducking uses program audio as its reference, while the separate voiceover is mixed later. Uploaded narration does not automatically become that music reference. Set the music level against it explicitly and review the mix.
Can I set different volumes for speech, music, and voiceover?
Yes. Valmera mixes three separate audio layers — the original speech in your footage, uploaded music, and voiceover — and each layer has its own gain. You can say things like "raise the voiceover a little and drop the music to 15%" and only those layers change.
Can Valmera remove background noise or do "studio sound" cleanup?
No. Valmera does not offer denoise, audio restoration, or per-speaker level balancing inside one source track. Prepare that audio in a suitable external tool when needed, then check the mix in Valmera. Changing a layer volume or adding music does not repair noise in the original recording.
Can I mute the original audio of my video?
Yes. Request muting of the original soundtrack while retaining the intended music and voiceover layers. Review the result: changing program audio can also change the signal that drives smooth music ducking.
Can Valmera generate music with AI?
No — Valmera does not generate music, songs or sound effects with AI, and the agent will say so if you ask. Music comes from a web search for real, license-labeled tracks, or files you provide. Sound effects are real recordings the agent finds on the web too: describe a sound — a whoosh, a soft click, rain on a window — and it finds a real recorded one and places it in the mix.