Music, Voiceover & Audio Mixing
Add music, layer a voiceover, and get a clean mix — by typing what you want.
How Audio Works in Valmera
Valmera mixes the audio layers of your edit — the original speech in your footage, background music (from the built-in library, your own uploads, or a pasted link), a voiceover track, and sound effects — and the AI agent controls the volume of each layer independently. You describe the mix in plain English ("add my track under the whole video, keep it quiet, duck it when the voiceover speaks") and the agent builds it, then renders a preview you can check before exporting.
Because every layer has its own gain, you never have to choose between "music on" and "music off." Speech can stay at full volume while music sits at 10–20% underneath it, and a voiceover can ride on top of both.
Music: Built-In Library or Your Own Tracks
Valmera ships a built-in music library: 24 royalty-free CC0 tracks across 8 moods — upbeat, chill, cinematic, corporate, dramatic, hip-hop, ambient, and inspiring. Ask for "something chill under the whole video" and the agent picks a track, sets the level, and ducks it under your speech. CC0 means no licensing surprises when you publish to YouTube, TikTok, or a client.
Prefer your own sound? Upload tracks up to 50MB each, or paste a link to bring a song into the project. Attach the file in the chat or drop it straight onto the timeline — music appears as a draggable block you can reposition by hand (see Timeline, Transcript & Player). Upload limits for every asset type are listed in Uploading Footage.
Voiceover and Automatic Ducking
Upload a recorded voiceover and tell the agent where it belongs — over the intro, across a montage, or under a specific section. The voiceover becomes its own layer on the timeline, draggable like any other block.
When a voiceover or your own speech plays over music, Valmera ducks the music with smooth sidechain compression: it dips under the voice and rises back naturally when the line ends — the ride a mix engineer would perform by hand, not a hard volume step. No keyframes, no envelope drawing — the mix stays intelligible by default.
Sound Effects & the Sound-Design Pass
Valmera includes 18 built-in one-shot sound effects — whooshes, impacts, risers, clicks, booms, dings — that the agent places wherever you ask. Or hand over the whole job: a one-request sound-design pass adds whooshes at the cuts, an impact on your strongest word, and a riser into the energy peak of the video.
Need a sound that is not in the pack? Describe it and the agent generates a one-shot with AI — from half a second up to 22 seconds — and drops it into the mix. See AI Generation for how generated media is charged.
Loudness Mastering & Audio Analysis
When the mix is set, ask for mastering: Valmera masters the final mix to -14 LUFS with -1.5 dBTP true-peak headroom — the loudness targets platforms expect — so your video does not play quieter than everyone else's.
The agent can also listen for you. Ask where the drop is in a music track and it analyzes the song's energy and tells you — useful for timing a beat drop to a reveal. The same audio perception powers beat-synced cuts, and it declines honestly when a track has no reliable tempo to snap to.
What You Can Ask For
Each request is a normal chat message, and a simple audio change costs only a few credits — see How Credits Work for the details.
Honest Limits
- No denoise or "studio sound." Valmera does not remove background noise, hum, or echo from your recording. If the source audio is rough, it stays rough — record clean when you can.
- No per-speaker balancing inside one track. If two people were recorded into the same audio track at different volumes, Valmera cannot level them against each other. Layer gain applies to the whole speech track.
- No separating music from speech. If music is already baked into your footage's own audio track, Valmera cannot split it out — layer control applies to the layers Valmera adds.
- No AI music generation. Valmera generates sound effects, not songs — music comes from the built-in royalty-free library or files you provide.
- The agent will say so. Valmera's honesty layer checks every reply against what the system actually did — if you ask for something outside its toolkit, it tells you instead of pretending it worked.
Common Questions
Does Valmera include a stock-music library?
Yes. Valmera ships 24 royalty-free CC0 tracks across 8 moods — upbeat, chill, cinematic, corporate, dramatic, hip-hop, ambient, and inspiring — that the agent can lay under your video on request. You can also upload your own tracks, each up to 50MB, or paste a link to pull a song into the project.
Can Valmera automatically duck music under a voiceover?
Yes. Valmera ducks music under speech with sidechain compression — the smooth ride a mix engineer would dial in, not a hard volume step. The music dips while the voiceover or dialogue plays and rises back naturally afterwards. You do not have to keyframe anything; ask for it in the chat and the agent handles the mix.
Can I set different volumes for speech, music, and voiceover?
Yes. Valmera mixes three separate audio layers — the original speech in your footage, uploaded music, and voiceover — and each layer has its own gain. You can say things like "raise the voiceover a little and drop the music to 15%" and only those layers change.
Can Valmera remove background noise or do "studio sound" cleanup?
No. Valmera does not offer denoise, audio restoration, or per-speaker level balancing inside a single track. If you ask for it, the agent tells you honestly that it cannot do it instead of pretending — that is Valmera's honesty layer at work.
Can I mute the original audio of my video?
Yes. Ask the agent to mute the original audio and it will silence the footage's own soundtrack while keeping your music and voiceover layers exactly as you set them — useful for b-roll montages or music-only cuts.
Can Valmera generate music with AI?
No — Valmera does not generate songs or music with AI, and the agent will say so if you ask. Music comes from the built-in royalty-free library or files you provide. What AI generation does cover is sound effects: describe a sound — a whoosh, a soft click, rain on a window — and the agent creates a short one-shot and places it in the mix.