← Home
TOOL

Published · Updated

Add Voiceover to Video

Upload your narration and Valmera mixes it the way an engineer would: the voiceover sits on its own layer, background music and the original audio duck underneath it automatically, and any layer's volume changes with a sentence. You describe the mix; the AI agent builds it.

It works on real footage and on canvas projects made from images — which is how you turn a folder of photos and a voice memo into a narrated video.

Add a Voiceover Free

Upload the video, upload the narration, describe the mix. No credit card required.

Start Free →
See pricing →

How to Add a Voiceover to a Video

  1. 1
    Upload your video and narration
    Drag in any video up to 2GB (up to 3 hours) — it's analyzed once with visible progress. Then add your recorded voiceover as an audio file, up to 50MB per track, or paste a link to bring the audio in.
  2. 2
    Type the request
    Type something like "add my voiceover over the whole video and duck everything else under it", "mute the original audio and use the voiceover instead", or "make the voiceover louder and the music quieter". The agent places the track and mixes the layers.
  3. 3
    Preview, refine, export
    Check the preview, adjust with follow-ups — "start the narration at 0:03", "master it to platform loudness" — then export a full-quality MP4 rendered from your original upload, with only a small corner mark on the free plan.

The upload is analyzed once with visible progress; after that, every mix change is a single chat message.

An Auto-Ducked Mix, Not a Volume Fight

The hard part of voiceover isn't recording it — it's the mix. Narration has to win over the music and the original location sound without the other layers lurching up and down. In Valmera, a voiceover automatically ducks the other layers under it, and music ducks under speech via sidechain compression — the smooth dip-and-recover behavior of a real mix engineer, not a two-level volume switch.

Every layer keeps independent gain on top of that. "Voiceover full, music at 15%, original audio at half" is one message. When the balance is right, one mastering request targets −14 LUFS with −1.5 dBTP true peak — the loudness platforms expect — so the finished mix holds up next to professional audio. Full details in the audio and music docs.

Replace the Original Audio with Your Narration

Screen recordings, b-roll, gameplay, product demos — often the original sound just needs to go. Ask the agent to mute the original audio and it drops to full silence, leaving your voiceover as the voice of the video. Add a bed underneath from the 24-track royalty-free music library and it ducks under the narration automatically.

Prefer to keep some of the original sound? Lower it instead of muting it — per-layer gain means "original audio at 20% under the voiceover" is a valid, one-sentence mix.

Narrated Slideshows: Voiceover Over Images

You don't need footage to need a voiceover. Valmera's canvas projects start with no video at all: images, clips, and music become a sequential timeline in any aspect ratio, stills get Ken Burns motion so they feel alive, and your narration plays over the top. Real-estate walkthroughs, photo essays, tutorials from screenshots — described in chat, assembled by the agent.

Honest scope: canvas projects don't get transcript-based captions, speed changes, or censor effects — manual captions, on-screen text, overlays, and music all work.

Bring Your Own Voice — Honest Limits

Three things this page won't pretend. Valmera does not generate AI voices or text-to-speech — the narration is yours, recorded by you. It does not offer denoise or "studio sound" processing — a noisy recording stays a noisy recording, so record as clean as you can. And it does not do per-speaker leveling within one track — layers are balanced against each other, not voices within a single file.

What you can trust is that every claim the agent makes about your mix is server-verified by the honesty layer — it cannot tell you it ducked the music if it didn't.

More Than a Voiceover Tool

Once the narration is in, the same chat finishes the video: cut the dead air, style word-accurate captions, add sound effects and a sound-design pass, reframe for vertical, grade the color. Browse all the AI video editing tools — one agent, every request.

Frequently Asked Questions

In Valmera you upload your video, add your recorded narration as an audio file (up to 50MB per track), and type what you want: "add my voiceover over the whole video and duck everything else under it". The agent places the narration on its own layer and mixes the other audio underneath. No timeline editing required — refinements are follow-up messages.
Yes. A voiceover in Valmera auto-ducks the other layers, and music ducks under speech with sidechain compression — the smooth, mix-engineer kind of ducking that dips while you speak and recovers in the pauses, not a hard volume switch. You can still override levels by asking: "music a touch louder in the intro".
Yes. Ask the agent to mute the original audio — it goes fully to silence — and your voiceover becomes the voice of the video, with music from the library or your uploads underneath if you want it. Each layer keeps independent volume control throughout.
No. Valmera doesn't offer AI voice generation or text-to-speech — you record the narration yourself and upload it. What Valmera automates is everything around the voice: placement, ducking the other layers under it, per-layer levels, and a one-request master to platform loudness.
Yes — Valmera supports canvas projects that start with no video at all: images, clips, and music arranged as a sequential timeline, in any aspect ratio, with Ken Burns motion on stills. Add your voiceover on top and you have a narrated slideshow. Honest note: canvas projects don't get transcript-based captions or speed changes — manual captions, text, overlays, and music all work.
Upload audio up to 50MB per track, or paste a link in chat and the agent pulls the audio into your project. Once imported, a track behaves the same however it arrived: the agent can place it, duck other layers under it, and adjust its volume on request.
Describe the balance and the agent sets it: every layer — voiceover, music, original audio — has independent gain, so "voiceover full, music at 15%" is one message. To finish, one mastering request targets −14 LUFS with −1.5 dBTP true peak, the loudness targets platforms expect, so the narration sits at a professional level.

Let the Mix Handle Itself

Your voice on top, everything else ducked underneath — by request. Free plan, no credit card.

Add a Voiceover Free →
See pricing →

Related Articles

Add Music to Video
24 CC0 tracks in 8 moods, auto-ducked under speech.
Mute a Video
Silence the original audio or lower any layer independently.
Docs: Audio & Music
Every audio capability, parameter by parameter.
All AI Video Editing Tools
The full Valmera toolset, one chat away.