Add B-Roll to a Video Automatically
The tedious half of b-roll is finding the exact moment and landing on it. That half is mechanical, and an agent that can read your transcript does it better than you can by scrubbing.
To add b-roll to a video in Valmera, upload your footage and describe the moment: "cover the jump cuts in the first two minutes", or "show a shot of a busy street when I say commute". The AI agent reads the word-level transcript of the current edit, resolves those words to a program time, sources a clip — from your uploads, a link, a stock search, or footage it generates — and lays it over the picture while your audio keeps running. The runtime does not change, your captions stay on top, and it tells you exactly which moments it covered and with what.
Cover Your Cuts by Describing Them
Upload once, say which moments to cut away on, review the preview. 50 free credits, no card.
Add B-Roll Free →How to Add B-Roll to a Video
- 1Upload your main footageDrag in any MP4, MOV, MKV or WebM up to 14GB or 3 hours. A one-time analysis builds a word-level transcript with speaker labels, detects shots and silences, and lays out labeled frame tiles the agent reads directly — so it knows what is on screen before you type anything.
- 2Add whatever b-roll you already haveUpload your own clips and stills into the same project, or paste a link to one. This matters more than anything else on this page: footage from your own shoot matches the A-roll's white balance, exposure and lens character by default, and stock usually does not.
- 3Say which moments to coverType the literal sentence: "cover the jump cuts in the first two minutes with the kitchen clips I uploaded", or "when I say the dashboard, show the screen recording — keep my sentence running". Naming the words is what makes the placement exact.
- 4Ask for the shots you do not have"Find a stock clip of city traffic at night for the section about commuting." The agent searches, reads the candidate descriptions, picks the one that actually depicts the subject, downloads it and places it — and tells you it came from a stock library rather than describing it as something you filmed.
- 5Review the covers, then exportWatch the preview and correct in plain English — "that cutaway is too long, cut back on my next sentence", "move the clip after the intro", "do not cover the punchline at 3:40". When it reads right, export a full-quality H.264 MP4 rendered from your original file.
The analysis happens once per upload. Every cover, splice, move and correction after that is a chat message.
- "cover the jump cuts in the first two minutes with b-roll"
- "when I say the dashboard, show my screen recording — keep my sentence running"
- "find a stock clip of city traffic at night for the commuting section"
- "splice the product clip between the intro and the demo"
- "generate a still of a hand-drawn funnel diagram and push in on it for 3 seconds"
- "that cutaway is too long — cut back on my next sentence"
The Rule Comes Before the Tool
B-roll does exactly three jobs: it covers a cut you had to make, it illustrates a claim the words just made concrete, or it changes the pace of a take that has gone static. Cutting away for any other reason is worse than not cutting away at all — a generic timelapse under a paragraph about hiring is wallpaper, and viewers feel it as filler even when they cannot name why.
Four rules follow from that, and they hold whether a person or an agent is doing the placing. Cut away on the word, not after it — landing half a second late is the most common placement error in amateur edits. Never cover a punchline or a moment of direct address; when the content of the shot is the person's face, the face is the shot. End the cover on a beat in the narration, not when the clip runs out, because a cover that ends because the source ended is visible. And keep the density at one to three purposeful cutaways per minute; above that the piece turns into a montage and the speaker stops being present in their own video.
The full craft background — where the name came from, how much footage to shoot, and why b-roll is not the same thing as stock — is in the b-roll glossary entry. The rest of this page is about the machinery.
Cover or Splice: the Choice That Decides Your Runtime
Two completely different operations are both called "adding b-roll", and most tutorials skip the distinction. A cover replaces the picture for a window of program time; the audio, the music and the burned-in captions keep running and the video does not get one frame longer. A splice puts the clip in as its own scene, which pauses the speech and extends the video by the clip's length. The agent picks between them by asking one question: should the sentence keep going?
You do not have to know which one you want. "Show the dashboard while I keep talking" is a cover; "put the product clip between the intro and the demo" is a splice; and if the agent guesses wrong, one sentence fixes it. But knowing the difference is what lets you diagnose the single most common surprise in this job — the video came back longer than you left it.
How the Placement Actually Works
The transcript is of the edit, not of the recording. This is the part that makes automatic b-roll placement work at all. After you have cut silences, filler words and a fluffed sentence, every timestamp in the raw transcript is wrong — the word "dashboard" is no longer at 4:12 of the finished video. The agent works from the transcript the current edit actually keeps, in program time with the matching source spans, so "show it when I mention the dashboard" resolves to where that word lands after your cuts. Word boundaries in ordinary speech are roughly 200 to 400 milliseconds apart, which makes "on the word" a real target rather than a slogan.
The joins are already known. A twelve-minute talking head stripped of dead air and filler easily carries forty removals, and on a locked-off single camera every one of them is a jump cut. The agent made those cuts, so the joins are scene boundaries in the edit's own program map rather than something it has to detect afterwards. "Cover the joins in the first two minutes" resolves to a precise list instead of a guess. This is the single job where an agent beats a person on quality and not merely on time.
It looks inside your clip before choosing a moment. Splicing a fifteen-minute screen recording whole is never what you meant. The agent samples frames across an uploaded asset and reads them directly — the same way it reads labeled frame tiles of your main footage every turn — then picks a start point inside the clip and a duration, rather than taking the first few seconds and hoping. That is also how it avoids covering your sentence with four seconds of somebody adjusting a tripod.
A cover is composited, not cut in. The overlay renders above the footage and below the captions, so your captions stay readable over the b-roll. Its own audio does not play. It can sit full-frame, or as a positioned picture-in-picture with a size, an opacity, a fade or slide entrance, and a slow keyframed drift across the frame. It does not track anything moving underneath it.
A splice lands at a word edge. Ask for a clip in the middle of a take and the take is split cleanly at a word boundary rather than mid-syllable. Once a spliced scene exists it stays editable in place: it can be re-windowed to a different part of the clip, sped up from 0.25x to 4x so a ten-second screen recording becomes a five-second scene with nothing lost, cropped so one region of the frame becomes the whole scene, muted, padded instead of cover-cropped when a portrait clip lands in a landscape edit, or moved between two other scenes — all without removing and re-adding it, which would cost two renders and make your clip visibly vanish and return.
And because every agent reply is verified server-side against the edits actually recorded, the list of moments it says it covered is checked against what the edit contains. It cannot report a cutaway it did not place.
Six Places the Footage Can Come From
The order matters. Same-shoot footage matches the A-roll's white balance, exposure and lens character by default, which is why self-shot b-roll beats stock even when the stock is prettier. The agent is told to exhaust your own material first, and to say so rather than fake it if nothing fits.
On stock specifically: the agent reads the candidate descriptions and is instructed to prefer a clip that actually depicts the subject over one that merely shares a keyword, and to download the smallest rendition that still covers your output frame rather than pulling 4K it will throw away. It is also required to tell you which shots came from a stock library — describing a stock clip as something you filmed is exactly the kind of quiet dishonesty the verification layer exists to prevent. More on the generated sources in the AI generation docs.
Where the Agent Is Strong, and Where It Is Not
It is worth being blunt about this, because "AI b-roll" is sold as though taste were the solved part.
The division of labour that actually works: you name the moments and the meaning, it does the placement and the mechanics. "Cover the joins in the first two minutes, use the workshop clips, and do not touch the story at 6:10" is a better instruction than "add b-roll", and it takes about the same time to type.
When It Goes Wrong, and What to Say
- The video came back longer. It spliced when you wanted a cover. Say "that should have been a cutaway — keep my sentence running underneath" and the runtime comes back.
- The cutaway landed a beat late. Name the word rather than a timecode: "start it on the word commute, not after it". Timecodes drift when you make further cuts; words do not.
- A portrait clip got beheaded in a landscape edit. A cover crops to fill the frame by default. Ask for it padded instead and the whole picture shows on bars.
- The stock clip shares a keyword but not the subject. This is the most common stock failure and it is a search-query problem. Short and visual works — "city traffic at night" beats "the pace of modern life". Say what should be physically visible in the frame.
- It covered a moment that should have stayed on your face. Say "never cover me between 3:38 and 3:52". This is the correction worth making first, because a covered punchline costs more than a missing cutaway.
- There is too much of it. Ask for a density instead of a list: "keep it to about two cutaways a minute and drop the rest".
- Nothing suitable exists. The stock search reports no results rather than returning near-misses, and a source that is not configured on a deployment reports itself unavailable rather than failing silently — both are shapes designed so the agent cannot paper over them. If it says it could not find a shot, that is the true answer, and the fix is to upload one.
Nothing here is destructive. Your original upload is never modified, every cover and splice lives as an instruction on top of it, and any of them can be moved, re-windowed or removed without re-uploading a file.
How People Do This Without Valmera
All of these work, and for a lot of editors they work well. They differ in how much of the placement you personally perform.
Premiere Pro. Put the A-roll on V1 and drop b-roll on V2 above it — a higher track covers the picture while V1's audio keeps playing, which is the cover operation exactly. The fast version is source patching: point the source patch at V2, mark in and out in the Source monitor, park the playhead and press the overwrite key, and the clip lands on V2 at the playhead without touching the audio. It is precise and quick once it is in your fingers, and it is entirely manual — you find every moment yourself.
DaVinci Resolve. The same track-above logic on the Edit page, with the Cut page's source tape view making it faster to scan a long b-roll clip for a usable few seconds. Resolve is free for this, and its colour tools are the best place to actually match a stock clip to your A-roll, which is the step most people skip.
ffmpeg. A cover is one overlay filter with a time predicate — overlay=enable='between(t,4,8)' — mapping the audio from the first input so the b-roll stays silent. A splice is the concat demuxer over three pieces, which needs matching codec, resolution and frame rate or a re-encode. It is exact, scriptable and free, and it gives you no help whatsoever with the only hard question: which four seconds.
Online layer editors — Kapwing, Veed, Clipchamp — give you a browser timeline with layers and a stock library, which is enough for a handful of cutaways and much gentler than a desktop NLE. You are still scrubbing to find each moment.
The AI b-roll features. Several editors now ship some version of this. Descript pairs an AI assistant with a royalty-free stock library, and script-to-video tools like Pictory and InVideo build the video by matching stock footage to each line of a script — but that is a different job, because there the b-roll is the video rather than a cover over footage you shot. Feature sets and clip limits in this corner move month to month, so check each vendor's own current docs before you rely on a specific behaviour. The distinction that matters when comparing: does the tool place b-roll against your edit's transcript, at the joins it made, or does it lay visuals against a script? See Valmera vs Descript for the longer version.
Honest Limits, Specific to This Job
- A cover's audio does not play. Cover footage is silent by design, so a cutaway cannot bring in its own location atmos. A spliced clip keeps its audio; a cover never does.
- Nothing motion-tracks. A picture-in-picture overlay holds its position, or drifts along a path you describe. It will not follow a moving subject underneath it.
- Covers render below captions. That is the right default — your captions stay readable over b-roll — but if you wanted the cutaway to obscure them, it will not.
- Spliced clips are not captioned. Captions come from the transcript of your main footage, so a spliced scene carries no captions of its own. On-screen text can be placed over it manually.
- There is no crossfade or dissolve. The change from A-roll to b-roll is a hard cut. That is the correct default for this move and it is also a real limit if you wanted a soft one; a cover can fade in and out, but the underlying footage does not dissolve into it.
- Generated video is short and priced per second. 5 or 10 seconds per clip. It is the most expensive source on the list and the slowest, which is why the sourcing order puts it last.
- Stock is a supplier, not a match. No automatic system fixes the fact that a library clip was shot on a different camera in a different room. If colour continuity matters to you, plan to grade it — or shoot your own.
Related Tools and Reading
B-roll is rarely the first request. Most people arrive here after removing silence and filler words left them with forty visible joins, which is exactly the case this job is best at. When you have no cutaway to hide a join with, a punch-in on the A-roll does it using only footage you already have — the professional answer is usually a mix of the two. For the craft background read the b-roll glossary entry and jump cut; for the machinery, text and overlays and AI generation.
If you want a demo of your own product as the b-roll, record a website demo and splice the capture in. If you have no main video at all, the slideshow builder assembles images, clips and music on a bare canvas. Everything else is in the tools hub, and the same 108 tools are drivable from a Claude conversation over the Valmera MCP server.
Try It on Footage You Already Have
Upload, name the moments, review the covers. 50 free credits, no credit card.
Start free →Frequently Asked Questions
Add B-Roll by Describing It
The agentic AI video editor — say which moments to cut away on, and it places them on the word. Free plan available.
Start free →