← Home
TOOL

Published · Updated

Add B-Roll to a Video Automatically

The tedious half of b-roll is finding the exact moment and landing on it. That half is mechanical, and an agent that can read your transcript does it better than you can by scrubbing.

To add b-roll to a video in Valmera, upload your footage and describe the moment: "cover the jump cuts in the first two minutes", or "show a shot of a busy street when I say commute". The AI agent reads the word-level transcript of the current edit, resolves those words to a program time, sources a clip — from your uploads, a link, a stock search, or footage it generates — and lays it over the picture while your audio keeps running. The runtime does not change, your captions stay on top, and it tells you exactly which moments it covered and with what.

Cover Your Cuts by Describing Them

Upload once, say which moments to cut away on, review the preview. 50 free credits, no card.

Add B-Roll Free →
See pricing →

How to Add B-Roll to a Video

  1. 1
    Upload your main footage
    Drag in any MP4, MOV, MKV or WebM up to 14GB or 3 hours. A one-time analysis builds a word-level transcript with speaker labels, detects shots and silences, and lays out labeled frame tiles the agent reads directly — so it knows what is on screen before you type anything.
  2. 2
    Add whatever b-roll you already have
    Upload your own clips and stills into the same project, or paste a link to one. This matters more than anything else on this page: footage from your own shoot matches the A-roll's white balance, exposure and lens character by default, and stock usually does not.
  3. 3
    Say which moments to cover
    Type the literal sentence: "cover the jump cuts in the first two minutes with the kitchen clips I uploaded", or "when I say the dashboard, show the screen recording — keep my sentence running". Naming the words is what makes the placement exact.
  4. 4
    Ask for the shots you do not have
    "Find a stock clip of city traffic at night for the section about commuting." The agent searches, reads the candidate descriptions, picks the one that actually depicts the subject, downloads it and places it — and tells you it came from a stock library rather than describing it as something you filmed.
  5. 5
    Review the covers, then export
    Watch the preview and correct in plain English — "that cutaway is too long, cut back on my next sentence", "move the clip after the intro", "do not cover the punchline at 3:40". When it reads right, export a full-quality H.264 MP4 rendered from your original file.

The analysis happens once per upload. Every cover, splice, move and correction after that is a chat message.

EXAMPLE PROMPTS
  • "cover the jump cuts in the first two minutes with b-roll"
  • "when I say the dashboard, show my screen recording — keep my sentence running"
  • "find a stock clip of city traffic at night for the commuting section"
  • "splice the product clip between the intro and the demo"
  • "generate a still of a hand-drawn funnel diagram and push in on it for 3 seconds"
  • "that cutaway is too long — cut back on my next sentence"

The Rule Comes Before the Tool

B-roll does exactly three jobs: it covers a cut you had to make, it illustrates a claim the words just made concrete, or it changes the pace of a take that has gone static. Cutting away for any other reason is worse than not cutting away at all — a generic timelapse under a paragraph about hiring is wallpaper, and viewers feel it as filler even when they cannot name why.

Four rules follow from that, and they hold whether a person or an agent is doing the placing. Cut away on the word, not after it — landing half a second late is the most common placement error in amateur edits. Never cover a punchline or a moment of direct address; when the content of the shot is the person's face, the face is the shot. End the cover on a beat in the narration, not when the clip runs out, because a cover that ends because the source ended is visible. And keep the density at one to three purposeful cutaways per minute; above that the piece turns into a montage and the speaker stops being present in their own video.

The full craft background — where the name came from, how much footage to shoot, and why b-roll is not the same thing as stock — is in the b-roll glossary entry. The rest of this page is about the machinery.

Cover or Splice: the Choice That Decides Your Runtime

Two completely different operations are both called "adding b-roll", and most tutorials skip the distinction. A cover replaces the picture for a window of program time; the audio, the music and the burned-in captions keep running and the video does not get one frame longer. A splice puts the clip in as its own scene, which pauses the speech and extends the video by the clip's length. The agent picks between them by asking one question: should the sentence keep going?

A cover leaves the runtime unchanged; a splice extends itThree versions of the same twelve-second program are drawn to the same time scale. The first is the original: one continuous block of A-roll picture over one continuous block of speech, twelve seconds long. The second is a cover: the picture between four and eight seconds is replaced by a b-roll block, while the speech bar underneath runs unbroken from end to end and the captions still sit on top of the b-roll. The program is still twelve seconds. The third is a splice: a five-second clip is inserted at four seconds as its own scene, so the A-roll picture and its speech are pushed to the right and resume afterwards, and the program is now seventeen seconds long.ORIGINAL12.0sA-ROLL PICTUREspeechCOVER12.0sunchangedB-ROLLCAPTIONS STAYspeech runs through, unbrokenSPLICE17.0s+5.0s longerCLIP AS ITS OWN SCENEA-ROLL RESUMESsentence paused0s4s8s12s17sAll three rows are drawn to the same time scale — the splice row is genuinely longer.

You do not have to know which one you want. "Show the dashboard while I keep talking" is a cover; "put the product clip between the intro and the demo" is a splice; and if the agent guesses wrong, one sentence fixes it. But knowing the difference is what lets you diagnose the single most common surprise in this job — the video came back longer than you left it.

How the Placement Actually Works

The transcript is of the edit, not of the recording. This is the part that makes automatic b-roll placement work at all. After you have cut silences, filler words and a fluffed sentence, every timestamp in the raw transcript is wrong — the word "dashboard" is no longer at 4:12 of the finished video. The agent works from the transcript the current edit actually keeps, in program time with the matching source spans, so "show it when I mention the dashboard" resolves to where that word lands after your cuts. Word boundaries in ordinary speech are roughly 200 to 400 milliseconds apart, which makes "on the word" a real target rather than a slogan.

The joins are already known. A twelve-minute talking head stripped of dead air and filler easily carries forty removals, and on a locked-off single camera every one of them is a jump cut. The agent made those cuts, so the joins are scene boundaries in the edit's own program map rather than something it has to detect afterwards. "Cover the joins in the first two minutes" resolves to a precise list instead of a guess. This is the single job where an agent beats a person on quality and not merely on time.

It looks inside your clip before choosing a moment. Splicing a fifteen-minute screen recording whole is never what you meant. The agent samples frames across an uploaded asset and reads them directly — the same way it reads labeled frame tiles of your main footage every turn — then picks a start point inside the clip and a duration, rather than taking the first few seconds and hoping. That is also how it avoids covering your sentence with four seconds of somebody adjusting a tripod.

A cover is composited, not cut in. The overlay renders above the footage and below the captions, so your captions stay readable over the b-roll. Its own audio does not play. It can sit full-frame, or as a positioned picture-in-picture with a size, an opacity, a fade or slide entrance, and a slow keyframed drift across the frame. It does not track anything moving underneath it.

A splice lands at a word edge. Ask for a clip in the middle of a take and the take is split cleanly at a word boundary rather than mid-syllable. Once a spliced scene exists it stays editable in place: it can be re-windowed to a different part of the clip, sped up from 0.25x to 4x so a ten-second screen recording becomes a five-second scene with nothing lost, cropped so one region of the frame becomes the whole scene, muted, padded instead of cover-cropped when a portrait clip lands in a landscape edit, or moved between two other scenes — all without removing and re-adding it, which would cost two renders and make your clip visibly vanish and return.

And because every agent reply is verified server-side against the edits actually recorded, the list of moments it says it covered is checked against what the edit contains. It cannot report a cutaway it did not place.

Six Places the Footage Can Come From

The order matters. Same-shoot footage matches the A-roll's white balance, exposure and lens character by default, which is why self-shot b-roll beats stock even when the stock is prettier. The agent is told to exhaust your own material first, and to say so rather than fake it if nothing fits.

1. Your own uploads
Clips and stills you drop into the project. Always the best match, and the only source that carries the meaning you actually intended.
2. A pasted link
A direct file link, or a YouTube, TikTok or Vimeo page. Pulled into the project as an asset — you should never be told to download something the editor could have fetched.
3. A stock library search
Returns candidates only; nothing enters the video until one is chosen. The search orientation is derived from your output frame, so a 9:16 edit gets vertical footage.
4. A generated still
From a prompt, or by restyling a frame of your own footage. Fast and cheap, and with a slow zoom or pan it reads as a deliberate insert rather than a frozen frame.
5. A generated clip
Real moving footage of 5 or 10 seconds, from a prompt or by animating a still. Slow and charged per second — spend it on the one shot that matters.
6. A recorded web page
A headless browser opens a public URL at your aspect ratio and pans down it, or clicks through steps you describe. The honest answer to "show my landing page here".

On stock specifically: the agent reads the candidate descriptions and is instructed to prefer a clip that actually depicts the subject over one that merely shares a keyword, and to download the smallest rendition that still covers your output frame rather than pulling 4K it will throw away. It is also required to tell you which shots came from a stock library — describing a stock clip as something you filmed is exactly the kind of quiet dishonesty the verification layer exists to prevent. More on the generated sources in the AI generation docs.

Where the Agent Is Strong, and Where It Is Not

It is worth being blunt about this, because "AI b-roll" is sold as though taste were the solved part.

Genuinely strong
Covering the jump cuts it created — it knows every join by position. Landing a cover on the exact word that names the thing. Choosing a usable moment inside a long clip instead of its first four seconds. Doing forty of these consistently without the fifteenth one drifting. Getting the orientation, the fit and the duration right every time.
Genuinely weak
Knowing which shot means something. It cannot tell that the clip at 4:12 is the one you travelled for, that a stock shot of a stranger typing says nothing about your product, or that the section you think is boring is the one your audience came for. Left to itself it will also cover too much, because covering more looks like doing more.

The division of labour that actually works: you name the moments and the meaning, it does the placement and the mechanics. "Cover the joins in the first two minutes, use the workshop clips, and do not touch the story at 6:10" is a better instruction than "add b-roll", and it takes about the same time to type.

When It Goes Wrong, and What to Say

  • The video came back longer. It spliced when you wanted a cover. Say "that should have been a cutaway — keep my sentence running underneath" and the runtime comes back.
  • The cutaway landed a beat late. Name the word rather than a timecode: "start it on the word commute, not after it". Timecodes drift when you make further cuts; words do not.
  • A portrait clip got beheaded in a landscape edit. A cover crops to fill the frame by default. Ask for it padded instead and the whole picture shows on bars.
  • The stock clip shares a keyword but not the subject. This is the most common stock failure and it is a search-query problem. Short and visual works — "city traffic at night" beats "the pace of modern life". Say what should be physically visible in the frame.
  • It covered a moment that should have stayed on your face. Say "never cover me between 3:38 and 3:52". This is the correction worth making first, because a covered punchline costs more than a missing cutaway.
  • There is too much of it. Ask for a density instead of a list: "keep it to about two cutaways a minute and drop the rest".
  • Nothing suitable exists. The stock search reports no results rather than returning near-misses, and a source that is not configured on a deployment reports itself unavailable rather than failing silently — both are shapes designed so the agent cannot paper over them. If it says it could not find a shot, that is the true answer, and the fix is to upload one.

Nothing here is destructive. Your original upload is never modified, every cover and splice lives as an instruction on top of it, and any of them can be moved, re-windowed or removed without re-uploading a file.

How People Do This Without Valmera

All of these work, and for a lot of editors they work well. They differ in how much of the placement you personally perform.

Premiere Pro. Put the A-roll on V1 and drop b-roll on V2 above it — a higher track covers the picture while V1's audio keeps playing, which is the cover operation exactly. The fast version is source patching: point the source patch at V2, mark in and out in the Source monitor, park the playhead and press the overwrite key, and the clip lands on V2 at the playhead without touching the audio. It is precise and quick once it is in your fingers, and it is entirely manual — you find every moment yourself.

DaVinci Resolve. The same track-above logic on the Edit page, with the Cut page's source tape view making it faster to scan a long b-roll clip for a usable few seconds. Resolve is free for this, and its colour tools are the best place to actually match a stock clip to your A-roll, which is the step most people skip.

ffmpeg. A cover is one overlay filter with a time predicate — overlay=enable='between(t,4,8)' — mapping the audio from the first input so the b-roll stays silent. A splice is the concat demuxer over three pieces, which needs matching codec, resolution and frame rate or a re-encode. It is exact, scriptable and free, and it gives you no help whatsoever with the only hard question: which four seconds.

Online layer editors — Kapwing, Veed, Clipchamp — give you a browser timeline with layers and a stock library, which is enough for a handful of cutaways and much gentler than a desktop NLE. You are still scrubbing to find each moment.

The AI b-roll features. Several editors now ship some version of this. Descript pairs an AI assistant with a royalty-free stock library, and script-to-video tools like Pictory and InVideo build the video by matching stock footage to each line of a script — but that is a different job, because there the b-roll is the video rather than a cover over footage you shot. Feature sets and clip limits in this corner move month to month, so check each vendor's own current docs before you rely on a specific behaviour. The distinction that matters when comparing: does the tool place b-roll against your edit's transcript, at the joins it made, or does it lay visuals against a script? See Valmera vs Descript for the longer version.

Honest Limits, Specific to This Job

  • A cover's audio does not play. Cover footage is silent by design, so a cutaway cannot bring in its own location atmos. A spliced clip keeps its audio; a cover never does.
  • Nothing motion-tracks. A picture-in-picture overlay holds its position, or drifts along a path you describe. It will not follow a moving subject underneath it.
  • Covers render below captions. That is the right default — your captions stay readable over b-roll — but if you wanted the cutaway to obscure them, it will not.
  • Spliced clips are not captioned. Captions come from the transcript of your main footage, so a spliced scene carries no captions of its own. On-screen text can be placed over it manually.
  • There is no crossfade or dissolve. The change from A-roll to b-roll is a hard cut. That is the correct default for this move and it is also a real limit if you wanted a soft one; a cover can fade in and out, but the underlying footage does not dissolve into it.
  • Generated video is short and priced per second. 5 or 10 seconds per clip. It is the most expensive source on the list and the slowest, which is why the sourcing order puts it last.
  • Stock is a supplier, not a match. No automatic system fixes the fact that a library clip was shot on a different camera in a different room. If colour continuity matters to you, plan to grade it — or shoot your own.

Related Tools and Reading

B-roll is rarely the first request. Most people arrive here after removing silence and filler words left them with forty visible joins, which is exactly the case this job is best at. When you have no cutaway to hide a join with, a punch-in on the A-roll does it using only footage you already have — the professional answer is usually a mix of the two. For the craft background read the b-roll glossary entry and jump cut; for the machinery, text and overlays and AI generation.

If you want a demo of your own product as the b-roll, record a website demo and splice the capture in. If you have no main video at all, the slideshow builder assembles images, clips and music on a bare canvas. Everything else is in the tools hub, and the same 108 tools are drivable from a Claude conversation over the Valmera MCP server.

Try It on Footage You Already Have

Upload, name the moments, review the covers. 50 free credits, no credit card.

Start free →
See pricing →

Frequently Asked Questions

Upload your footage to Valmera and type what you want covered — for example "cover the jump cuts in the first two minutes with b-roll" or "show a shot of a busy city when I say the word commute". The agent reads the word-level transcript of the current edit, finds the program time of those words, sources a clip (from your uploads, a link you paste, a stock search, or footage it generates), and lays it over the picture while your audio keeps running. It then tells you which moments it covered and with what.
It can place b-roll accurately and it can source it, but choosing which shot carries meaning is judgement, and judgement is where an AI is weakest. The agent is strong at the mechanical half: it knows the exact program time of every word, it knows where every cut it made is, and it can start a cover on the word rather than half a second after it. It does not know that the shot at 4:12 is the one you drove two hours to get, or that the clip of a stranger typing says nothing about your product. Name the moments that matter and let it do the placement.
Five sources, in the order the agent should try them: clips and images you uploaded to the project; a link you paste (a direct file, or a YouTube, TikTok or Vimeo page); a stock library search that returns candidates for the agent to choose between; an AI-generated still with a slow Ken Burns move on it; or a short AI-generated video clip of 5 or 10 seconds. It can also record a live web page as footage, which is the honest answer for "show my landing page here". If none of those fit, it is supposed to say so and ask you for a clip rather than covering the moment with something irrelevant.
Only if you splice. A cover swaps the picture for a window of program time and leaves the runtime untouched — your sentence keeps playing underneath and the video is exactly as long as it was. A splice inserts the clip as its own scene, which pauses the speech and makes the video longer by the clip's duration. Both are called "adding b-roll". If your cutaway lengthened the video when you did not expect it to, you spliced when you meant to cover; say "that should have been a cutaway, keep my sentence running" and it is re-done as a cover.
Yes, and this is the strongest case for an agent doing it. When silence removal or filler-word removal cuts a locked-off single-camera take, every removal leaves a visible jump. The agent made those cuts, so their positions are scene boundaries in the edit rather than something it has to hunt for — "cover the joins in the first two minutes" resolves to a precise list. It is still worth triaging: cover the joins that fall mid-sentence, punch in on the ones at a paragraph break, and leave the rest as honest cuts.
No. Cover footage is silent by design, which is the standard the whole craft uses — two audio sources under one sentence is mush. That means a cutaway cannot bring in its own location atmos. A spliced clip is different: it becomes its own scene and its audio plays, and you can silence it by asking for that clip to be muted.
One to three purposeful cutaways per minute reads as edited. Past roughly three a minute the piece turns into a montage and the speaker stops being present in their own video. On screen, plan for b-roll to cover roughly 20 to 30 percent of the runtime, with each cover on screen for two to six seconds — under about 1.5 seconds a viewer registers a flash rather than an image, and past about eight seconds over someone still talking the shot starts competing with the words.
Yes, two ways, and they cost differently. A generated still is fast and cheap, and with a slow zoom or pan on it, it reads as a deliberate insert rather than a frozen frame — this is the right choice for a diagram, a concept, or a product shot that does not exist. A generated video clip is real moving footage of 5 or 10 seconds, and it is slow and charged per second of output, so it is worth spending deliberately on the one shot that matters rather than on every cutaway in the video.
Say so in plain English. "Move that clip after the intro" reorders a spliced scene in place, "only show the last four seconds of it" re-windows it, "speed it up so it takes five seconds instead of ten" changes its rate with nothing cut, and "that cutaway is too long, cut back on my next sentence" shortens a cover. Nothing is destructive — your original upload is never modified, and any placement can be undone or replaced without re-uploading anything.

Add B-Roll by Describing It

The agentic AI video editor — say which moments to cut away on, and it places them on the word. Free plan available.

Start free →
See pricing →

Related Articles

Glossary: B-Roll
The definition, the 16mm origin of the name, and the rule of when to cut away.
Add Zoom Effects
The punch-in — what hides a cut when you have no cutaway to hide it with.
Remove Silence from Video
Where most of the joins that need covering come from in the first place.
Docs: AI Generation
Generated stills, short clips and link import — sourcing a cutaway you never filmed.