← Home
TOOL

Published · Updated

Remove Silences, Filler Words and Add Captions — Automatically

These are three jobs everyone does to a podcast video, and doing them in three separate tools is how an afternoon disappears. Valmera does all three from one sentence — and, more importantly, in the right order, so the captions are timed against the edit rather than the raw recording.

“Cut all the silences and the ums, and add captions.”

Clean Up Your Recording

50 free credits on signup, no card required. Upload an episode and type one sentence.

Start free →
See pricing →

The Three Jobs, and Where Tools Get Them Wrong

Silences are detected across the whole recording and cut at their midpoints, snapped to word boundaries so nothing sounds clipped and no breath is severed halfway. You control how aggressive it is by saying so — "tighter", "leave the pauses before my punchlines".

What to watch for: Crude tools cut on raw waveform level, which means a quiet moment where you are still talking counts as silence. Detection here works from the transcript, so "silence" means nobody is speaking — not merely that the meter dropped.

Ums, uhs, ers and hmms come out in the same pass, plus any custom words you name — "you know", "basically", "like". Each removal is a cut at word-level timestamps, so the sentence closes cleanly around the gap.

What to watch for: Filler removal is only as good as the transcript's word timings. If a tool works from sentence-level timings it has to estimate where the um sat, and estimates clip the word after it.

Captions are generated from the same transcript and timed word by word, then burned in. 11 presets, 12 bundled fonts, sizes, positions, per-word emphasis and karaoke word-pop. Restyling later does not retime them.

What to watch for: Captions must be timed against the EDITED timeline, not the original. A tool that captions first and cuts second leaves every caption after the first cut drifting out of sync.

How to Auto-Edit a Podcast Video

  1. 1
    Upload the recording
    Up to 14GB or 3 hours. A one-time analysis produces the word-level transcript, the silences and the shot boundaries, with progress visible throughout.
  2. 2
    Ask for all three at once
    Type: "Cut all the silences and the ums, and add captions." One request, one pass — the agent sequences the cuts before the captions so nothing drifts.
  3. 3
    Correct anything you disagree with
    Restore a pause you wanted, restyle the captions, tighten it further — each in a sentence, without disturbing the rest of the edit. Then export a source-quality MP4.

Analysis time scales with recording length and shows progress; every edit afterward reuses it rather than re-analyzing.

The Same Sentence Can Ask for More

Because the agent sequences dependent steps itself, there is no reason to stop at three. Requests that work just as well in one go:

  • “…and remove the repeated takes, keeping the last one.”
  • “…and punch in on the words I stress.”
  • “…and put a chill track under my voice, quiet.”
  • “…and clip the section about pricing into a 9:16 version for Reels.”
  • “…and master it to the loudness platforms expect.”

Frequently Asked Questions

Yes — Valmera does all three from a single request. Upload the recording, say "cut the silences and the ums and add captions", and one pass removes every detected pause, strips every filler word, and burns in word-accurate captions timed against the edited timeline. You do not run three separate tools or line anything up by hand.
Because captions have to be timed against the finished edit, not the raw recording. If a tool adds captions first and cuts second, every caption after the first cut drifts. Valmera's agent sequences the dependency itself: cuts first, then captions against the resulting timeline, so they stay locked regardless of how much was removed.
Sometimes, and that is why restoring is cheap. A deliberate beat before a punchline looks identical to dead air in the audio. If it takes one you wanted, say so — "put back the pause before my closing line" — and it is restored without disturbing any of the other cuts, because each operation is an independent reversible version.
Up to 14GB or 3 hours per upload. Long recordings take a while to analyze on the first pass, with progress visible the whole way, and every edit after that reuses the analysis rather than redoing it.
It is built for video. The recording is analyzed for its transcript, silences, shot boundaries and what is on screen, so beyond the three jobs here you can also ask for punch-ins on your emphasized words, a 9:16 clip for Reels, music under the intro, or a colour grade — in the same conversation.
Yes, and on a podcast recording that is often the bigger win. It detects repeated phrasing — the three attempts at the same sentence — so you can keep the best one. Ask for it explicitly: "remove the repeated takes and keep the last one".
A downloadable H.264 MP4 rendered from your original upload at source quality — previews use a faster proxy, exports never do. Captions are burned into the video; there is no SRT or VTT export. Free-plan exports carry a small Valmera mark in the corner and every paid plan exports with no watermark, with a brief (~2.5-second) end card after your content on every plan.
Every account gets 50 free credits with no card required, which covers real edits on your own footage. Paid plans are Creator $30/mo (2,000 credits per billing cycle), Pro $50/mo (4,000) and Frontier $100/mo (10,000), each opening with a 3-day free trial. Requests charge in proportion to the work actually done, so a cleanup pass costs less than a full restyle.

One Sentence, Three Jobs Done

50 free credits, no card required. Upload an episode and see how much shorter it gets.

Start free →
See pricing →

Related Articles

Remove Silence From Video
The silence-cutting job on its own, in detail.
Remove Filler Words
Ums, uhs and your own custom tics.
Remove Repeated Takes
Keep the best of three attempts at the same line.
Valmera for Podcasters
The whole workflow, from a long recording to clips.