Remove Silences, Filler Words and Add Captions — Automatically
These are three jobs everyone does to a podcast video, and doing them in three separate tools is how an afternoon disappears. Valmera does all three from one sentence — and, more importantly, in the right order, so the captions are timed against the edit rather than the raw recording.
“Cut all the silences and the ums, and add captions.”
Clean Up Your Recording
50 free credits on signup, no card required. Upload an episode and type one sentence.
Start free →The Three Jobs, and Where Tools Get Them Wrong
Silences are detected across the whole recording and cut at their midpoints, snapped to word boundaries so nothing sounds clipped and no breath is severed halfway. You control how aggressive it is by saying so — "tighter", "leave the pauses before my punchlines".
What to watch for: Crude tools cut on raw waveform level, which means a quiet moment where you are still talking counts as silence. Detection here works from the transcript, so "silence" means nobody is speaking — not merely that the meter dropped.
Ums, uhs, ers and hmms come out in the same pass, plus any custom words you name — "you know", "basically", "like". Each removal is a cut at word-level timestamps, so the sentence closes cleanly around the gap.
What to watch for: Filler removal is only as good as the transcript's word timings. If a tool works from sentence-level timings it has to estimate where the um sat, and estimates clip the word after it.
Captions are generated from the same transcript and timed word by word, then burned in. 11 presets, 12 bundled fonts, sizes, positions, per-word emphasis and karaoke word-pop. Restyling later does not retime them.
What to watch for: Captions must be timed against the EDITED timeline, not the original. A tool that captions first and cuts second leaves every caption after the first cut drifting out of sync.
How to Auto-Edit a Podcast Video
- 1Upload the recordingUp to 14GB or 3 hours. A one-time analysis produces the word-level transcript, the silences and the shot boundaries, with progress visible throughout.
- 2Ask for all three at onceType: "Cut all the silences and the ums, and add captions." One request, one pass — the agent sequences the cuts before the captions so nothing drifts.
- 3Correct anything you disagree withRestore a pause you wanted, restyle the captions, tighten it further — each in a sentence, without disturbing the rest of the edit. Then export a source-quality MP4.
Analysis time scales with recording length and shows progress; every edit afterward reuses it rather than re-analyzing.
The Same Sentence Can Ask for More
Because the agent sequences dependent steps itself, there is no reason to stop at three. Requests that work just as well in one go:
- “…and remove the repeated takes, keeping the last one.”
- “…and punch in on the words I stress.”
- “…and put a chill track under my voice, quiet.”
- “…and clip the section about pricing into a 9:16 version for Reels.”
- “…and master it to the loudness platforms expect.”
Frequently Asked Questions
One Sentence, Three Jobs Done
50 free credits, no card required. Upload an episode and see how much shorter it gets.
Start free →