Every podcast, webinar, and stream contains a week of social content — if someone does the clipping. Here's how to get captioned, vertical, hook-first clips by describing them instead of cutting them.
Clip Your Last Recording
Upload the long video, describe the clips, post them today.
These are exactly the things you can ask for. A request like "hook first, karaoke captions, 45 seconds, end on the laugh" is a complete clip spec — the agent executes it, down to the caption style, size, and position.
How the Agent Reads a Long Video
When you upload, Valmera analyzes the whole file before you type a word: it builds a word-level transcript (every word timestamped), detects shot changes, maps silences, and generates vision descriptions of what's on screen. That analysis is why "clip the pricing story" works — the agent can find the pricing story.
Word-level transcript
Every word timestamped — so a quote, a topic, or "around 22:00" all resolve to exact cut points.
Shot detection
Camera and scene changes are mapped, so clips start and end on clean visual boundaries.
Silence map
Dead air is located up front — so "cut the silences" resolves instantly, and clips come back tight.
Vision descriptions
The agent knows what's on screen, not just what was said — "the part where the dashboard is up" is findable.
Main uploads go up to 2GB — MP4, MOV, MKV, WebM and similar — long recordings fit at moderate export bitrates. See upload formats and limits for the full list. Analysis runs once per upload — longer files take longer, and progress is visible while it runs; every clip you request afterwards reuses it.
The Workflow
1
Upload the long video
The full podcast, webinar, or stream VOD — up to 2GB. Valmera transcribes it word by word and maps shots and silences across the whole file.
2
Point at a moment
Name it ("the pricing story"), timestamp it ("around 22:00"), or ask the agent to propose the strongest five and pick one. One clip per request — repeat for each moment.
3
Specify platform and style
"9:16 for Reels, blurred-pad framing, 45 seconds, karaoke captions, hook first." One sentence sets the whole format.
4
Review and refine the clip
"Start two seconds earlier." "Different hook line." "Captions bigger." Notes apply to the clip in front of you — and the agent's replies are checked against what it actually rendered, so it can't claim a change it didn't make.
5
Export, then request the next clip
Download the clip as its own MP4, rendered from your original full-quality file — previews use a fast proxy, final exports don't. Then ask for the next moment; every clip reuses the same analysis.
Going Vertical: Crop, Pad, or Blurred-Pad
Converting 16:9 footage to a vertical clip means deciding what happens to the sides of the frame. Valmera gives you three ways, and you pick per clip:
Crop
Fills the vertical frame by trimming the sides. Best when the speaker sits near the center.
Pad
Keeps the full widescreen image with bars above and below. Nothing gets cut off.
Blurred-pad
Fills the empty space with a blurred copy of the footage — the classic podcast-clip look.
All four social ratios are supported — 9:16 for TikTok, Reels, and Shorts, 1:1 and 4:5 for feed posts, and 16:9 stays available for the long cut. Valmera doesn't auto-track subjects; if a crop lands off the speaker, one follow-up message moves it. Details in the reframing guide. And unlike batch clippers that score and export on their own — see Valmera vs Opus Clip — you direct every clip, on a free tier of 20 daily credits.
Example Requests to Copy
Podcast to TikTok
From this 60-minute podcast, make a 45-second vertical clip of the most quotable moment. Blurred-pad framing, karaoke captions, hook first, clean cuts. Skip the intro chatter.
Webinar to teaser
Turn this webinar into one 60-second teaser for LinkedIn. Use the strongest insight as the hook, add captions, and end on the end-card image I uploaded with a fade to black.
Stream highlights
From this 90-minute stream VOD, cut a 10-minute highlight reel of the best plays. After I review it, I'll ask for 30-second vertical clips of the funniest moments with big captions.
Upload the long video to Valmera (up to 2GB — MP4, MOV, MKV, WebM), let it analyze the footage, then describe one clip at a time: platform, length, moment, and caption style. The agent cuts it, reframes it to 9:16, captions it from the word-level transcript, and returns it as a preview you refine by chat — then you ask for the next clip.
Yes — that's the difference between directed and batch clipping. Tell the agent "clip the story about the first customer" or "the demo section around 22 minutes" and that's the clip you get. You can also ask it to propose the strongest moments first.
30–60 seconds for TikTok, Reels, and Shorts is the sweet spot. Ask for a hook-first structure — the agent will open on the strongest line rather than the chronological start.
Yes, on request, three ways: crop (fills the 9:16 frame by trimming the sides), pad (keeps the full 16:9 image with bars), or blurred-pad (fills the empty space with a blurred copy of the footage — the podcast-clip look). It does not auto-track subjects, so if a crop lands off the speaker, say "shift the crop" in a follow-up and the agent re-renders. 16:9, 9:16, 1:1, and 4:5 are all supported.
The limit is file size, not duration: main uploads go up to 2GB — long recordings fit at moderate export bitrates. Analysis runs once per upload — longer files take longer, and progress is visible while it runs; every clip request after that reuses the same analysis.
As many as the footage supports — three, five, or ten, requested one at a time. Each clip is its own request on the same upload, charged in credits proportional to the actual AI work (a simple clip request is a few credits). The free tier gives you 20 credits every day plus a one-time 150-credit welcome bonus.
Start Clipping with Valmera
The agentic AI video editor for long-form and clips alike. Free tier available.