A lot of social video is watched on mute — captions aren't decoration, they're the difference between a view and a skip. Here's how to caption any video in one request, styled exactly how you want.
Caption a Video Now
Upload it and ask. Word-accurate, styled, burned in. Free to try.
The old cost was real: typing captions by hand or wrestling subtitle files and timing. With an agentic editor, captions are one line of a request — and they arrive styled, timed, and burned in. If your captioned clips start life as long recordings, see how to turn long videos into clips.
What Caption Styles Does Valmera Support?
Valmera captions support karaoke word-pop, fade / pop / slide-up entrance animations, sizes from S to XL on a continuous scale, bottom / middle / top placement, and any color — all timed word-by-word from a Deepgram nova-3 transcript.
Karaoke word-pop
Each word highlights the instant it's spoken — the bold social-clip look.
Entrance animations
Captions fade, pop, or slide up into the frame.
Sizes S–XL + in between
Named sizes plus a continuous scale — "a bit bigger" works.
Position control
Bottom, middle, or top of the frame.
Any color
Name the color in the request and it renders that way.
Word-accurate timing
Word-level timestamps from Deepgram nova-3, with a whisper fallback.
Every option is a sentence, not a settings panel — the full reference lives in the captions docs. And because every agent reply is verified against what the system actually rendered, the agent can't tell you it captioned a video it didn't touch — if a request fails, the reply says so.
The Workflow
1
Upload the video
Any talking video up to 2GB: a Reel, a tutorial, a podcast clip, a product demo. MP4, MOV, MKV, or WebM.
2
Describe the caption style
Color, size (S–XL), position (bottom, middle, top), karaoke word-pop, and an entrance animation — fade, pop, or slide-up. Copy an example request below to start.
3
Review the preview
Check accuracy and look. The agent transcribed, timed, styled, and rendered in one pass — previews use a fast proxy, and the final export renders from your original full-quality file.
4
Fix anything by talking — or in the transcript
"Make the captions 20% bigger." "Move them to the middle." A wrong word gets fixed right in the editable transcript panel, and the captions re-render.
5
Export
Download the captioned MP4, ready for any platform. Ask for a 9:16 vertical version in the same session if you need one.
Example Requests to Copy
Social-style captions
Add karaoke captions to this clip — bold white, size L, at the bottom of the frame, with each word popping as it's spoken.
Clean YouTube captions
Add clean, readable captions to this video. White text, size M, bottom center, with a subtle fade-in. Nothing flashy — it's a tutorial.
Fix a word everywhere
The captions say 'Val Mira' at 1:32 — it's 'Valmera'. Fix it everywhere it appears and re-render the captions.
Got a Word Wrong? Edit the Transcript
You never juggle SRT or VTT files with Valmera — the transcript itself is the subtitle source, and it's editable. Open the transcript panel, fix the word the model misheard — a product name, a technical term, a mumbled number — and the captions re-render with the correction, already timed to the audio. No exporting, re-syncing, or re-uploading a subtitle file.
Captions also pair naturally with a reframe: ask for a 9:16 version and the same captions render on the vertical cut. A straightforward captioning request costs a few credits, and the free tier's 20 daily credits cover it — see plans and pricing for the Plus and Pro tiers.
Never Type a Caption Again
Styled, word-accurate captions from one request. Free tier — 20 daily credits.
Upload the video to Valmera (up to 2GB — MP4, MOV, MKV, WebM) and ask: "add captions." The system builds a word-level transcript with Deepgram nova-3 speech recognition, then the agent styles the captions and burns them into the video. Add style direction in the same request — "white karaoke captions, size L, bottom of the frame."
Captions are timed from a word-level transcript generated by Deepgram nova-3 (with a whisper fallback), so each caption lands on the exact word being spoken. If a name or technical term comes out wrong, fix it in the editable transcript panel — or tell the agent "at 1:32 it's 'Valmera', not 'Val Mira'" — and the captions re-render with the correction.
Yes. You choose the color, a size from S to XL (or anywhere on a continuous scale between them), and the position — bottom, middle, or top of the frame. You can turn on karaoke-style word-pop, where each word highlights the moment it's spoken, and animate captions in with fade, pop, or slide-up.
On social feeds a large share of viewers watch with the sound off — captions are the difference between a skip and a view. They also make videos accessible to deaf and hard-of-hearing viewers and easier to follow for non-native speakers, on every platform.
No — Valmera generates captions from its own word-level transcript rather than from subtitle files. In practice that replaces the SRT workflow: instead of re-timing a file when a line is wrong, you edit the word in the transcript panel and the captions re-render, already synced to the audio. The export you download has the captions burned in.
Simple edits cost the least — you're charged for the work actually done, and a straightforward captioning request sits at the simple end. The free tier includes 20 credits every day (they reset daily) plus a one-time 150-credit welcome bonus; Plus is $20/month for 800 monthly credits and Pro is $50/month for 2,400.
Caption Everything You Post
Valmera captions, cuts, and polishes from plain-English requests. Free tier available.