AI Captions & Subtitles
Word-accurate captions generated straight from your video's transcript — styled, positioned, and animated by asking in plain English.
How Valmera Captions Work
Valmera generates captions from a word-level transcript of your footage, so every caption appears exactly when the word is spoken. When you upload a video, the system transcribes it with Deepgram nova-3 speech recognition (with a whisper fallback), producing a timestamp for each individual word — not just each sentence.
Because captions come from the same transcript the agent uses for cuts and silence removal, they stay in sync no matter how much you edit. Cut a rambling minute out of the middle and the captions simply follow the new timeline — there is nothing to re-time by hand.
Caption Styles You Can Ask For
11 PRESETS
Podcast, beast, karaoke, elegant, stacked, iridescent, chrome, editorial, fashion, luxe, and impact — a complete look in one word.
12 FONTS
Inter Display (Black, ExtraBold, Bold), Anton, Bebas Neue, Archivo Black, Poppins Black, Syne ExtraBold, Playfair Display Black, Instrument Serif, DM Serif Display, and Montserrat.
COLOR
Any caption color you name — white, yellow, brand red — plus a separate highlight color for karaoke word-pop.
SIZE & POSITION
Four named sizes from S to XL plus a continuous scale (“a touch bigger” works), at the bottom, middle, or top of the frame.
9 ANIMATIONS
Caption entrances: fade, pop, slide-up, punch, blur-in, whip, flash, rise, and drop.
KARAOKE WORD-POP
Each word pops as it is spoken, up to 6 words per line — the style short-form viewers expect on TikTok, Reels, and Shorts.
PER-WORD EMPHASIS
Make individual words land: big, huge, accent, pop, box, serif, chrome, glow, or chroma — with an adjustable emphasis scale.
WORDS PER CAPTION
Anywhere from 1 to 16 words on screen at a time — one word at a time for punchy verticals, longer lines for talking-head video.
ALWAYS IN SYNC
Word-level timestamps mean captions land on the exact word — even after heavy cutting.
The 11 Caption Presets
A preset sets the font, weight, colors, emphasis behaviour, and animation together, so you can get a finished caption look by naming it — "use the beast preset" — instead of specifying six settings. Every part of a preset stays adjustable afterwards: ask for a different color, a smaller size, or another font and only that changes.
Podcast
Clean, readable lines built for long-form talking-head video.
Beast
Loud, oversized, high-contrast — the aggressive short-form look.
Karaoke
Word-by-word highlight as each word is spoken.
Elegant
Restrained and typographic, for calmer edits.
Stacked
Short lines stacked for vertical framing.
Iridescent
Colour-shifting highlight treatment.
Chrome
Metallic, glossy word styling.
Editorial
Magazine-style captions with serif weight.
Fashion
Sparse, high-style lettering.
Luxe
Premium look with accent-coloured emphasis.
Impact
Heavy display type for maximum punch.
The same 12 fonts and animation set also power on-screen text and overlays — titles, lower thirds, callouts, and quotes.
Edit the Transcript, Captions Follow
Speech recognition occasionally misses a product name or an unusual word. You do not fix that caption by caption — Valmera has an editable transcript panel. Click the sentence, correct the text, and the captions re-render with the fix everywhere that word appears.
This is also how Valmera's honesty layer applies to captions: the agent's reply is checked against what was actually rendered, so it will never claim it styled or fixed captions it did not touch. If a request could not be done, the reply says so plainly.
Example Requests
"Add karaoke captions, large, bottom"
The classic short-form setup: word-pop captions, size L, anchored at the bottom of the frame.
"Make the captions yellow and move them to the top"
Restyle existing captions without touching anything else in the edit.
"Captions a bit smaller, with a slide-up animation"
The continuous size scale means “a bit smaller” is a real instruction, not a guess.
"Use the beast preset in Anton, three words at a time"
A preset, one of the 12 bundled fonts, and the words-per-caption count in a single sentence.
"Emphasize the key words with a glow and make them bigger"
Per-word emphasis — the agent picks the words that carry the line and styles those only.
"White captions, middle of the frame, pop animation"
Combine color, position, and animation in one request — the agent applies all three.
Honest Limits
A few things Valmera captions do not do today, so you are never surprised:
- No font uploads — you choose from the 12 bundled fonts, you cannot supply your own typeface file.
- No stored brand kits — caption styling is set per project, not saved as a reusable brand preset.
- No stickers or emoji decorations attached to captions.
- No SRT/VTT subtitle upload or export — captions are burned into the exported video and always come from the transcript, which you can edit directly.
- Transcript captions need spoken audio from a main video, so they do not apply to canvas projects started without one — use on-screen text there instead.
Frequently Asked Questions
How accurate are Valmera's AI captions?
Captions are generated from a word-level transcript produced by Deepgram nova-3 speech recognition, with a whisper fallback, so every caption is timed to the exact word being spoken. If a name or technical term is transcribed wrong, you can correct it in the editable transcript panel and the captions re-render with the fix.
What caption styles can I choose?
There are 11 style presets — podcast, beast, karaoke, elegant, stacked, iridescent, chrome, editorial, fashion, luxe, and impact — plus 12 bundled fonts, any colors you name (including the karaoke highlight color), sizes from S to XL or anywhere on a continuous scale, bottom/middle/top placement, 1 to 16 words per caption, and 9 entrance animations: fade, pop, slide-up, punch, blur-in, whip, flash, rise, and drop.
Can I emphasize specific words in the captions?
Yes. Per-word emphasis lets individual words stand out with big, huge, accent, pop, box, serif, chrome, glow, or chroma treatments, and you can dial the emphasis scale up or down. Ask for something like "make the key words pop" and the agent applies emphasis to the words that carry the line.
Can I upload my own SRT or VTT subtitle file?
No. Valmera does not accept SRT or VTT uploads, and does not export a subtitle file either — captions are burned into the video. They are always generated from the word-level transcript of your footage, which keeps them perfectly synced. If the transcript has an error, edit that sentence in the transcript panel and the captions update.
Can I use my own brand font?
You cannot upload a font file, but you are not stuck with one typeface either: 12 fonts are bundled and any of them can be used for captions and on-screen text — Inter Display in Black, ExtraBold and Bold, Anton, Bebas Neue, Archivo Black, Poppins Black, Syne ExtraBold, Playfair Display Black, Instrument Serif, DM Serif Display, and Montserrat. Name the one you want in your request.
Do captions work on vertical and square video?
Yes. Captions scale to the frame in every supported aspect ratio — 16:9, 9:16, 1:1, and 4:5 — so a video reframed for Shorts or Reels keeps readable, correctly placed captions.
How much does adding captions cost?
Edits charge credits in proportion to the actual AI work in the turn, and a straightforward caption pass is a few credits. The Free plan includes 20 credits per day plus a one-time 150-credit welcome bonus; Plus is $20/month for 800 monthly credits and Pro is $50/month for 2,400.