AI Captions: Edit Text, Timing & Styles in Valmera
Generate captions from your footage, correct the words, choose a look and review the result. Separate a text problem from a timing or layout problem so the next change fixes the right thing.
Documentation reviewed .
Add Captions and Review Them
- Open the intended project and version. Use footage with speech and wait for its transcript. Check the project and selected edit before changing captions. A canvas without a speech-bearing main video needs authored on-screen text instead of transcript captions.
- Check the words against the audio. Listen to names, numbers, negatives and overlapping speech. Correct the intended transcript line before styling it. Keep the wording faithful to the recording; changing text does not change what the speaker says.
- Request a caption style. Name the look, color, position and approximate amount of text to show together. Start with one clear request, such as podcast captions in white near the bottom, then inspect the result.
- Review the revised preview. Play the beginning, cut joins and difficult phrases. Check whether each caption belongs to the words currently heard, whether line breaks make sense and whether the text stays on screen long enough to read.
- Check each output shape. Review landscape, vertical or square versions separately. Check faces, graphics, frame edges and the destination player's controls. Ask for a specific size or placement correction where needed.
- Export and inspect the complete file. Create the final MP4 in Studio, download it and inspect the captions in that file. Captions are burned into the picture. Studio does not provide SRT/VTT import or export; arrange a separate workflow if those files are required.
The first-edit guide covers upload and project readiness. For a short social clip, stabilize the kept footage and target shape before spending time on detailed typography.
Fix the Text at the Right Level
| Problem | Use | Review afterwards |
|---|---|---|
| A transcript sentence is wrong or needs a different word count | Edit transcript line, then save the corrected sentence | Refreshed text, word timing and the completed caption preview |
| A displayed spelling or capitalization needs a replacement with the same word count | Ask for a caption spelling fix; the relevant tool is set_caption_fixes | Every affected occurrence and its surrounding phrase |
| The words are right but too large, misplaced or over-animated | A caption-style request | The complete shot in each output shape |
| The recording itself says the wrong thing | A content decision: retain, cut or replace the source passage | Agreement between delivered audio and captions |
A transcript-line edit targets that sentence. It is not a global replacement of every matching word in the project. Caption spelling replacements are a separate operation; review their scope before using a common word as the search term. Manually authored caption items also need their own text changed.
Illustrative correction: the recording says “Try Valmera before Friday.” If the transcript says “Try Val Mira before Friday,” edit that sentence: “Val Mira” has two words, while “Valmera” has one. A same-word-count caption replacement is not the right operation. A capitalization change from “valmera” to “Valmera” keeps one word on each side and fits that narrower spelling operation.
After a transcript edit, word positions may be reconstructed within the sentence span. Listen again instead of assuming the corrected spelling proves exact timing. A preview or final downloaded before the correction remains an older file.
Why Source Time and Caption Time Differ After Cuts
The transcript describes the source recording. The viewer sees only the retained material, in its edited order. At normal speed with hard cuts, no inserts and no overlaps, a word's output time is the duration already kept plus its offset inside the current source section.
| Kept source range | Kept duration | Output range |
|---|---|---|
| 00:10.0–00:14.0 | 4.0 seconds | 00:00.0–00:04.0 |
| 00:20.0–00:24.0 | 4.0 seconds | 00:04.0–00:08.0 |
Hypothetical timing example: a word at source 00:21.2–00:21.6 appears at output 00:05.2–00:05.6: 4.0 seconds already kept, plus offsets of 1.2 and 1.6 seconds. It should not remain at output 21.2 seconds. This calculation illustrates the clocks; it is not a measured captioning result.
Speed changes, added scenes and other timing operations change the mapping. When reporting a fault, identify the project version and whether the timestamp refers to source or preview. Include the words just before and after the problem. See timeline and transcript controls.
Choose a Look, Then Check Readability
Named looks include podcast, beast, karaoke, elegant, stacked, iridescent, chrome, editorial, fashion, luxe and impact. Treat these as examples of supported presets, not a fixed count of every option. Start from the look that suits the footage, then ask for a specific revision.
| Control | Examples | Decision to make |
|---|---|---|
| Bundled typefaces | Inter Display Black, ExtraBold or Bold; Anton; Bebas Neue; Archivo Black; Poppins Black; Syne ExtraBold; Playfair Display Black; Instrument Serif; DM Serif Display; Montserrat | Choose a face that remains legible at the intended viewing size; custom font uploads are unsupported |
| Color and emphasis | White text, a named highlight color, emphasis on selected words | Preserve contrast and meaning; highlighting a number should not obscure its unit |
| Size and position | S, M, L or XL; a size adjustment; bottom, middle or top | Keep the text in frame and away from faces, graphics and player controls |
| Grouping | One word at a time or short groups; normal caption groups can use 1–16 words | Prefer readable phrases over maximizing the configured word limit |
| Motion | Karaoke word-pop; entrances such as fade, pop, slide-up, punch, blur-in, whip, flash, rise or drop | Inspect motion in playback; preset and legacy word-pop paths can have different limits |
A font change can wrap a line differently. A larger caption can cover a face. A reframe can move the usable area. Check each delivered aspect ratio, especially long names and the busiest shot, rather than assuming scaled text always fits. Read the reframing guide for the picture itself.
Requests That Make Corrections Easier
- “Add podcast captions in white near the bottom. Keep them clear of the speaker's face.”
- “In the captions, capitalize Valmera. Preserve the wording and review the affected occurrences.”
- “At preview 00:05.2 in this version, the caption arrives late on ‘Valmera.’ Inspect that phrase and the preceding cut.”
- “In the vertical version, make the captions smaller and lift them above the destination's bottom controls. Preserve the approved cuts and music.”
These are example briefs. Inspect the resulting version before considering the correction finished. For titles, lower thirds and silent footage, use authored text and overlays.
Burned-In Captions and Separate Subtitle Files
Valmera Studio burns captions into the MP4 picture. Viewers cannot toggle that burned text off. Studio does not import or export SRT/VTT files. If a client needs a separate caption track, editable subtitle file or multilingual player choices, confirm an additional delivery workflow before editing.
The free subtitle timing utility adjusts an existing subtitle file. It does not generate a Studio transcript or convert the Studio export into a separate subtitle track.
W3C's caption guidance explains that captions communicate speech and meaningful non-speech audio, and automatic text needs accuracy review. Check speaker identification and important sounds as well as words. A stylish transcript overlay alone does not establish that an accessibility requirement has been met.
Use the export guide for version selection, the final MP4 and its complete runtime. Use current plans and credits documentation for account access.
Caption Questions
How do I add AI captions in Valmera?
Upload footage with speech, wait for its transcript, then ask for captions with a named look and placement. Review the text, timing and line breaks in the preview. Export the accepted version in Studio and check the downloaded MP4.
How accurate are Valmera's captions?
They start from speech recognition and word timestamps. Names, noise, accents and overlapping speech can cause errors. Review against the audio and correct the result; this guide does not claim a measured accuracy percentage or perfect synchronization.
How do I correct a transcript line?
Open the Transcript panel, use Edit transcript line on the affected sentence, change its text and save. Studio refreshes the transcript and requests a new preview when transcript captions are active. Confirm that the new preview has finished and shows the correction; an older downloaded file does not update.
What is the difference between transcript editing and a caption spelling fix?
A transcript edit changes the selected sentence used for future transcript-based work. The caption spelling tool applies replacements to displayed transcript captions while retaining their word slots. Its replacement pairs require the same number of words. A correction that changes word count belongs in the transcript-line workflow, followed by a timing check.
Does changing captions change the spoken audio?
No. A text correction does not replace speech or synthesize a new voice. If the speaker actually said the wrong name or number, decide whether to retain it, cut that passage or supply corrected source audio. Keep the caption faithful to the delivered soundtrack.
Do captions stay synchronized after cuts?
Transcript captions use the kept source words and map them to the edited timeline. Their preview timestamps can therefore differ from source timestamps. Inspect cut boundaries, speed changes and inserted material; automatic mapping is not proof that every word or transition is correctly timed.
Can I use karaoke captions or emphasize particular words?
Supported looks include karaoke and other named presets, with color, size, placement, animation and word-emphasis controls. Request the treatment and inspect the rendered result. A named preset and legacy dynamic word-pop can behave differently; do not assume every setting combination has the same grouping or animation.
Can I upload a custom font?
Custom font-file uploads are unsupported. You can request a bundled face such as Anton, Bebas Neue, Poppins Black or Montserrat. Verify the resulting type, line breaks and size in the preview, particularly after changing the aspect ratio.
Can I import or export SRT or VTT subtitles?
Studio does not import or export SRT/VTT subtitle files. Its captions are burned into the exported video. Valmera's separate free subtitle utilities work with existing subtitle files; they do not add subtitle-file delivery to Studio.
Why are my captions clipped or covering a face?
Size and placement depend on the output frame and the shot. Review the actual crop, reduce oversized text or move it, and shorten an awkward group if needed. A vertical reframe does not guarantee that the existing caption layout remains readable or clears the destination's interface.
Can an MCP assistant work on captions?
The checked public catalog includes add_captions, set_caption_style, set_caption_fixes, set_caption_mutes and audit_captions. Use the schemas and results available in the authenticated session. A completed tool call still needs preview review; final export remains in Studio under the checked public contract.
Is AI captioning free?
Account creation and uploads are free, while indexing and AI editing require a paid subscription and use credits. The amount depends on processing and revisions. Check current plans and account usage rather than assuming a fixed price per captioned video.
This guide is based on Valmera's documented tools and implementation, with the public catalog checked on September 5, 2026. It does not report a new authenticated captioning test. The MCP reference lists advertised tools; inspect the actual session for available schemas and results.
← Back to Docs