Getting Started
Go from raw footage to a finished, full-quality video. No editing experience required.
How Valmera Works
Valmera is an agentic AI video editor. You upload real footage, the system analyzes it — building a word-level transcript, detecting shots, and mapping silences — and then you describe the edit you want in plain English. An AI agent performs it: cuts, captions, music, reframing, and polish. You refine the result with follow-up messages, just like working with a human editor.
You never have to touch a timeline to get a finished video — but there is one. Valmera includes a visual timeline and editable transcript — draggable blocks for inserts and music, cut markers, a scrubbable playhead — for the moments when you want direct control.
One more thing that makes Valmera different: every reply from the agent is checked against what the system actually did. The agent cannot claim an edit it did not make — if nothing changed, the reply says so.
Step-by-Step
Create Your Account
Sign up at valmera.io with your email or Google account. No credit card required — every account starts with 20 free daily credits plus a one-time 150-credit welcome bonus.
Upload Your Footage
Attach the video you want edited — MP4, MOV, MKV, WebM and similar formats, up to 2 GB. Podcasts, screen captures, gameplay, vlogs, and long recordings all work.
Let the Analysis Run
Before any editing happens, Valmera builds a word-level transcript of everything said, detects shot changes, and maps the silences. Analysis runs once per upload — longer files take longer, and you can watch the progress in the Studio.
Describe the Edit
Type what you want in plain English. For example: "Cut the silences, add captions, and tighten the pacing — keep it under 10 minutes." Be as specific or as vague as you like.
The Agent Edits, You Preview
The agent performs the edit and renders a fast preview in the player. Its reply is verified against the actual changes — it can only tell you about edits it really made.
Refine — by Chat or by Hand
Send follow-ups like "Make the intro shorter" or "Use bolder captions" and the agent re-edits. Or take direct control: drag blocks on the timeline, scrub the playhead, or fix a sentence in the transcript panel and the captions re-render.
Export at Full Quality
Previews render from a fast lower-resolution proxy, but Export renders the final video from your original full-quality file. Download the MP4 and post it anywhere.
Full details on formats and size limits are in Uploading Footage & Assets, and how edits are charged is covered in Understanding Credits.
Start Without a Video
You do not need footage to start. In a canvas project you pick an aspect ratio — 16:9, 9:16, 1:1, or 4:5 — and build a sequential timeline out of images, clips, music, and media you generate inside the edit. That is the path for a photo slideshow, a product carousel, or an image-plus-voiceover video.
Everything else works the same way: describe what you want, the agent assembles it, and you refine by chat or by dragging blocks. Stills get Ken Burns motion, on-screen text and overlays are available, and music mixes with automatic ducking.
The honest trade-off: with no main video there is no speech transcript, so canvas projects have no transcript captions, no speed changes, and no censor regions or reframing — details in Timeline & Transcript.
Tips for Better Results
- Say who the video is for. "A punchy 60-second vertical clip" and "a relaxed 15-minute YouTube video" get very different — and better — edits than "edit this."
- Iterate in small steps. Instead of rewriting your whole request, send one change at a time. The agent edits the existing cut rather than starting over.
- Upload your own music. Attach a track (up to 50 MB) and the agent lays it under your video with automatic ducking, so it never drowns out speech.
- Fix words in the transcript. If the transcript got a name or term wrong, edit the sentence in the transcript panel — the captions re-render from your correction.
- Be specific about caption style. Eleven caption presets, 12 bundled fonts, karaoke word-pop, bottom/middle/top position, sizes from S to XL, colors, per-word emphasis, and nine entrance animations are all one sentence away.