AI Agent That Edits Video
Yes, this exists, and you are looking at one. You upload footage you actually shot, type the change you want in ordinary English, and an AI agent does the work: it reads and watches the video first, makes the cuts, writes and times the captions, places the music, reframes the picture, renders a preview, looks at the frames it just produced, and hands you a finished MP4 cut from your original file.
What follows is the unglamorous version of that sentence. One real request, followed from the keystroke to the download. What the agent knows about your footage and how it learned it. What it does when it is unsure. The things it genuinely cannot do. What an edit costs. And what happens when it gets something wrong, which it will. If you would rather read the same product described in category language, that page is agentic video editor.
Give It Something to Edit
50 free credits on every account, no card. Upload a real video, type one sentence, watch what comes back.
Start free →Yes — and Here Is Exactly What It Can and Cannot Do
These are requests typed the way people actually type them. The ones marked with a cross are not features we are about to launch; they are things the agent will decline to your face rather than attempt badly.
| If you ask for… | Does it? | What actually happens |
|---|---|---|
| “Cut the dead air out of this 40-minute recording” | ✓ | Silences are found before the edit starts; cuts land on word boundaries, not mid-syllable. |
| “Take out every um and every you know” | ✓ | Filler words are removed from the picture and the audio together, including words you name yourself. |
| “Caption it, karaoke style, nice and big” | ✓ | Captions are timed from word-level timestamps and burned into the frame. |
| “Give me a vertical version for Shorts” | ✓ | Reframes to 9:16, 1:1 or 4:5 with a subject-aware crop, and rescales the captions to the new frame. |
| “I said that line three times — keep the best one” | ✓ | Repeated takes are detected across the transcript so the retries can be dropped. |
| “Put something calm under my voice” | ✓ | Music from the built-in library, your own file or a link, ducked under speech automatically. |
| “Blur the licence plate in the driveway shot” | ✓ | The blurred region follows the footage through every later cut and reframe. |
| “Generate me a 60-second ad from this script” | ✗ | This is not a text-to-video generator. It edits footage you shot; it does not invent a whole video. |
| “Export the subtitles as an SRT” | ✗ | Captions are burned in. No subtitle file goes in and none comes out. |
| “Crossfade between these two shots” | ✗ | Transitions are hard-cut styles — dips, whips, a zoom punch. There is no true dissolve. |
| “Post it to my TikTok when it is done” | ✗ | It renders a file and hands it to you. Publishing and scheduling are yours. |
How One Sentence Becomes a File
The part worth understanding is the arrow that points backwards. Most tools that call themselves AI editors run a pipeline: input goes in, output comes out, and whether the output was any good is your problem. An agent closes that loop on itself — it renders what it made, looks at the actual frames, and goes back to the decision list when they are wrong. Everything else on this page follows from that one arrow.
Stage 3 is the one with consequences. The agent does not modify frames; it writes an edit decision list (EDL) — the same object a professional edit has always been, and the same three letters a Premiere or Resolve editor would use for it: a record of which ranges of which source are kept, in what order, with what applied to them. Every guarantee further down this page is a consequence of that: it is why nothing is destructive, why any cut can come back, why the export is source quality, and why an agent can safely be handed the only copy of a shoot.
One Request, From the Keystroke to the Download
Say you have a 41-minute talking-head recording. One camera, no B-roll, a lav mic, several fluffed lines you restarted, and a long stretch in the middle where you looked something up on your phone in silence. This is the most common thing anyone uploads.
Before you type anything: the upload is read
Files up to 14 GB or 3 hours, as MP4, MOV, MKV or WebM. Indexing runs once, with visible progress, and every edit you ever make on that project reuses it. This is a deliberate up-front cost: an agent that has not read your footage can only guess at timestamps, and a guessed cut lands in the middle of a word. Details are in the upload documentation.
The sentence
“Cut the silences and the ums, drop the takes where I restarted, caption it karaoke style, put something calm under my voice, and give me a vertical cut for Shorts.” Nobody writes that in tool language. It is five jobs with dependencies running between them, expressed the way you would say it to a person.
The ordering, which is the actual work
Those five jobs cannot be done in the order you said them. The cuts have to happen first, because they change the timeline everything else is measured against. The captions must be timed against the edited program, not the raw recording — caption a 41-minute file and then remove nine minutes from it and every caption after the first cut is wrong. The music has to be fitted to a length that nothing in the system knows until the cuts are finished. The reframe has to come after the captions exist, so they can be rescaled into the narrower frame instead of running off the sides. Sequencing that correctly is not a nicety; it is the difference between a finished video and four features that fought each other.
The stringout
In a traditional edit, the first thing anyone builds is a stringout: every usable take laid end to end, before anyone worries about polish. What the agent produces at stage 3 is exactly that object, arrived at from the other direction — it starts from the whole recording and removes, rather than starting from nothing and assembling. The dead air goes. The filler words go, matched in the transcript and removed from picture and sound together. The three attempts at the line about pricing are recognised as three attempts, and two of them go. Every cut is snapped to a word boundary, so nothing clips a consonant.
The preview, and the agent watching it
A preview renders from a fast proxy so it comes back quickly, and the agent looks at real frames of it. This is the step that has no equivalent in a feature-based editor. If a caption is sitting across your mouth in the vertical crop, that is a visual fact, invisible in any transcript and invisible in any log — the only way to know is to look at the frame, which is what it does before it tells you the edit is done.
Your correction
You watch it and something is off. “The cuts are too tight, leave me a breath at the end of each sentence.” That is not a manual fix you perform; it is another full pass. The agent adjusts the keep list, re-times nothing else because the captions follow the edit automatically, re-renders, and looks again. Two or three of these is normal. It is where your taste gets into the video, and there is no way to skip it, because the agent has no way of knowing you like loose cuts until you say so.
The export
The final render goes back to your original upload and cuts from it directly. Not from the proxy the previews used, and not from a re-compressed intermediate: the export is the first time the source file is read at full quality, which is why it takes longer than the preview did and why the result does not look like a copy of a copy. Out comes an H.264 MP4 you download. Free-plan exports carry a small Valmera mark in the corner; on Creator, Pro and Frontier there is no mark at all, and upgrading re-renders an already-marked export clean. Every export, on every plan, ends with a brief Valmera end card after your content finishes.
What the Agent Knows About Your Footage, and How It Learned It
“The AI understands your video” is the least falsifiable sentence in this industry, so here is the specific list of what gets extracted during indexing and what each item makes possible.
- A transcript with a timestamp on every single word, plus labels for who is speaking when more than one person talks. Word-level timing is what makes “cut from where I start talking about pricing” a real range rather than an approximation, and it is what karaoke captions are built from — you cannot highlight a word as it is spoken if you only know which sentence it was in.
- Every silence, with the words on either side of it. Finding a gap is the easy half. Knowing which words border it is what lets a cut close cleanly instead of swallowing the breath before the next sentence.
- Where the shots change. Without this an editor cannot tell a jump cut inside one continuous take from a genuine scene change, which is how tools end up dropping forty transitions into a single unbroken shot.
- Labelled frame tiles of every clip in the project, refreshed on every turn. This is the part that is unusual. Rather than one written summary of the video, generated once by another model and then trusted forever, the agent is handed actual still frames with their timestamps and reads them itself, each time it does anything. It can also call for more frames of any moment, or of the assembled program, whenever it needs to check something — and it does exactly that before aiming a zoom, placing text, or disagreeing with you about what is on screen.
- Measured audio: tempo, the beat grid, and where your voice puts its stress. Beat-aligned cuts and automatic punch-ins on emphasised words are only honest if the beat and the stress were measured from the waveform rather than guessed from the words.
None of this is generated per request. It is computed once when you upload and reused by every edit after it, which is why the second change to a project is faster than the first.
What It Does When It Is Not Sure
An agent that guesses confidently is worse than one that does nothing, because you cannot tell the difference from the reply. Four behaviours cover the uncertain cases.
It asks. Asking you a question is a tool it can call, the same as making a cut. If “remove the bit about the launch” matches three separate passages, the honest move is to say which three and let you pick, not to delete the longest one and report success.
It looks before it aims. Anything with a position in the frame — a zoom target, a blur region, where a lower third sits — is decided from frames it has actually pulled, with coordinates read off the picture. A coordinate is a measurement, not an impression.
It refuses out loud. Ask for a crossfade and you are told there is no true dissolve, and what the available transitions are instead. Ask it to generate a video from a script and you are told this is an editor, not a generator. A refusal you can read is worth more than a silent approximation you have to catch yourself.
It cannot overstate what it did. Every reply is checked server-side against the edit decisions that were actually recorded before it reaches you. If the agent describes a change that is not in the record, the claim does not survive; if nothing changed, the reply says nothing changed. When you are reading a report instead of watching your own hands on a mouse, this is the property that makes the report usable.
What It Cannot Do
Named individually, because “some limitations apply” is not information. If one of these is the thing you needed, you have saved yourself a signup.
- It does not generate video from a prompt. There must be real footage. It can splice in a generated still, a generated sound effect or a short generated clip inside an edit, but there is no path from a script to a finished video, and a project with nothing in it cannot be edited into one.
- No subtitle files, in or out. Captions are burned into the picture. There is no SRT or VTT import and no export, and no chapter metadata.
- No true crossfade or dissolve, and one transition style applies to the whole video rather than being chosen per cut.
- Nothing is motion-tracked. A blur or an overlay holds a region and follows the footage through later cuts, but it does not chase a moving face across the frame.
- No uploaded fonts and no stored brand kit. You get the bundled families and whatever colours you name, every time.
- Audio repair is not in scope. No denoise, no “studio sound”, no levelling each speaker separately, and no pulling music back out of a mixdown that was already baked with the speech. Original music generation is not offered either.
- One deliverable per request. There is no batch mode that returns ten clips from one recording in a single pass; five clips are five requests.
- It does not publish. No posting or scheduling to YouTube or TikTok, no share links. It produces a file and hands it over.
- Single-account. No team seats, no collaboration, no review workflow.
- No native mobile app, though mobile browsers work.
- It never upscales, and slow motion duplicates frames rather than synthesising new ones, so a heavy slowdown on 30fps footage will look like a heavy slowdown on 30fps footage.
- English is the best-tested transcription path. Other languages are not blocked, but the accuracy that everything transcript-driven depends on is not equally proven across them.
- Every export ends with a brief Valmera end card after your content, on every plan. Free-plan exports additionally carry a small mark in the corner during the video; paid exports do not, and previews are never marked on any plan.
The honest general limit sits above all of those: an agent can perform an edit, but it cannot know your taste until you tell it. The first pass is a proposal. Treating it as a draft to react to, rather than a lottery ticket to accept or reject, is the whole skill of working with one.
What a Real Edit Costs
Every account starts with 50 credits, no card, and they are spendable on your own footage rather than on a sample project. Paid plans are Creator at $30/month for 2,000 credits, Pro at $50/month for 4,000, and Frontier at $100/month for 10,000 on a stronger model for both reasoning and vision. Each paid plan opens with a 3-day free trial, and subscribers also get 20 daily credits that are spent before anything you paid for.
There is no per-edit price list, because a flat price would be a lie in one direction or the other. A turn charges in proportion to the AI work actually done in it: recolouring your captions or nudging the music volume is a few credits, while cutting silences across a long recording, captioning all of it and reframing it in one request costs meaningfully more, because considerably more happened. Generation inside an edit is priced the same honest way, so a still image and a ten-second clip are not charged alike.
Previews and exports are renders of decisions you already made and paid for, so they are not charged again. The practical consequence is that watching your preview properly before asking for the next change is the cheapest habit available. The full breakdown is in the credits documentation, and the plans are on the pricing page.
What Happens When It Gets It Wrong
It will get things wrong. The question that matters is what a wrong edit costs you, and the answer is a sentence, because nothing the agent does is destructive.
- Your upload is never modified. Not on the first edit, not on the fiftieth. Every render reads it; nothing writes to it.
- The decision list is versioned. Each operation produces a new version rather than overwriting the previous one, so the history of the edit is a record and not a rumour.
- Anything cut can come back. Ask for a removed range to be restored and it is restored. Cuts are entries in a list, not deletions from a file.
- Anything applied can be removed by name. A zoom, an overlay, a blur, a music bed, a speed change — each comes off individually without disturbing the rest of the edit.
- The whole edit can be reset. One request returns the project to the untouched upload, which is occasionally the fastest route out of an edit that went somewhere strange.
None of this is generosity. It is what makes the loop at the top of this page workable at all: an edit you cannot cheaply reverse is an edit you would never let an agent attempt, and you would end up asking for the timid version of everything. Reversibility is what buys the agent permission to try things.
How This Compares to Opening Premiere
The comparison people actually have in mind is not another AI tool, it is the non-linear editor they either use or avoid — Premiere Pro, DaVinci Resolve, Final Cut Pro. These are different classes of object and it is worth being blunt about where each one wins.
| AI agent (Valmera) | Premiere / Resolve / Final Cut | |
|---|---|---|
| Who performs the cuts | The agent | You |
| Where it runs | A browser tab | A workstation, with the media on a local disk |
| Knows what is in the footage before you start | ✓ | ✗ |
| Multi-cam sync, nested sequences, dedicated grading pages or panels | ✗ | ✓ |
| Colour management, LUTs and RAW workflows | ✗ | ✓ |
| Frame-by-frame manual control | Partial — timeline, player, editable transcript | Total |
| Turning one recording into a vertical cut | One sentence | A second sequence, built by hand |
If your work involves multi-cam interviews, colour-managed RAW, or a client who sends frame-accurate notes, a conventional NLE is still the correct tool and no amount of agent will change that. If your work is a recording of yourself talking that needs to be tightened, captioned, scored and cut vertical — which is most video made by most people — describing the outcome is a faster route to the same file. The two are not really competing for the same afternoon.
Where an Editing Agent Fits in the Category
In “It’s time for agentic video editing” (Justine Moore, a16z, 21 January 2026), the work an editing agent could take on is split into five parts: Process, Orchestrate, Polish, Adapt and Optimize. It is a useful frame precisely because it makes gaps visible, so here is where this agent sits against it, including the parts where it does nothing.
- Process — sorting through the footage and deciding what to use. This is the core of it: indexing, silence and filler detection, repeated-take detection, and cutting from what was said.
- Orchestrate — partial, and honestly so. Generated images, generated sound effects, short generated clips and media pulled from a link can all be placed inside an edit, but this is not a conductor for a fleet of external models. The centre of gravity is your footage.
- Polish — partial. Captions, music that ducks under speech, loudness mastering, colour grades, zooms, punch-ins on emphasised words and finishing effects are all here. The audio-repair half of that bucket is not: no denoise, no “studio sound”, no levelling each speaker separately.
- Adapt — partial. Reframing to 16:9, 9:16, 1:1 and 4:5 with the captions rescaled, so one recording becomes the version each platform wants. See aspect ratio if that is unfamiliar territory. The other half a16z means by this word — translating and dubbing a video into other languages — is not offered.
- Optimize — not offered, and this is the honest gap. The piece described there is taste: an agent exercising editorial judgement on its own and handing you a few draft versions to react to. One request here produces one cut, and the judgement arrives through your corrections rather than ahead of them. It is the part of the frame worth being sceptical about whenever a tool claims to have closed it.
Programmatic access, for anyone who wants to build on top of it, is the MCP server described below rather than a separate REST product.
The Same Agent, Running Inside Claude
If the agent is the editor, then the agent is replaceable — which is a strange sentence until you see what it buys you. All 108 tools are published as a remote Model Context Protocol server (97 for editing, 11 for projects, uploads, indexing, rendering, export and download) over streamable HTTP, with OAuth 2.1, dynamic client registration and PKCE, so there is no API key to copy anywhere. Add it as a connector in the Claude app, sign in, and edit your video from the conversation you were already having.
The tools are not a re-declared subset written for the connector. The editing engine publishes one registry and the connector serves it verbatim, so the model in your Claude session gets what the in-house agent gets: the same names, the same schemas, the same refusals. Long operations return a job id and a tool to wait on it rather than a fabricated completion, and editing the same project from the studio and from Claude at the same time is refused in both directions instead of quietly racing.
Setup, including the errors worth recognising, is on the Claude connection page; the complete tool list is on the tool reference. ChatGPT is a narrower fit — its read-only connector path cannot drive an editor at all, and the developer mode that does load full read-and-write tool surfaces is documented as web-only and on paid accounts. That path, and the four places it stops, is written up on the ChatGPT page.
How to Get an AI Agent to Edit Your Video
- 1Upload the footage you shotUp to 14 GB or 3 hours per video, as MP4, MOV, MKV or WebM. It is indexed once with visible progress — word-level transcript, silences, shot changes and frame tiles — and every later edit reuses that index.
- 2Describe the finished video, not the stepsType what you want to be true at the end: “cut the dead air and the ums, caption it karaoke style, put calm music under my voice, and give me a vertical cut”. The agent works out the order the jobs have to happen in.
- 3Watch the preview and say what is offThe agent renders a preview and inspects the frames it produced. React in plain English — “leave a breath after each sentence”, “captions lower, they are on my face”, “swap the music for something with drums”. Each correction is another full pass, not a manual repair you perform.
- 4Export from the originalWhen it is right, export. The final render is cut from your original upload at source quality and downloads as an H.264 MP4. Nothing about the upload has been altered, and anything you cut can still be restored.
Indexing is a one-time cost per upload and scales with length; editing turns are short by comparison, and the final export takes longer than a preview because it reads the original file rather than the proxy.
Frequently Asked Questions
Keep Reading
If you want the individual jobs rather than the whole workflow: removing silence, karaoke captions and resizing for a platform each have their own page. For the wider toolset there is the AI video editor overview, and for how this compares to the other tools claiming an agent, the agentic field. If you have not used the product before, getting started is four paragraphs long.
Hand It a Video and See
50 credits on signup, no card. Or connect it to Claude and edit from the chat you already have open.
Start free →