← Home
EXPLAINER

Published · Updated

AI Agent That Edits Video

Yes, this exists, and you are looking at one. You upload footage you actually shot, type the change you want in ordinary English, and an AI agent does the work: it reads and watches the video first, makes the cuts, writes and times the captions, places the music, reframes the picture, renders a preview, looks at the frames it just produced, and hands you a finished MP4 cut from your original file.

What follows is the unglamorous version of that sentence. One real request, followed from the keystroke to the download. What the agent knows about your footage and how it learned it. What it does when it is unsure. The things it genuinely cannot do. What an edit costs. And what happens when it gets something wrong, which it will. If you would rather read the same product described in category language, that page is agentic video editor.

Give It Something to Edit

50 free credits on every account, no card. Upload a real video, type one sentence, watch what comes back.

Start free →
See pricing →

Yes — and Here Is Exactly What It Can and Cannot Do

These are requests typed the way people actually type them. The ones marked with a cross are not features we are about to launch; they are things the agent will decline to your face rather than attempt badly.

If you ask for…Does it?What actually happens
“Cut the dead air out of this 40-minute recording”Silences are found before the edit starts; cuts land on word boundaries, not mid-syllable.
“Take out every um and every you know”Filler words are removed from the picture and the audio together, including words you name yourself.
“Caption it, karaoke style, nice and big”Captions are timed from word-level timestamps and burned into the frame.
“Give me a vertical version for Shorts”Reframes to 9:16, 1:1 or 4:5 with a subject-aware crop, and rescales the captions to the new frame.
“I said that line three times — keep the best one”Repeated takes are detected across the transcript so the retries can be dropped.
“Put something calm under my voice”Music from the built-in library, your own file or a link, ducked under speech automatically.
“Blur the licence plate in the driveway shot”The blurred region follows the footage through every later cut and reframe.
“Generate me a 60-second ad from this script”This is not a text-to-video generator. It edits footage you shot; it does not invent a whole video.
“Export the subtitles as an SRT”Captions are burned in. No subtitle file goes in and none comes out.
“Crossfade between these two shots”Transitions are hard-cut styles — dips, whips, a zoom punch. There is no true dissolve.
“Post it to my TikTok when it is done”It renders a file and hands it to you. Publishing and scheduling are yours.

How One Sentence Becomes a File

The part worth understanding is the arrow that points backwards. Most tools that call themselves AI editors run a pipeline: input goes in, output comes out, and whether the output was any good is your problem. An agent closes that loop on itself — it renders what it made, looks at the actual frames, and goes back to the decision list when they are wrong. Everything else on this page follows from that one arrow.

How a plain-English request becomes an exported video fileA typed request enters at the top. The agent indexes the footage into a word-level transcript, silences, shot cuts and labelled frame tiles; turns the words of the request into real timecodes; edits an edit decision list; renders a fast preview; looks at the frames it produced and loops back to the decision list when they are wrong; and finally exports from the original upload at source quality. A separate rail down the left shows the original file passing through untouched to the export.YOU TYPE“cut the dead air, caption it,put music under it, make it vertical”1It reads the footageword-level transcript · silencesshot cuts · labelled frame tiles2It turns your words into timecodes“the pricing bit” → a real rangeno timestamp is ever estimated3It edits the decision listcuts, captions, music, reframedecisions — never your pixels4It renders a previewa fast proxy, never watermarkedso you can watch it immediately5It looks at the frames it madereads its own output backand revises what it got wrong6It exports from your originalsource-quality H.264 MP4the upload itself is untouchedWRONG? REVISE AND RE-RENDERYOUR ORIGINAL FILE, NEVER MODIFIEDThe finished MP4, downloadedcut from the original, at source quality
The request path. Stages 3 and 5 are the loop: the agent inspects what it rendered and edits the decision list again, which is the difference between an agent and a pipeline.

Stage 3 is the one with consequences. The agent does not modify frames; it writes an edit decision list (EDL) — the same object a professional edit has always been, and the same three letters a Premiere or Resolve editor would use for it: a record of which ranges of which source are kept, in what order, with what applied to them. Every guarantee further down this page is a consequence of that: it is why nothing is destructive, why any cut can come back, why the export is source quality, and why an agent can safely be handed the only copy of a shoot.

One Request, From the Keystroke to the Download

Say you have a 41-minute talking-head recording. One camera, no B-roll, a lav mic, several fluffed lines you restarted, and a long stretch in the middle where you looked something up on your phone in silence. This is the most common thing anyone uploads.

Before you type anything: the upload is read

Files up to 14 GB or 3 hours, as MP4, MOV, MKV or WebM. Indexing runs once, with visible progress, and every edit you ever make on that project reuses it. This is a deliberate up-front cost: an agent that has not read your footage can only guess at timestamps, and a guessed cut lands in the middle of a word. Details are in the upload documentation.

The sentence

“Cut the silences and the ums, drop the takes where I restarted, caption it karaoke style, put something calm under my voice, and give me a vertical cut for Shorts.” Nobody writes that in tool language. It is five jobs with dependencies running between them, expressed the way you would say it to a person.

The ordering, which is the actual work

Those five jobs cannot be done in the order you said them. The cuts have to happen first, because they change the timeline everything else is measured against. The captions must be timed against the edited program, not the raw recording — caption a 41-minute file and then remove nine minutes from it and every caption after the first cut is wrong. The music has to be fitted to a length that nothing in the system knows until the cuts are finished. The reframe has to come after the captions exist, so they can be rescaled into the narrower frame instead of running off the sides. Sequencing that correctly is not a nicety; it is the difference between a finished video and four features that fought each other.

The stringout

In a traditional edit, the first thing anyone builds is a stringout: every usable take laid end to end, before anyone worries about polish. What the agent produces at stage 3 is exactly that object, arrived at from the other direction — it starts from the whole recording and removes, rather than starting from nothing and assembling. The dead air goes. The filler words go, matched in the transcript and removed from picture and sound together. The three attempts at the line about pricing are recognised as three attempts, and two of them go. Every cut is snapped to a word boundary, so nothing clips a consonant.

The preview, and the agent watching it

A preview renders from a fast proxy so it comes back quickly, and the agent looks at real frames of it. This is the step that has no equivalent in a feature-based editor. If a caption is sitting across your mouth in the vertical crop, that is a visual fact, invisible in any transcript and invisible in any log — the only way to know is to look at the frame, which is what it does before it tells you the edit is done.

Your correction

You watch it and something is off. “The cuts are too tight, leave me a breath at the end of each sentence.” That is not a manual fix you perform; it is another full pass. The agent adjusts the keep list, re-times nothing else because the captions follow the edit automatically, re-renders, and looks again. Two or three of these is normal. It is where your taste gets into the video, and there is no way to skip it, because the agent has no way of knowing you like loose cuts until you say so.

The export

The final render goes back to your original upload and cuts from it directly. Not from the proxy the previews used, and not from a re-compressed intermediate: the export is the first time the source file is read at full quality, which is why it takes longer than the preview did and why the result does not look like a copy of a copy. Out comes an H.264 MP4 you download. Free-plan exports carry a small Valmera mark in the corner; on Creator, Pro and Frontier there is no mark at all, and upgrading re-renders an already-marked export clean. Every export, on every plan, ends with a brief Valmera end card after your content finishes.

What the Agent Knows About Your Footage, and How It Learned It

“The AI understands your video” is the least falsifiable sentence in this industry, so here is the specific list of what gets extracted during indexing and what each item makes possible.

  • A transcript with a timestamp on every single word, plus labels for who is speaking when more than one person talks. Word-level timing is what makes “cut from where I start talking about pricing” a real range rather than an approximation, and it is what karaoke captions are built from — you cannot highlight a word as it is spoken if you only know which sentence it was in.
  • Every silence, with the words on either side of it. Finding a gap is the easy half. Knowing which words border it is what lets a cut close cleanly instead of swallowing the breath before the next sentence.
  • Where the shots change. Without this an editor cannot tell a jump cut inside one continuous take from a genuine scene change, which is how tools end up dropping forty transitions into a single unbroken shot.
  • Labelled frame tiles of every clip in the project, refreshed on every turn. This is the part that is unusual. Rather than one written summary of the video, generated once by another model and then trusted forever, the agent is handed actual still frames with their timestamps and reads them itself, each time it does anything. It can also call for more frames of any moment, or of the assembled program, whenever it needs to check something — and it does exactly that before aiming a zoom, placing text, or disagreeing with you about what is on screen.
  • Measured audio: tempo, the beat grid, and where your voice puts its stress. Beat-aligned cuts and automatic punch-ins on emphasised words are only honest if the beat and the stress were measured from the waveform rather than guessed from the words.

None of this is generated per request. It is computed once when you upload and reused by every edit after it, which is why the second change to a project is faster than the first.

What It Does When It Is Not Sure

An agent that guesses confidently is worse than one that does nothing, because you cannot tell the difference from the reply. Four behaviours cover the uncertain cases.

It asks. Asking you a question is a tool it can call, the same as making a cut. If “remove the bit about the launch” matches three separate passages, the honest move is to say which three and let you pick, not to delete the longest one and report success.

It looks before it aims. Anything with a position in the frame — a zoom target, a blur region, where a lower third sits — is decided from frames it has actually pulled, with coordinates read off the picture. A coordinate is a measurement, not an impression.

It refuses out loud. Ask for a crossfade and you are told there is no true dissolve, and what the available transitions are instead. Ask it to generate a video from a script and you are told this is an editor, not a generator. A refusal you can read is worth more than a silent approximation you have to catch yourself.

It cannot overstate what it did. Every reply is checked server-side against the edit decisions that were actually recorded before it reaches you. If the agent describes a change that is not in the record, the claim does not survive; if nothing changed, the reply says nothing changed. When you are reading a report instead of watching your own hands on a mouse, this is the property that makes the report usable.

What It Cannot Do

Named individually, because “some limitations apply” is not information. If one of these is the thing you needed, you have saved yourself a signup.

  • It does not generate video from a prompt. There must be real footage. It can splice in a generated still, a generated sound effect or a short generated clip inside an edit, but there is no path from a script to a finished video, and a project with nothing in it cannot be edited into one.
  • No subtitle files, in or out. Captions are burned into the picture. There is no SRT or VTT import and no export, and no chapter metadata.
  • No true crossfade or dissolve, and one transition style applies to the whole video rather than being chosen per cut.
  • Nothing is motion-tracked. A blur or an overlay holds a region and follows the footage through later cuts, but it does not chase a moving face across the frame.
  • No uploaded fonts and no stored brand kit. You get the bundled families and whatever colours you name, every time.
  • Audio repair is not in scope. No denoise, no “studio sound”, no levelling each speaker separately, and no pulling music back out of a mixdown that was already baked with the speech. Original music generation is not offered either.
  • One deliverable per request. There is no batch mode that returns ten clips from one recording in a single pass; five clips are five requests.
  • It does not publish. No posting or scheduling to YouTube or TikTok, no share links. It produces a file and hands it over.
  • Single-account. No team seats, no collaboration, no review workflow.
  • No native mobile app, though mobile browsers work.
  • It never upscales, and slow motion duplicates frames rather than synthesising new ones, so a heavy slowdown on 30fps footage will look like a heavy slowdown on 30fps footage.
  • English is the best-tested transcription path. Other languages are not blocked, but the accuracy that everything transcript-driven depends on is not equally proven across them.
  • Every export ends with a brief Valmera end card after your content, on every plan. Free-plan exports additionally carry a small mark in the corner during the video; paid exports do not, and previews are never marked on any plan.

The honest general limit sits above all of those: an agent can perform an edit, but it cannot know your taste until you tell it. The first pass is a proposal. Treating it as a draft to react to, rather than a lottery ticket to accept or reject, is the whole skill of working with one.

What a Real Edit Costs

Every account starts with 50 credits, no card, and they are spendable on your own footage rather than on a sample project. Paid plans are Creator at $30/month for 2,000 credits, Pro at $50/month for 4,000, and Frontier at $100/month for 10,000 on a stronger model for both reasoning and vision. Each paid plan opens with a 3-day free trial, and subscribers also get 20 daily credits that are spent before anything you paid for.

There is no per-edit price list, because a flat price would be a lie in one direction or the other. A turn charges in proportion to the AI work actually done in it: recolouring your captions or nudging the music volume is a few credits, while cutting silences across a long recording, captioning all of it and reframing it in one request costs meaningfully more, because considerably more happened. Generation inside an edit is priced the same honest way, so a still image and a ten-second clip are not charged alike.

Previews and exports are renders of decisions you already made and paid for, so they are not charged again. The practical consequence is that watching your preview properly before asking for the next change is the cheapest habit available. The full breakdown is in the credits documentation, and the plans are on the pricing page.

What Happens When It Gets It Wrong

It will get things wrong. The question that matters is what a wrong edit costs you, and the answer is a sentence, because nothing the agent does is destructive.

  • Your upload is never modified. Not on the first edit, not on the fiftieth. Every render reads it; nothing writes to it.
  • The decision list is versioned. Each operation produces a new version rather than overwriting the previous one, so the history of the edit is a record and not a rumour.
  • Anything cut can come back. Ask for a removed range to be restored and it is restored. Cuts are entries in a list, not deletions from a file.
  • Anything applied can be removed by name. A zoom, an overlay, a blur, a music bed, a speed change — each comes off individually without disturbing the rest of the edit.
  • The whole edit can be reset. One request returns the project to the untouched upload, which is occasionally the fastest route out of an edit that went somewhere strange.

None of this is generosity. It is what makes the loop at the top of this page workable at all: an edit you cannot cheaply reverse is an edit you would never let an agent attempt, and you would end up asking for the timid version of everything. Reversibility is what buys the agent permission to try things.

How This Compares to Opening Premiere

The comparison people actually have in mind is not another AI tool, it is the non-linear editor they either use or avoid — Premiere Pro, DaVinci Resolve, Final Cut Pro. These are different classes of object and it is worth being blunt about where each one wins.

AI agent (Valmera)Premiere / Resolve / Final Cut
Who performs the cutsThe agentYou
Where it runsA browser tabA workstation, with the media on a local disk
Knows what is in the footage before you start
Multi-cam sync, nested sequences, dedicated grading pages or panels
Colour management, LUTs and RAW workflows
Frame-by-frame manual controlPartial — timeline, player, editable transcriptTotal
Turning one recording into a vertical cutOne sentenceA second sequence, built by hand

If your work involves multi-cam interviews, colour-managed RAW, or a client who sends frame-accurate notes, a conventional NLE is still the correct tool and no amount of agent will change that. If your work is a recording of yourself talking that needs to be tightened, captioned, scored and cut vertical — which is most video made by most people — describing the outcome is a faster route to the same file. The two are not really competing for the same afternoon.

Where an Editing Agent Fits in the Category

In “It’s time for agentic video editing” (Justine Moore, a16z, 21 January 2026), the work an editing agent could take on is split into five parts: Process, Orchestrate, Polish, Adapt and Optimize. It is a useful frame precisely because it makes gaps visible, so here is where this agent sits against it, including the parts where it does nothing.

  • Process — sorting through the footage and deciding what to use. This is the core of it: indexing, silence and filler detection, repeated-take detection, and cutting from what was said.
  • Orchestrate — partial, and honestly so. Generated images, generated sound effects, short generated clips and media pulled from a link can all be placed inside an edit, but this is not a conductor for a fleet of external models. The centre of gravity is your footage.
  • Polish — partial. Captions, music that ducks under speech, loudness mastering, colour grades, zooms, punch-ins on emphasised words and finishing effects are all here. The audio-repair half of that bucket is not: no denoise, no “studio sound”, no levelling each speaker separately.
  • Adapt — partial. Reframing to 16:9, 9:16, 1:1 and 4:5 with the captions rescaled, so one recording becomes the version each platform wants. See aspect ratio if that is unfamiliar territory. The other half a16z means by this word — translating and dubbing a video into other languages — is not offered.
  • Optimize — not offered, and this is the honest gap. The piece described there is taste: an agent exercising editorial judgement on its own and handing you a few draft versions to react to. One request here produces one cut, and the judgement arrives through your corrections rather than ahead of them. It is the part of the frame worth being sceptical about whenever a tool claims to have closed it.

Programmatic access, for anyone who wants to build on top of it, is the MCP server described below rather than a separate REST product.

The Same Agent, Running Inside Claude

If the agent is the editor, then the agent is replaceable — which is a strange sentence until you see what it buys you. All 108 tools are published as a remote Model Context Protocol server (97 for editing, 11 for projects, uploads, indexing, rendering, export and download) over streamable HTTP, with OAuth 2.1, dynamic client registration and PKCE, so there is no API key to copy anywhere. Add it as a connector in the Claude app, sign in, and edit your video from the conversation you were already having.

The tools are not a re-declared subset written for the connector. The editing engine publishes one registry and the connector serves it verbatim, so the model in your Claude session gets what the in-house agent gets: the same names, the same schemas, the same refusals. Long operations return a job id and a tool to wait on it rather than a fabricated completion, and editing the same project from the studio and from Claude at the same time is refused in both directions instead of quietly racing.

Setup, including the errors worth recognising, is on the Claude connection page; the complete tool list is on the tool reference. ChatGPT is a narrower fit — its read-only connector path cannot drive an editor at all, and the developer mode that does load full read-and-write tool surfaces is documented as web-only and on paid accounts. That path, and the four places it stops, is written up on the ChatGPT page.

How to Get an AI Agent to Edit Your Video

  1. 1
    Upload the footage you shot
    Up to 14 GB or 3 hours per video, as MP4, MOV, MKV or WebM. It is indexed once with visible progress — word-level transcript, silences, shot changes and frame tiles — and every later edit reuses that index.
  2. 2
    Describe the finished video, not the steps
    Type what you want to be true at the end: “cut the dead air and the ums, caption it karaoke style, put calm music under my voice, and give me a vertical cut”. The agent works out the order the jobs have to happen in.
  3. 3
    Watch the preview and say what is off
    The agent renders a preview and inspects the frames it produced. React in plain English — “leave a breath after each sentence”, “captions lower, they are on my face”, “swap the music for something with drums”. Each correction is another full pass, not a manual repair you perform.
  4. 4
    Export from the original
    When it is right, export. The final render is cut from your original upload at source quality and downloads as an H.264 MP4. Nothing about the upload has been altered, and anything you cut can still be restored.

Indexing is a one-time cost per upload and scales with length; editing turns are short by comparison, and the final export takes longer than a preview because it reads the original file rather than the proxy.

Frequently Asked Questions

Yes. Valmera is an AI agent that edits real footage you upload. You describe the change in ordinary English — “cut the dead air, caption it, make it vertical” — and the agent performs the whole edit: it analyses the video first, makes the cuts, writes and times the captions, places and ducks the music, reframes the picture, renders a preview, looks at the frames it produced, and exports a finished MP4 from your original file. You do not touch a timeline unless you want to. It is not a text-to-video generator: it edits footage that already exists, and it will tell you so if you ask it to invent a video.
Cut anything you can describe, including silences, filler words and repeated takes. Burn word-accurate captions with 11 presets, including karaoke word-pop and per-word emphasis. Add music from a royalty-free library, your own file or a link, ducked under your voice. Change speed from 0.25x to 4x. Add zooms, punch-ins on emphasised words, colour grades, film-grain-style finishing effects, transitions at scene changes, text templates, title cards and picture-in-picture overlays. Reframe to 16:9, 9:16, 1:1 or 4:5. Blur or black out a region that keeps following the subject through later cuts. Splice in generated images, generated sound effects or short generated clips. Master the loudness. All of it from sentences rather than menus.
Yes — a complete edit can be produced and exported without ever opening one. The timeline, the player and the editable transcript are there for when you want to check something or nudge it by hand, but nothing in the workflow requires them. In practice you type a request, watch what comes back, and reply with whatever is bothering you about it. Each of those replies is another full pass by the agent rather than a repair you carry out yourself.
No. The agent edits an edit decision list — a record of which ranges to keep and what to apply to them — and never touches the pixels of your upload. The renderer reads that list. Your original file stays exactly as you uploaded it, which is what makes the final export source quality: it is cut from the original, not from a compressed copy. Previews render from a faster proxy so they come back quickly, and previews are never watermarked on any plan.
Nothing it does is destructive, so getting it wrong costs a sentence rather than a file. Every operation writes a new version of the decision list, anything cut can be restored by asking for it back, any individual effect can be removed by name, and the whole edit can be reset to the untouched upload. Because the original is never modified, the worst case is that you ask for the change again with better words. There is also a server-side check on every reply: the agent's message is verified against the edit decisions actually recorded, so it cannot report a change it did not make, and when nothing changed it says nothing changed.
It can find the moment and build the vertical cut. Ask for the section you want — by what was said, not by timecode — and the agent resolves it from the transcript, trims to it, reframes to 9:16 with a subject-aware crop, and captions it. What it does not do is emit a batch: one request produces one deliverable, so five clips are five requests rather than one button that returns a folder. If you want a whole recording sliced automatically into every viable clip, a dedicated clipping tool is the better shape for that job.
There are two waits and they are different. Indexing happens once per upload, scales with the length of the video, and shows progress while it runs — every later edit reuses it, so you pay that cost once. An editing turn is then the agent thinking, writing the decisions and rendering a preview from a fast proxy, which is a short wait rather than a full export. The final export is a real render from your original file at source quality, so it takes longer than the preview did.
Valmera gives every account 50 credits on signup with no credit card, and they are spendable on real edits of your own footage rather than a locked demo. Free exports carry a small Valmera mark in the top-left corner; previews are never marked, and upgrading re-renders an already-marked export clean. Paid plans are Creator at $30/month (2,000 credits), Pro at $50/month (4,000) and Frontier at $100/month (10,000, on a stronger model), and each opens with a 3-day free trial.
Inside Claude, yes. Valmera publishes all 108 of its tools as a remote Model Context Protocol server, so you add it as a connector in the Claude app, sign in through OAuth, and then upload, cut, caption, mix, render and download from the conversation you are already in. ChatGPT is a narrower fit. Its default connector path — the one used for company knowledge and deep research — is read-only search and fetch, which cannot drive an editor at all. Its developer mode does load remote MCP servers with both read and write tools, with write actions confirmed as they run, but it is documented as being on the web only and on paid accounts. On its own, with no connector attached, ChatGPT can plan an edit and write you a shot list; it cannot open your footage or produce a rendered file.
Valmera specifically cannot generate a video from a text prompt, import or export SRT or VTT files, do a true crossfade or dissolve, use a different transition on each cut, motion-track a sticker or blur onto a moving subject, accept an uploaded font, denoise or “studio sound” your audio, level each speaker separately, pull music back out of an already-baked mixdown, generate original music, produce a batch of clips from one request, publish to YouTube or TikTok, share a link, or host team seats. Every export also closes with a short Valmera end card. Transcription is best tested in English. More broadly, no agent can know your taste until you tell it — the first pass is a proposal, and the corrections are where your video actually gets made.

Keep Reading

If you want the individual jobs rather than the whole workflow: removing silence, karaoke captions and resizing for a platform each have their own page. For the wider toolset there is the AI video editor overview, and for how this compares to the other tools claiming an agent, the agentic field. If you have not used the product before, getting started is four paragraphs long.

Hand It a Video and See

50 credits on signup, no card. Or connect it to Claude and edit from the chat you already have open.

Start free →
See pricing →

Related Articles

Agentic Video Editor
The same product in category language: what makes an editor agentic, and the test that separates it from an AI-assisted one.
Edit Video by Chatting With AI
A real exchange with the agent, corrections included, with no timeline involved at any point.
Connect Claude to Video Editing
Step-by-step setup for driving the same agent from the Claude app or Claude Code.
Can AI Edit Videos?
The wider question, answered without the marketing: what AI editors genuinely do and where they still fall over.