← Home
LIBRARY

Published · Updated

AI Video Editing Prompts: What to Actually Type

A prompt in a video editor is not a search query and not an incantation. It is an instruction to something that has already read your footage. Every word carries a timestamp, every silence has been measured, every shot boundary found, and the agent has looked at labeled frame tiles of the picture itself. Prompts fail for exactly two reasons: they name something that is not in that index, or they ask for a thing the tool genuinely cannot do.

Below: the five rules that decide whether a request resolves, 96 copy-paste prompts grouped by the job you sat down to do, and an honest section on the requests that will never work no matter how you phrase them. Every prompt here maps to a shipped capability of Valmera's AI video editor — nothing aspirational, and nothing that was quietly withdrawn.

Paste Any Prompt Into a Real Editor

50 free credits on signup, no card. Upload footage, paste a prompt, watch the agent perform the edit.

Start free →
See pricing →

What Makes a Video Editing Prompt Work

Five rules, each with the weak version and the rewrite. The weak versions are not strawmen — they are the four or five phrasings that show up most often, and every one of them is a reasonable thing to say to a human editor who has watched the footage with you. The difference is that a human fills the gap with context they already have. An agent fills it with a guess.

1. State an outcome, not an operation

WEAK
Add some cuts to tighten this up.
STRONG
Cut every pause longer than half a second so the whole thing runs under four minutes.

“Tighten” names a direction with no finish line, so the agent has to invent one and you have no way to check it. The rewrite carries a threshold (half a second) and a stopping condition (under four minutes). Both are measurable off the index, which means the agent knows when it is done and you know what to look at.

2. Anchor to something the agent can locate

WEAK
Zoom in on the important bit.
STRONG
Punch in when I say “this is the part everyone gets wrong.”

There are exactly three anchors that resolve: a spoken phrase, a timestamp, and a described moment (“the part where I knock the mic”). All three land on the same index — a word-level transcript, measured silences, shot boundaries and labeled frame tiles. “The important bit” is not in any of them, so it resolves to a guess.

3. One deliverable per request

WEAK
Give me five clips from this podcast, each under a minute.
STRONG
Cut my answer about pricing into a 9:16 clip under 60 seconds with karaoke captions.

Stacking several operations in one message is normal and good — that is what the multi-step section below is for. Stacking several separate outputs is different: there is one program per project and one preview per turn, so a five-clip request has no shape to come back in. Ask for the best one, then ask for the next.

4. Say the platform when framing matters

WEAK
Make it vertical.
STRONG
Reframe to 9:16 for Reels, aim the crop at whoever is speaking, and keep the captions clear of the bottom UI.

A platform name is shorthand for a bundle of constraints — aspect ratio, safe area, how big captions have to be to survive a phone screen, what loudness the mix should hit. Naming Reels imports all of it in one word. “Vertical” imports one of them.

5. Correct by describing what you saw

WEAK
That is not quite right, try again.
STRONG
The cuts are too tight — leave about a third of a second of breath after each sentence.

Every turn edits the existing decision list rather than starting from the source, so a correction is cheap, local, and keeps the parts that already worked. “Try again” throws away information you have and the agent does not: which specific thing in the frame you did not like.

How to Write a Video Editing Prompt That Works

  1. 1
    Write the finished state, with a stopping condition
    Say what the video should be when the turn is over, not which button you imagine being pressed. A threshold or a target length is what lets both of you tell when it is done — "under four minutes", "pauses longer than half a second", "60 seconds for Shorts".
  2. 2
    Anchor it to a phrase, a timestamp or a described moment
    Three anchors resolve against the index: a line you actually said, a timecode, or a moment you can describe ("the part where the doorbell goes"). Mixing them in one sentence is fine. Anything else — "the important bit", "the boring part" — resolves to a guess.
  3. 3
    Keep it to one deliverable, and stack operations freely
    Several operations in one message is the point: cuts, captions, music and a reframe together let the agent sequence the dependencies. Several separate outputs is the thing that has no shape — ask for one clip, then ask for the next.
  4. 4
    Name the platform when framing or loudness matters
    "For Reels" or "for Shorts" imports the aspect ratio, the safe area, the caption size that survives a phone screen and the loudness target, all in one word. Leave it out only when the video is not going anywhere in particular.
  5. 5
    Watch the preview, then correct by describing what is wrong
    The follow-up turn is where an agentic editor beats a one-shot tool. "The captions are covering my hands", "the cuts are too tight", "the grade is too warm" all edit the existing decision list rather than starting over, so everything that already worked survives.

Indexing runs once per upload and shows progress; every prompt after that reads from the index and is fast. Nothing you ask for is destructive — the original file is never modified, and anything cut can be restored by asking.

96 Prompts, Grouped by Job

Grouped by the job you sat down to do rather than by the feature that implements it, which is why reframing appears under short-form and again under framing. Copy any of them verbatim — none of it is syntax, so rewriting a prompt in your own words costs nothing and changes nothing.

Two things to know before you start. Anything that refers to a file — a logo, a track, a screen recording — needs that file in the project first, by upload or by pasting a link in the chat. And when a request turns on a taste call the footage cannot settle, the agent asks one specific question rather than picking for you and hoping. If you want the shorter, feature-ordered list instead, the 50-prompt version is a completely different set.

Cleanup 19

The pass that pays for itself. Cuts snap to word boundaries so a removal never clips a syllable, every cut is reversible, and every word in the recording already carries a timestamp — so you name what to remove and skip the hunt for its timecode.

1
Take out every pause longer than half a second, but leave a breath before each new point.
2
Cut the ums and uhs, and treat “sort of” and “right?” as filler too.

The filler list is extensible — name any crutch word you overuse and it goes in the same pass.

3
I flubbed the intro four times. Keep the last clean take of each line and drop the rest.
4
Everything before I say “let’s get into it” is throat-clearing. Cut it.
5
Remove the section where the doorbell goes and I lose my train of thought.
6
You cut too hard around 6:20. Put that back and leave the pauses in that section alone.

Restoring one range does not disturb any other cut you have already accepted.

7
Trim the dead air after my last word so it ends on the sentence.
8
Tighten the middle third only. Leave the intro and the outro at the pace they are.
9
Show me the transcript of what is actually left after the cuts so I can check nothing repeats.

This reads the kept program, not the raw recording, and flags repeated phrasing that survived a tightening pass.

Captions 1018

Captions are the spoken words, timed from the transcript and burned into the picture. 11 presets, 12 bundled font families, 1–16 words per caption, three positions, per-word emphasis and 9 entrance animations — all addressable by name.

10
Caption the whole thing in the beast preset, size L, sitting just above the bottom third.
11
Karaoke captions, two words at a time, white text with a red pop on the word being spoken.
12
Use Anton for the captions, all caps, with a punch entrance on every line.
13
The captions hear “Kubernetes” as “cooper netties.” Fix that everywhere and re-render them.

Correcting the transcript re-renders the captions from the corrected words — you never retype a caption.

14
Cap the captions at five words a line so they never wrap on a phone.
15
Emphasise every number that appears in the captions — bigger and in the accent colour.
16
Move the captions to the top for the section where I am writing on the whiteboard.
17
Hide the captions while the full-frame quote card is up, then bring them back.
18
Switch the caption preset to editorial and drop the size one step. They are fighting the footage.

Short-form and clipping 1926

One clip per request, directed rather than batched. That is the trade: you do not get ten candidates ranked by a virality score, you get the one clip you asked for, cut where you said and framed for where it is going.

19
Find the strongest 45 seconds in this, cut everything else, and make it 9:16.
20
Start the clip on the sentence “nobody tells you this.” No preamble before it.
21
Cut this under 60 seconds for Shorts. Keep the argument intact and end on the punchline.
22
Take my answer to the pricing question out as a standalone vertical clip with karaoke captions.
23
The first three seconds have to carry the hook. Cut straight to the claim and put it on screen as text as well.
24
Run the setup at 1.5x and leave the payoff at normal speed.
25
Punch in on the most vocally emphasised words across the whole clip.
26
This is going on TikTok. Reframe to 9:16 aimed at whoever is talking, and master the audio for phone speakers.

Naming the platform imports the aspect, the caption scale and the loudness target in one go.

Podcast 2733

The transcript carries speaker labels, so two people in one file can be cut, searched and levelled separately by naming who is talking. What it does not carry is a diarization export or an automatic speaker balancer — levelling is volume automation over spans.

27
Two speakers in this one. Cut the dead air and the filler across both of them and leave everything else alone.
28
Drop the first eight minutes of small talk and open on the first real question.
29
Find where she talks about hiring her first employee and give me the timestamps.
30
Put a lower third with the guest’s name and company the first time she speaks.
31
Lay an ambient track under the whole episode, low, ducked under the voices.
32
I am noticeably louder than my guest. Pull my volume down across the spans where I am the one talking.

This is volume automation over the spans the transcript attributes to you — not a per-speaker levelling engine, which Valmera does not have.

33
Master the episode to standard loudness, fade up from black at the top and out at the end.

YouTube 3440

Long-form talking head is where the mechanical work is worst and the agent is strongest: pacing cuts, punch-ins on emphasis, a consistent grade, a logo that stays put, and section cards that do not need a designer.

34
Cut the silences, then punch in on my most emphasised words so it stops feeling static.
35
Open with a full-frame title card reading “I WAS WRONG ABOUT THIS” before the first frame of me.
36
Put my channel logo in the top-right at 8% of the frame width for the whole video.
37
Splice the screen recording I uploaded over the part where I explain the settings.
38
Add a chapter card at each new section — The Problem, The Fix, What I’d Do Differently.
39
Grade it cinematic, add a light vignette, and fade in and out.
40
Bring the music in loud over the first four seconds, then duck it under my voice for the rest.

Course and tutorial 4147

Teaching video has a rhythm that clipping tools get wrong: the pauses after an instruction are load-bearing and should survive, the typing should be slow enough to follow, and every module should look identical to the last one.

41
Caption the whole lesson in the elegant preset, bottom, four words a line — same as my other modules.
42
Cut the pauses where I am reading ahead, but leave the pauses after each instruction.
43
Every time I say “important,” put a callout on screen with that sentence.
44
Slow the part where I type the command to 0.75x so it is followable.
45
Add a title card between each of the five steps with the step number and its name.
46
Zoom into the terminal whenever I am typing and pull back out when I talk to camera.
47
Make a 4:5 promo cut from the two minutes where I explain why this matters.

Screen recordings and demos 4854

A screen recording is unwatchable at full frame because the thing that matters is 200 pixels wide. The fixes are mechanical: a floating window treatment, a bigger cursor, a zoom that travels with the pointer, and speed over the parts where nothing happens.

48
This is a screen recording. Put it in the floating window treatment on a dark gradient.
49
Make the cursor bigger and steadier, it is hard to follow at this size.

The pointer is found in the actual frames rather than assumed from where it was supposed to be.

50
Follow the cursor with a travelling zoom as it moves from the sidebar to the settings panel.
51
Run the loading stretches at 4x and leave the rest at normal speed.
52
Zoom into the top-left panel from 0:20 to 0:35. The text is too small to read at full frame.
53
Record my landing page and open the video with it.
54
Record a browser signing up on my site, then cut it like a product demo with a zoom on every click.

The recorder captures real click timestamps, so the zooms land on the clicks rather than near them.

Music and sound 5562

Four independent layers — the speaker, music, voiceover and placed sound effects — each with its own gain. Music ducks under speech by sidechain compression rather than a static level cut, so it swells back in the gaps.

55
Show me the music library and pick something cinematic that suits a two-minute piece.
56
Put the chill track under the whole thing, ducked, and fade it out over the last five seconds.
57
Swap the music for something with more energy but keep the level and timing where they are.
58
Analyse the track I uploaded and tell me the tempo and where the biggest energy rise is.
59
Snap the cuts in the montage onto the beat of that track.

If the track has no reliably detectable tempo the agent says so instead of nudging cuts at random.

60
Generate a low sub-bass impact and land it on the title reveal.

Sound effects are generated from your description, 0.5 to 22 seconds. There is no bundled effects library to pick from.

61
Take the audio out of the clip I uploaded and use it as the music bed.
62
Mute my original audio over the b-roll stretch and let the music carry it.

Look and colour 6370

Grades apply to the whole program; the eight finishing effects are windowed. The two compose, which is why “cinematic grade plus grain over the flashback only” is one sentence rather than two conflicting requests.

63
Apply the luxury look — captions, grade, transitions, all of it — and show me a preview.
64
Grade it cool and pull the saturation back about 15%.
65
It is underexposed. Lift the exposure a little without crushing the contrast.
66
Film grain and a vignette over the flashback from 2:40 to 3:05, everything else clean.

Finishing effects take a time window. Colour grades do not — a grade is the whole program.

67
Give the drop one frame of flash and a small camera shake.
68
Add VHS over the intro and drop it at the first cut.
69
Make it look like a meme edit. The whole aesthetic, not just the captions.
70
The grade came out too warm. Pull the temperature back towards neutral.

Framing and reframing 7177

Four output frames — 16:9, 9:16, 1:1, 4:5 — reached by crop, pad or blurred pad. Cropping never upscales, captions rescale to the new frame, and the crop can be aimed at the person actually speaking rather than at the centre of the picture.

71
Reframe to 9:16 and aim the crop at whoever is actually speaking, not the middle of the frame.
72
Make it 4:5 with blurred padding. I do not want to lose the top of the whiteboard.
73
Go vertical for the section where I read the message out, then back to 16:9.
74
There is too much headroom for the whole video. Push in slightly and hold it there.
75
Same edit, give me it at 1:1 for the feed.
76
The 9:16 crop is cutting off my hands when I gesture. Pad it instead of cropping.
77
Reframe to 9:16 and make sure the captions scale down with the new frame.

Text and titles 7884

On-screen text is a separate system from captions: seven designed templates — title, subtitle, lower third, callout, big number, quote, chapter — each with its own time window, position, colours and entrance, including a typewriter.

78
Open with a title reading “THE 5-MINUTE FIX” with a typewriter entrance.
79
Put a quote card up with the line I just said, middle of frame, for four seconds.
80
Add a callout saying “don’t do this” over the mistake at 1:12.
81
Big number “$0” on screen when I say the price, then let it fall away.
82
Lower third: my name on the first line, Founder, Northwind on the second, first six seconds.
83
Cut to a full-frame black card that just says PART TWO between the two halves.
84
The title is sitting over my face. Move it to the lower third instead.

The agent renders a preview and looks at the frames, so a collision like this is usually caught before you have to say it.

Censoring and privacy 8590

Two different jobs that get confused. Censoring puts a visible mark over a region on purpose and it stays visible. Erasing repaints the pixels so the thing is gone and the background behind it is reconstructed.

85
Pixelate the customer name in the CRM for the whole screen-recording section.
86
Black bar over the bottom-right where my email is showing.
87
Blur the number plate in the driveway shot. It is in frame from 0:12 to 0:19.

A censor region is fixed to a rectangle. It follows the footage through later cuts and reframes, but it does not chase a moving subject.

88
There is a watermark from the app I recorded with. Actually remove it, do not cover it.
89
The bottom of this file has burned-in subtitles from an old export. Take them out and reconstruct the background.
90
Remove the censor on the whiteboard, I checked and there is nothing sensitive on it.

Multi-step “do the whole thing” 9196

The requests worth stacking are the ones with dependencies. Cuts change the timeline the captions have to be timed against; the music has to be fitted to a program length that nothing knows until the cuts are done. Sequencing that is the agent’s job, not yours.

91
Cut the dead air and the filler, add karaoke captions, chill track under it ducked, reframe to 9:16, master the audio.
92
Turn this raw talking head into a finished video: tighten the pacing, punch in on the emphasis, cinematic grade, logo top-right, fade in and out.
93
Course module. Clean it up, caption it in the podcast preset, title card per section, give me it at 16:9.
94
Podcast episode: cut silences and filler for both speakers, lower thirds on first appearance, ambient bed ducked under the voices, master to standard loudness.
95
Screen demo: floating window treatment, bigger cursor, a zoom on every click, a title card at the top, speed up the loading.
96
Make the whole thing feel like an ad for a watch. You pick the grade, the captions and the music, then tell me what you chose so I can override it.

Handing over the taste call and asking to be told what was chosen is the fastest way to converge when you do not have the vocabulary yet.

Every Prompt Above Runs on the Free Plan

50 credits on signup, no card. Upload a video and paste one in.

Start free →
See pricing →

Prompts That Do Not Work, and Why

Most prompt libraries stop at the list. That is the half that is easy to write and the half that is easy to find elsewhere. This is the other half: four requests people type constantly, why each one fails, and what to type instead. They fail in four different ways — one is unmeasurable, one is ambiguous, one has no reference, and one is simply addressed to the wrong category of software.

Make it go viral.

There is nothing in the footage that measures this. Every other request in this library resolves to something the agent can find — a word, a silence, a region, a loudness figure. “Viral” resolves to a hope, and the honest response is a question rather than an edit.

Say this instead. Say the mechanics you actually believe cause retention, because those are all real operations: open on the strongest claim with no intro, cut every pause, captions on, 9:16, the hook on screen as text in the first three seconds. If you want the agent to choose, ask it to choose and report back — “you pick the opening line and tell me why” works.

Fix the audio.

“Fix” covers about five different jobs and Valmera does three of them. Loudness mastering, ducking music under speech, and per-span volume automation are real. Denoise, “studio sound” and pulling music back out of an already-mixed track are not, and never will be from this tool. A vague request makes it impossible to tell which half you meant, so you find out by being disappointed rather than by being told.

Say this instead. Name the job. “Master the whole thing to standard loudness.” “Duck the music under my voice.” “Mute my original audio and keep the music.” “Pull my volume down over the spans where I am talking.” If the actual problem is hiss, room tone or a phone microphone, no prompt fixes it — that is a re-record or a dedicated audio repair tool before you upload.

Make it look professional.

Professional is a reference, not a setting. The word describes a hundred incompatible looks — a bank explainer and a streetwear ad are both professional and share no single parameter. The agent can pick one, but it is picking blind, and you will spend the next three turns walking it back.

Say this instead. Either name the parameters (“cinematic grade, light vignette, fade in and out, captions in the editorial preset”) or name a feeling and hand over the taste call explicitly: “make it feel like a luxury ad, then tell me what you chose.” The second is faster when you do not have the vocabulary yet, because the reply gives you the vocabulary.

Generate a video about remote work.

Wrong category, and it is the single most common mismatch. Valmera edits footage you already have. It can generate stills and 5- or 10-second clips to splice into an edit, and it can generate a sound effect — but there is no text-to-video that produces a whole video, and asking for one is asking a film editor to shoot the film.

Say this instead. Upload the recording and describe the edit. If you have no footage at all, the honest alternatives are a different category of tool for generation, or a canvas project built from images, a voiceover and music — which is a real thing Valmera does, with no main video required.

Twelve more that are simply out of scope

These are not phrasing problems. They are documented limits, and no rewording gets past them — so each line names the nearest thing that is real. A refusal without an alternative is just an error message.

Crossfade between the shots.

No true crossfade or dissolve. Seven junction styles exist; dip to black is the closest feel.

Use a different transition on each cut.

One transition style applies to every cut. Pick the one that suits the video.

Keep the blur on the bystander while he moves through the shot.

Censoring covers a fixed rectangle and does not follow a moving subject. Cover a larger area over that window, or cut the shot.

Export the captions as an SRT.

Captions are burned into the picture. There is no SRT or VTT import or export; the editable transcript is the closest thing.

Use my brand font.

No custom font uploads. Twelve families ship with the editor; Inter Display and Montserrat are the two that read closest to a generic brand sans.

Post it to my TikTok when it is done.

No publishing or scheduling to any platform. Export is a file you download and post yourself.

Give me ten clips from this recording.

One deliverable per request. Direct them one at a time.

Add chapter markers for YouTube.

There is a chapter text template, which is a visual. There is no chapter metadata written into the file.

Separate the music from the speech in this track.

Not possible on a single already-mixed track.

Write me a track for this.

No AI music generation. Use the 23-track library, upload your own, or paste a link. Sound effects can be generated; songs cannot.

Balance the two speakers automatically.

No per-speaker levelling and no diarization output. Volume automation over the spans each person speaks is the real version of this.

Interpolate the slow motion so it is smooth.

Slow motion duplicates frames rather than synthesising them; below about 0.6x the movement visibly steps.

You do not have to memorise any of this, because the agent will tell you. Every reply is checked server-side against the edits actually written to the edit decision list, so it cannot confirm a request it did not carry out. That verification is the reason a prompt library for this editor can be honest about its edges — the alternative is a page that promises what the product declines, and the product wins that argument every time.

The Same Prompts, Typed at Claude

Valmera publishes its complete editing toolset as a remote MCP server with OAuth, so you can add it as a connector in the Claude app and paste these prompts into the conversation you are already in — upload, cut, caption, mix, render and download without opening the studio. It is the same registry the in-house agent uses, served verbatim, so there is no second tool list that could drift and no prompt that works in one place and not the other. The setup guide covers the Claude app and Claude Code; the tool reference lists what each tool does and what it refuses.

One thing genuinely changes about prompting over MCP, and it is worth knowing. Claude is doing the sequencing, which means you can hand it context the studio has no field for: a document of brand rules, last month's caption settings, a transcript you already annotated. "Edit this the way we did the January episode, the notes are in the doc above" is a prompt that only makes sense when the editor lives inside a conversation that already has the doc in it. The trade in the other direction is that the studio's preview player is right there, and over MCP you download a file to watch it.

Where to Go Deeper

Each of these covers one job in detail — the options, the defaults, and the edges the prompts above skate over: silence removal, filler words, captions and karaoke captions, reframing, censoring a region, and cutting clips out of a long recording.

For the vocabulary rather than the task, the caption docs list every preset, font and animation by name, and the audio docs cover ducking, gain and the loudness target. If you have not run a first edit yet, getting started is upload, describe, preview, export in four screens — and word-level timestamps explains the index that makes "cut the part where I say…" resolve at all.

Frequently Asked Questions

State the outcome you want rather than the operation, and anchor it to something the editor can locate in the footage — a spoken phrase, a timestamp, or a described moment like “the part where I knock the mic.” An agentic editor has already indexed your upload into a word-level transcript, measured silences, shot boundaries and frame tiles it has looked at, so those anchors resolve to exact positions. Keep each request to one deliverable, name the platform when framing matters, and correct by describing what you saw rather than what to try instead. “Cut every pause longer than half a second, add karaoke captions two words at a time, and reframe to 9:16 for Reels” is a good prompt. “Tighten it up and make it look professional” is not, because neither half of it resolves to anything measurable.
Cleanup, before anything else, and for a mechanical reason rather than a stylistic one: cuts change the timeline that everything downstream has to be timed against. Captions timed to the raw recording drift once you remove the pauses, and music fitted to the raw length is suddenly too long. Ask for silence and filler removal first, look at the result, then layer captions, music, framing and look on top of the edit you accepted. If you would rather not think about it, stack the whole thing in one message — the agent sequences the dependencies itself.
The five rules transfer to any agentic editor — they are about how a request resolves against an index, not about one product. The specific prompts are written against Valmera’s shipped capabilities, and each one names things Valmera actually has: the 11 caption presets, the 12 bundled fonts, the 7 transition styles, the 6 grades, the 8 finishing effects, the four output frames. In a different tool, the phrasing survives and the vocabulary may not.
Yes, and the multi-step ones are usually the better prompts. “Cut the dead air, add captions, put music under it and reframe to 9:16” is four operations with real dependencies between them — the cuts change what the captions are timed against, and the music has to be fitted to a program length nothing knows until the cuts are done. Sequencing that is exactly what the agent is for. The thing you cannot stack is separate outputs: several operations producing one video is fine, several videos from one request is not.
It says so, and the saying is enforced rather than trusted. Before a reply reaches you it is checked against the decision list the system actually wrote, so it cannot claim work that never happened and cannot confirm an impossible request. Ask for a crossfade and what comes back is a refusal with the closest real option attached — a dip to black — instead of a cheerful confirmation over an unchanged video.
Specific about the target, loose about the craft. Naming the moment precisely is always worth it, because a wrong target wastes the whole turn. Naming the aesthetic precisely is optional, because the agent proposes and you correct in one sentence — “the grade is too warm” costs less than getting the temperature right in advance. The exception is when you already know exactly what you want; then say it, because “karaoke captions, two words a line, red pop on the spoken word” lands on the first try where “add captions” takes three turns.
Yes, and mixing the two in one sentence is fine: “cut from 2:10 to 2:48, and remove the bit after that where I repeat the same point.” Timestamps are exact and descriptions are robust to you misremembering, so use whichever you actually have. What no tool accepts is a timestamp the model invented — every timing an edit uses comes from the index or from you.
Ask for it back in plain English. Nothing removed is thrown away, so “put back what you took out around 6:20” returns that footage and leaves every other cut where it is, and anything you styled — the captions, the grade, the track, a censor region — can be swapped or dropped in a follow-up. The agent edits a decision list rather than pixels: your original upload is never modified, and each operation is a new version.
No, and the prompts here deliberately mix registers to show it. “Lower third” and “4:5 with blurred padding” are the vocabulary; “it is too small to read” and “the cuts are too tight” are not, and both work. When you do not have the word for something, describe the problem you can see. The vocabulary comes back in the reply, which is a reasonable way to learn it.
Yes. Valmera publishes the same tool registry its own agent uses as a remote MCP server with OAuth, so you can add it as a connector in the Claude app and type these prompts in the conversation you are already in — upload, cut, caption, mix, render and download without opening the studio. The prompts do not change. What changes is that Claude is doing the sequencing, and you can hand it context it can read for itself, like a document of brand rules.

Try a Prompt on Your Own Footage

50 free credits on signup, no card required. Or connect Valmera to Claude and paste the same prompts into the chat you already use.

Start free →
See pricing →

Related Articles

50 AI Video Editing Prompts That Actually Work
The shorter original list, organised by feature rather than by job. A different 50 — no overlap with the prompts on this page.
Edit Video by Chatting with AI
A real exchange end to end, including the corrections — what the turns after the first prompt actually look like.
Agentic Video Editor
Why the prompt is the whole interface: the loop the agent runs, and the return path that separates it from automation.
Valmera MCP Server
Type the same prompts at Claude. The full editing toolset as a remote MCP server with OAuth.
All Editing Tools
Every job Valmera performs, one page each — the capability list these prompts are checked against.