← Home
TOOL

Published · Updated

Text-Based Video Editing

Text-based video editing means editing a video through words instead of a timeline: you read a transcript of what was said, then either correct the text or type the change you want. Valmera supports both — an editable transcript, and typed instructions that an AI agent actually executes, from cuts to captions to the final mix.

IN SHORT

There are two schools of text-based editing. In transcript-deletion editors like Descript, you delete words and the video cuts follow — you still drive every decision. In Valmera, you type the outcome — "remove the silences and the repeated takes" — and the agent performs the edit, with every reply verified against what actually changed.

Edit Your Next Video by Typing

Upload footage, type the edit, review the preview. Free plan — no credit card required.

Start Editing Free →
See pricing →

How to Edit a Video by Typing

  1. 1
    Upload your video
    Upload footage up to 2GB or 3 hours. Valmera analyzes it with visible progress, building a word-level transcript with timings — the foundation every text-based edit runs on.
  2. 2
    Type the edit
    Tell the agent what you want: "remove all the silences and filler words", "cut everything before I say welcome back", "add karaoke captions and duck the music under my voice". One message can carry several jobs.
  3. 3
    Preview, refine, export
    Watch the preview, then steer with follow-ups — "tighten it more", "bring back the story at 2:10". When it's right, export an H.264 MP4 rendered from your original upload at source quality.

The analysis after upload is a one-time step and runs with visible progress — longer videos take longer. After that, each edit is a typed message and each change renders a fresh preview.

What Text-Based Editing Actually Is

Every text-based editor starts the same way: the software transcribes your video into a word-level transcript, where each word carries its exact timing in the footage. That transcript becomes the interface. Instead of scrubbing a timeline hunting for the moment a sentence ends, you work with the words themselves — which is a far more natural way to edit anything built on speech: talking-head videos, podcasts, tutorials, interviews, course lessons.

What differs between editors is what the text controls. In the transcript-deletion school, the text is the timeline: deleting a word deletes that slice of video. In the instruction school — the agentic model Valmera uses — the text is a conversation: you describe the result you want and the agent decides and executes the cuts, snapped to word boundaries so speech is never clipped mid-word.

Word Deletion vs Typed Instructions

Descript popularized transcript-deletion editing, and for hand-picking exact words to remove it remains excellent — a document-like view, multitrack audio, screen recording, and an AI co-editor called Underlord to assist. Its Creator plan runs $24/month billed annually, with limited Underlord use on the free tier. The catch is scale: deleting words one by one is still manual labor, and you are still the editor.

Typed instructions scale differently. "Remove every um and uh" is one sentence whether the video is four minutes or two hours — Valmera's agent applies it across the whole transcript at once, the same way filler-word removal handles a built-in list plus any custom words you add, and repeated-take removal finds the retakes in a recording session and keeps the best one. You review the result instead of performing it. The full head-to-head is in Valmera vs Descript.

Video Edits You Can Type

Anything the editor can do, you reach by typing. A sample of real requests, verbatim:

  • "Cut out all the silent parts" — dead air removed in one request, word-boundary-safe.
  • "Cut from 0:12 to 0:47" or "cut the part where the delivery guy interrupts" — cuts by timestamp or by describing the moment.
  • "Add karaoke captions, white with a yellow highlight" — word-accurate captions, styled in the same sentence.
  • "Put a lower third with my name at the start" — designed on-screen text templates.
  • "Add an upbeat track and duck it under my voice" — music from a 24-track royalty-free library, mixed under speech automatically.
  • "Speed up the setup section to 1.5x" — any range from 0.25x to 4x, voices keeping natural pitch.
  • "Punch in whenever I make the key point" — automatic zoom punch-ins on your most emphasized words.
  • "Make it 9:16 for Shorts" — reframed by crop, pad, or blurred pad, captions rescaling with the frame.

Each request produces one updated edit and one fresh preview, so you always know exactly which change you are judging.

The Editable Transcript

The transcript in Valmera is not read-only. If the analysis mishears a name or a technical term, you fix the word directly and the captions re-render from the corrected text — no re-upload, no manual caption editing. The same word-level timings are what make typed cuts precise: when the agent removes a sentence, the cut lands on the word boundary, not somewhere inside a syllable.

Deliberately, editing transcript text does not delete footage. Content removal always goes through a typed request, which means every cut is one you asked for and can see in the reply — and any cut range can be restored later by asking. How the transcript, timeline, and versions fit together is covered in the timeline and transcript docs.

Honest Limits of Editing by Text

Text-based editing in Valmera has real boundaries, and knowing them saves you time. Captions render into the video itself — there is no subtitle-file import or export. Each request produces one deliverable: one program, one preview per turn, one export — the agent will not generate ten clips from one message. The one-time analysis on a long upload takes a while, with visible progress, before the conversational part begins. And the agent only claims what it can verify: a server-side honesty layer checks every reply against the edit actually made, so if a request is outside its abilities — a true crossfade dissolve, say, where it offers one of its 7 transition styles instead — it tells you rather than pretending.

More Than a Text-Based Editor

The same typed interface drives the whole finishing stack: transitions, color grades and one-request looks, targeted zooms, speed changes, music, sound effects, overlays, and on-screen text. Text-based editing is the front door — the editor behind it does the complete job. Browse all AI video editing tools to see every job it handles, or check pricing when the free credits stop being enough.

Type Your First Edit

Upload a video and tell the agent what to change. Free plan — no credit card required.

Start Editing Free →
See pricing →

Frequently Asked Questions

Text-based video editing means editing a video through words instead of a timeline. The editor transcribes what was said, and that text becomes the interface: in transcript-deletion editors like Descript you delete words and the video cuts follow, while in Valmera you type instructions — "remove all the silences", "add karaoke captions" — and an AI agent performs the edit on the actual footage.
Descript pioneered transcript-deletion editing: you hand-pick the words to delete, and its Underlord AI co-editor assists along the way — but you still drive the edit. Valmera flips the division of labor: you describe the outcome and the agent makes every cut itself, then a server-side honesty layer verifies each reply against what actually changed. Descript's Creator plan is $24/month billed annually with limited Underlord use on the free tier; Valmera's free plan includes 20 credits daily plus a one-time 150-credit bonus.
In Valmera, editing the transcript corrects the text — fix a misheard word and the captions re-render from the correction. Removing content works through typed requests instead: "cut the part where I repeat myself" or "remove everything before the intro". That split is deliberate — footage never disappears because of a stray text edit; every cut is something you asked for, and anything cut can be restored.
The analysis builds a word-level transcript with timings, and every cut snaps to word boundaries, so edits don't clip the start or end of a word. If a cut lands wrong, you say so in plain English — "you cut too much, bring back the joke about the dog" — and the agent restores it.
Yes. Valmera's free plan includes 20 credits every day (they reset daily and don't accumulate) plus a one-time 150-credit welcome bonus, with no credit card required. Edits charge credits in proportion to the work actually done, so simple edits cost the least. Paid plans are Plus at $20/month for 800 credits per billing cycle and Pro at $50/month for 2,400.
No. You describe outcomes, not operations — "make the intro punchier", "it drags in the middle", "caption it for people watching muted". The agent reads the transcript, looks at actual frames before and after its changes, and translates your intent into concrete edits. If something is outside what it can do, it says so instead of pretending.

Edit Video by Typing

The agentic AI video editor: describe the edit, review the preview, export at source quality. Free plan available.

Try Valmera Free →
See pricing →

Related Articles

Remove Silence from Video
The classic text-based edit: one request cuts every silent stretch, word-boundary-safe.
Add Captions to Video
Word-accurate captions with 11 presets, 12 fonts, and per-word emphasis — styled by request.
Timeline & Transcript Docs
How the editable transcript, timeline blocks, and version history work.
All AI Video Editing Tools
Every job Valmera does — cutting, captions, audio, motion, and framing — in one editor.