← Home
GLOSSARY

Published · Updated

Video Editing Glossary

69 terms from the working vocabulary of video editing, defined precisely and grouped by the part of the craft they belong to. Every definition is a sentence you can use as-is. Nothing here is a placeholder, and nothing is padded out to look thorough.

Most glossaries in this field define the easy half and quietly skip the words that actually cost people time — handles, true peak, chroma subsampling, variable frame rate, the two unrelated meanings of keyframe. Those are the entries worth having, so they are here, and they are why an export sometimes looks worse than the preview did.

Sixteen terms carry enough history or detail to need a page of their own, and their names link through. The rest are defined here and nowhere else, deliberately: a reference that links to pages it has not written is worse than one that just answers.

Use the Vocabulary Instead of the Timeline

Type the edit in the words above — cut the silences, add karaoke captions, duck the music, reframe to 9:16 — and an AI agent performs it. 50 free credits, no card.

Start free →
See pricing →

Cutting & structure

The vocabulary of what plays, in what order, and where the joins fall. Most of it predates computers, and several words are fossils of physical tape.

Edit decision list (EDL)
A machine-readable list that describes a finished video as ordered events — each naming a source, an in and out point in that source's own clock, and where those frames land in the program. It holds decisions, never pixels.
Non-destructive editing
An editing model in which the original media is never altered. Every trim, effect and grade is stored as an instruction over untouched files, so any decision can be undone, revised, or re-applied to the originals at full quality.
Timeline
The horizontal view of a program, with time running left to right and stacked tracks for picture, dialogue, music and graphics. It is a drawing of the edit's underlying decisions, not the decisions themselves.
Cut
The instantaneous change from one shot to the next — the default join, and the only one that costs no screen time. Cutting well is almost entirely about where the join falls, not about what you put between the shots.
Jump cut
A cut between two shots of the same subject whose framing is too similar to read as a change of shot, so the subject appears to jump. Once treated as an error, now the native grammar of talking-head video.
J-cut and L-cut
Split edits, where audio and picture cut at different moments. In a J-cut the next scene's sound arrives before its picture; in an L-cut the outgoing sound continues underneath the incoming shot.
B-roll
Supplementary footage cut over the main take: the thing being described, the hands, the screen, the street. The name is a fossil of tape-era editing, where a dissolve physically required a second reel.
Cutaway
A shot that leaves the main subject for a moment — a listener, a detail, an object — and then returns. Every cutaway is B-roll, but its particular job is hiding a cut in the audio underneath it.
Stringout
A first assembly that keeps everything worth keeping, in rough order, with nothing polished. It exists to make the structure of a long recording visible, so the real cutting starts from something rather than from nothing.
Handles
The extra frames kept beyond a clip's in and out points, commonly a second at each end. Transitions consume them, and trimming a join later is impossible without them. Media delivered with no handles is media you cannot adjust.
Ripple trim
A trim that shifts everything after it and changes the program's total length, as opposed to a roll trim, which moves the join between two clips and leaves the duration alone. Most surprise sync problems are a ripple nobody expected.
Multi-cam
An edit built from several cameras covering one event, synchronised by timecode or by matching their audio waveforms, then cut between as angles. The synchronisation is the hard part; the cutting afterwards is ordinary.
Shot detection
Automatic detection of the boundaries already present in a video, by measuring how much consecutive frames differ. It finds hard cuts reliably and gradual dissolves less so, and it is how a tool knows an upload has already been edited.
Text-based video editing
Editing a video by editing its transcript: delete a sentence from the text and the matching footage leaves the program. It suits speech-driven video and is useless where the words are not the structure.

Sound

Audio is where amateur edits are identified fastest, and where most of the measurable vocabulary lives. Loudness, unlike framing, has a correct answer.

LUFS and loudness normalization
LUFS measures perceived loudness over time rather than instantaneous signal level. Normalization brings a program to a target — broadcast uses −23 LUFS, most streaming platforms sit near −14 — so listeners stop reaching for the volume between videos.
True peak
The highest level the reconstructed analogue waveform reaches between digital samples, measured in dBTP. It can exceed the sample peak, which is why loudness standards leave a margin: EBU R 128 asks for no more than −1 dBTP.
Sidechain ducking
Automatically lowering one signal whenever another is present — music dipping under speech. A compressor on the music is triggered by the dialogue track rather than by the music itself, so the dip follows the voice exactly.
Silence detection
Measuring where a recording sits below a loudness threshold for longer than a minimum duration, producing candidate ranges to cut. The threshold and the minimum gap are the entire craft: too aggressive and breaths disappear with the dead air.
Room tone
The sound of a silent location — air, traffic, the building itself. Recorded deliberately for thirty seconds on set, it patches gaps so that a cut in dialogue does not also cut to digital silence, which the ear notices immediately.
Filler words
The sounds that hold a speaking turn without carrying meaning: um, uh, er, like, you know. Removing every one of them makes speech sound machine-processed; removing the ones at sentence boundaries is usually enough.
Beat detection and BPM
Finding a track's tempo and the position of its beats, expressed in beats per minute. Cuts that land on beats read as intentional. The failure mode is the cut that lands near a beat but not on it, which reads as sloppy.
Gain staging
Setting the level of a signal at each stage so that nothing clips and nothing is buried in noise. Gain is the level entering a process; a fader sets the level leaving it. Confusing the two is the usual cause of a distorted mix.
Dynamic range compression
Reducing the distance between the loudest and quietest parts of a signal, controlled by threshold, ratio, attack and release. It makes quiet speech audible without making loud speech painful, at the cost of some liveliness.
Sound effects and Foley
Sound effects are library or generated sounds placed against picture. Foley is everyday sound performed and recorded to picture by a person — footsteps, cloth, handling. Both exist to fill the gap between what a camera records and what a scene needs.

Picture & colour

Two stages that get collapsed into one word. Correction has a right answer and grading has a taste, and doing them in the wrong order is why footage looks graded but still wrong.

Color grading
The creative stage of colour work: deciding what a program should look like and applying it consistently — contrast shape, palette, mood. It follows correction, and it is the difference between footage that is right and footage that has a look.
Color correction
The technical stage that comes first: neutralising casts, matching shots to one another, and putting exposure where it belongs. Its success condition is objective — skin reads as skin, white reads as white — which is exactly why it is not grading.
LUT
A lookup table mapping every input colour to an output colour. Technical LUTs convert between colour spaces; creative LUTs carry a look. A LUT is a fixed transform, so it cannot adapt to a shot it was not built for.
White balance and colour temperature
Colour temperature describes the colour of a light source in kelvin — roughly 3200K for tungsten, 5600K for daylight. White balance is the compensation applied so that a white object reads as white under that particular light.
Exposure
How much light the sensor received, and therefore where the image sits on the brightness scale. Detail lost to clipped highlights cannot be recovered; detail crushed into the shadows sometimes can, at the cost of visible noise.
Contrast
The distance between the darkest and brightest parts of an image. Raising it deepens shadows and brightens highlights, which reads as punchy and also destroys detail at both ends. Most grading decisions are contrast decisions in disguise.
Saturation
The intensity of colour, independent of brightness. Pushing it is the quickest way to make an image look cheap, because skin tones saturate faster than anything else in the frame and go orange well before the rest of the picture moves.
Log footage
Footage recorded with a logarithmic transfer curve so that more of the camera's dynamic range survives into the file. It looks flat and grey until it is graded, and that flatness is headroom rather than a defect.
Rec. 709
The colour space of HD video: the primaries, white point and transfer function a normal screen assumes. Grade outside it and deliver anyway, and every viewer sees something other than what you approved.
Film grain
The visible texture of photochemical film, now added digitally. It hides banding in gradients, gives compression something to bite on other than faces, and makes an over-clean digital image read as photographed rather than rendered.

Framing & motion

Where the frame is, what it contains, and how either changes over time. This is the section where the rise of vertical video changed the working vocabulary most.

Aspect ratio
The ratio of a frame's width to its height — 16:9 for landscape, 9:16 for vertical, 1:1 square, 4:5 for feed video. Changing it is not a resize: something in the original frame has to be given up, or padded back in.
Auto-reframe
Automatically re-cropping a wide frame into a taller one by tracking whatever matters, usually the speaker, so the subject stays in frame. The hard cases are two people at once, and subjects who sit near the edge.
Letterbox and pillarbox
Bars added to fit one aspect ratio inside another: horizontal bars above and below when a wide image sits in a taller frame, vertical bars at the sides when a tall image sits in a wider one.
Safe area
The region of the frame guaranteed to survive display — historically the inner 90% for action and 80% for titles on CRT screens. On social platforms the modern equivalent is whatever the app's own buttons and captions do not cover.
Headroom and lead room
Headroom is the space between the top of a subject's head and the top of the frame. Lead room is the space in front of the direction they face or move. Too much of either reads as an accident rather than a choice.
Punch-in
A cut or push to a tighter framing of the same shot, used to add emphasis or to hide a cut. In a single-camera talking-head edit it is the only way to change shot size, so it does most of the visual work.
Ken Burns effect
Slow pan and zoom applied to a still image so that it holds screen time as motion. Named for the documentary technique of moving a camera across photographs, it is what stops a still from reading as frozen video.
Speed ramp
A speed change that eases in and out rather than switching abruptly — normal into slow motion and back. The ramp is what makes it read as an effect instead of a glitch, and audio is usually dropped through it.
Slow motion
Playing footage more slowly than it was shot. Genuine slow motion comes from a high frame rate at capture; slowing ordinary footage either duplicates frames, which judders, or synthesises new ones, which invents detail that was never recorded.
Transition
Any join between shots other than a cut: dissolve, fade, wipe, whip. Transitions consume handles and screen time, and reaching for one is usually a sign that the cut underneath it is not working.

Text & captions

Captions are the spoken words; text elements are designed graphics. They are separate systems that compete for the same part of the frame, and the words for them are routinely swapped.

Burned-in captions vs subtitle files
Burned-in captions are rendered into the picture and cannot be switched off, restyled or translated. A subtitle file is separate text with timings that the player draws — accessible and searchable, but at the mercy of the player's own styling.
Word-level timestamps
A transcript in which every individual word carries its own start and end time, rather than a whole sentence carrying one. This is what makes cutting to a word boundary, karaoke captions and per-word emphasis possible at all.
SRT and WebVTT
The two common subtitle file formats. SRT is numbered blocks with start and end timecodes plus plain text; WebVTT adds styling, positioning and metadata, and is the format HTML5 video expects.
Karaoke captions
Captions that reveal or highlight each word as it is spoken, instead of showing a whole line at once. They require word-level timestamps, and they hold attention mostly by giving the eye something that moves.
Lower third
A text graphic in the lower portion of the frame identifying who is speaking or what is being shown. Named for the area it occupies, and one of the very few graphics with a genuinely conventional position.
Title card
A full-frame text card, usually on a plain background, marking a beginning or a section boundary. It buys a beat of silence and separates chapters more cleanly than any transition does.
Kinetic typography
Text that animates as a designed element rather than as a caption — scaling, sliding, snapping onto a beat. It is a graphic that happens to be made of words, and it competes with captions for the same screen space.
Reading speed
How fast subtitle text can actually be read, measured in characters per second. Broadcast guidelines cluster around 15 to 20 CPS for adults. A line that is accurate but on screen too briefly is still unreadable.

AI and automation

The newest vocabulary, and the least stable. Several of these words are used by vendors to mean whatever is most flattering, so each definition below states the mechanism rather than the marketing.

Transcription and ASR
Automatic speech recognition converts speech to text with timings. Accuracy depends on audio quality, accent, vocabulary and overlapping speakers far more than on the length of the recording, and proper nouns are what fail first.
Word error rate (WER)
The standard accuracy measure for transcription: substitutions plus insertions plus deletions, divided by the number of words in the reference. It weighs every word equally, which is why a low WER can still ruin somebody's name.
Speaker diarization
Working out who spoke when, and labelling the transcript accordingly. It is a separate problem from recognising the words, it degrades badly on overlapping speech, and it usually knows there are two speakers without knowing who they are.
Indexing
The one-off analysis pass a tool runs over an upload before any edit — transcript, silences, shot boundaries, sampled frames. It is paid for once and reused by every later request, which is why the first wait is the long one.
Agentic video editing
Editing in which an AI agent performs the work itself: it plans a sequence of operations from a described outcome, executes them, inspects the result, and revises. The distinguishing feature is the return path, not the presence of AI.
Model Context Protocol (MCP)
An open protocol for exposing tools to AI assistants over one standard interface, so a model can operate an external system directly. For video it means the editor becomes something an assistant can drive from inside a conversation.
Text-to-video generation
Synthesising footage that was never filmed, from a written prompt. It is a different category from editing, which rearranges and treats footage you already have, and the two are routinely conflated in marketing copy.
Inpainting
Reconstructing a region of an image or video from its surroundings, so that whatever was there appears never to have been. In video the reconstruction has to stay consistent across frames, which is what makes it much harder than in stills.
Clip ranking
Automatically scoring segments of a long recording for how well they might perform as standalone clips. The score is a prediction from surface features — hooks, pacing, sentiment — and it has no access to your actual audience.

Delivery & formats

What the file is, once the editing is over. This is the section people skip until an export looks worse than the preview, which is almost always explained by one of these eight words.

Container and codec
A codec decides how picture and sound are compressed; a container decides how the compressed streams are packaged together with timing and metadata. MP4, MOV and MKV are containers. H.264 and AAC are codecs that live inside them.
H.264
The video codec almost everything can play, also called AVC. Newer codecs compress better at equal quality, but H.264 remains the safe delivery choice because hardware decoding for it is effectively universal.
Bitrate
How much data each second of video is allowed. Constant bitrate spends the same everywhere; variable bitrate spends more on complex shots and less on simple ones, which is why quality per megabyte is a scene-dependent question.
Keyframe (I-frame) and GOP
A keyframe is a fully self-contained frame; the frames after it are stored as differences from it, and the group between two keyframes is a GOP. Cutting anywhere except on a keyframe requires re-encoding, which is why some cuts cost quality.
Frame rate and VFR
How many frames each second holds. Constant frame rate fixes that number; variable frame rate lets it change during recording, which phones and screen recorders do routinely, and which causes audio drift when a tool assumes otherwise.
Resolution
The pixel dimensions of a frame, such as 1920x1080 or 3840x2160. It bounds detail without creating it: an out-of-focus 4K shot carries less real information than a sharp 1080p one, and upscaling adds pixels rather than detail.
Chroma subsampling
Storing colour at lower resolution than brightness, because eyes are less sensitive to it. 4:2:0 halves colour in both directions and is what nearly all delivery uses. It is also why saturated red text edges look ragged.
Proxy editing
Editing against small, easy-to-decode stand-in files and applying the finished decisions back to the originals for the final render. Speed while you work, full quality when you deliver — the decisions are the thing that transfers.

Six Words That Mean Two Different Things

The genuinely confusing part of this vocabulary is not the rare words. It is the six common ones that carry two unrelated meanings inside the same workflow, where reading the wrong one costs an hour before anybody notices.

Keyframe

1. In animation: a value pinned at a moment in time, with the software interpolating between pins. “Keyframe the zoom” means this one.

2. In compression: a fully self-contained frame that later frames are stored as differences from. “Cut on a keyframe” means this one.

Cut

1. The join itself — the instant one shot becomes the next.

2. The act of removing material, and also the noun for a whole version of the edit (a rough cut, a final cut). Three meanings, one syllable.

Gain

1. The level of a signal entering a process, set before anything else happens to it.

2. Loosely used for volume, which is the level leaving. Turning down the wrong one either wastes headroom or leaves the distortion already baked in.

Render

1. Producing a preview so you can watch what you have decided.

2. Producing the deliverable. If a tool uses one word for both, ask which file the second one starts from — a preview and an export should not read from the same source.

Frame

1. One still picture out of the sequence.

2. The visible rectangle itself, as in framing, reframe and out of frame. “In the frame at frame 240” is a legitimate sentence.

Track

1. A lane on the timeline holding clips.

2. A piece of music. “Swap the track” is ambiguous in every editor ever built.

The Two Clocks

Almost every timing mistake in editing is one clock being read as the other. Source time is a position inside an original file, on that file's own timeline. Program time — record time, in the tape-era term — is a position in the finished video. They agree only until the first cut, and after that they never agree again.

This is why an edit decision list carries four timecodes per event rather than two: source in, source out, record in, record out. It is why a transcript search returns a source timestamp that has to be translated before it means anything on the timeline. And it is why "cut the bit at 4:20" is an ambiguous instruction the moment a single earlier cut exists — 4:20 of the raw recording and 4:20 of the current edit are different moments.

Valmera's agent keeps both clocks explicitly. The keep list, the transcript, speed spans and volume all live in source seconds; the assembled program is addressed in output seconds; and any timestamp used in an edit has to come from a tool result rather than be invented. Cuts then snap to word-level timestamps, so a range is a word boundary instead of an approximation.

Where This Vocabulary Came From

A surprising amount of it is fossilised hardware. B-roll is named for the second tape reel a dissolve physically required, because you cannot dissolve a reel into itself. Handles exist because a machine needed frames to overlap during a transition. Reel names in an EDL are capped at eight characters because that is what the hardware could store. The J-cut and L-cut are named after the shapes the audio and video blocks make on a timeline, which is a computer-era name for a much older technique.

Some of it inverted. The jump cut was a continuity error for most of a century, then became a stylistic signature, and is now simply what talking-head video looks like — a term whose definition stayed constant while its verdict flipped completely.

And a slice of it is only a few years old. Word-level timestamps, silence detection, auto-reframe and text-based editing all describe capabilities that did not exist as consumer features before machine transcription got cheap. They are the terms that separate a modern editing vocabulary from a 2015 one.

How Valmera Uses These Words

Valmera is an agentic AI video editor: you upload real footage — up to 14 GB or 3 hours, in MP4, MOV, MKV or WebM — describe the edit in plain English, and an AI agent performs it. The vocabulary above is not decoration on that; it is the interface. Terms map to operations directly.

  • Indexing runs once per upload — a word-level transcript, silence measurement, shot boundaries, and labeled frame tiles the agent actually reads. Every later request reuses it.
  • The edit decision list is what the agent writes to. Nothing it does touches pixels, your original file is never modified, and anything cut can be restored by asking.
  • Proxy and export are separate. Previews render from a small proxy so they are fast; the deliverable is rendered from the original upload at source quality as an H.264 MP4.
  • Loudness and ducking are measured, not eyeballed. Music ducks under speech by sidechain compression, and one request masters the program to a −14 LUFS target.
  • Reframing crops or pads to 16:9, 9:16, 1:1 or 4:5 and never upscales, with captions rescaled to the new aspect ratio.
  • Colour is two stages. Six grade presets plus custom exposure, contrast, saturation, temperature and tint, with finishing effects such as grain and vignette applied over windows rather than the whole program.
  • Speed spans run 0.25x to 4x with pitch preserved. Slow motion duplicates frames rather than synthesising new ones, which is the honest version of the term.

The same toolset is published as an MCP server, so Claude can perform the edit from inside a conversation using exactly these words.

Terms Valmera Does Not Deliver

Several words defined above describe things Valmera genuinely cannot do. Defining a term is not a claim to support it, and it is worth being explicit about which is which.

  • Multi-cam. No synchronisation and no angle switching. A multi-camera shoot has to arrive as one cut file.
  • Transitions. No true crossfade or dissolve, and one transition style applies to every cut rather than being chosen per cut.
  • SRT and WebVTT. Captions are burned into the picture. There is no subtitle file import or export, and no chapter metadata. A standalone SRT to VTT converter is published separately as a free utility.
  • Speaker diarization output and per-speaker levelling. Neither is produced as a deliverable.
  • Audio repair. No denoise or "studio sound", and no separating music out of a track that was already baked in.
  • Motion tracking. Overlays and stickers do not follow a moving subject.
  • Text-to-video. Valmera edits footage you upload. Short generated clips and stills can be spliced in, but a whole video from a prompt is a different category of tool.
  • Clip ranking. One deliverable per request, not a scored batch of candidate clips.
  • Custom fonts and AI music generation. Bundled font families only, and a bundled royalty-free music library or your own uploaded track.

The agent knows these edges too — out-of-scope requests are refused rather than faked, and every reply is verified server-side against the edits actually recorded. If you need any of the above, a traditional non-linear editor is the right tool and this is not it.

Frequently Asked Questions

A reference of the working vocabulary of video editing: the terms editors use for cuts and structure, sound, colour, framing, captions, automation and file delivery. This one defines 69 terms across seven sections, each in a single sentence, and links to longer pages for the sixteen terms that carry enough history or detail to need one.
Six carry most of the weight. Cut (the join between shots), B-roll (footage laid over the main take), jump cut (a cut between two too-similar framings of the same subject), aspect ratio (the shape of the frame you are delivering), LUFS (the loudness target your program should hit), and non-destructive editing (the principle that the original file is never altered). Understanding those six makes the rest of the vocabulary read as variations rather than as new ideas.
Every cutaway is B-roll, but not every use of B-roll is a cutaway. B-roll is the category: supplementary footage that is not the main take. A cutaway is a specific job that footage does — leaving the subject for a moment and coming back — and its most common purpose is not illustration at all, but hiding a cut in the audio underneath it so the join is never seen.
Correction is technical and comes first: neutralise casts, match shots to each other, put exposure where it belongs. Its success condition is objective — white reads as white, skin reads as skin. Grading is creative and comes second: deciding what the program should look like and applying it consistently. Doing them in the wrong order is why footage can look heavily graded and still look wrong, because a look applied over an unmatched sequence just makes the mismatch louder.
LUFS is loudness units relative to full scale — a measure of perceived loudness over time, rather than the instantaneous level a peak meter shows. Broadcast standards target −23 LUFS; most streaming and social platforms normalise playback to somewhere near −14 LUFS, which is why a video mastered much louder than that gets turned down anyway and only loses its dynamics in exchange. Keep true peak at or below −1 dBTP.
They are better at different things. A subtitle file is accessible, searchable, translatable and switchable, and it is the right answer on a platform whose player renders it. Burned-in captions cannot be switched off — which is the point on feeds that autoplay muted with no caption track, and the only way to guarantee a specific type design survives to the viewer. The cost is real: no translation, no toggle, and no text for a search engine to read.
Several, deliberately listed here rather than omitted. Valmera has no multi-cam synchronisation, no true crossfade or dissolve (and one transition style applies to every cut rather than per-cut choice), no SRT or WebVTT import or export because captions are burned in, no speaker diarization output or per-speaker levelling, no audio denoise, no motion-tracked overlays, no custom font uploads, and no AI music generation. It produces one deliverable per request rather than a ranked batch of clips.
Yes, and where a word is ambiguous the tool asks rather than guesses. Valmera's agent edits an edit decision list, previews render from a proxy while the final export is rendered from the original upload at source quality, cuts snap to word-level timestamps, music ducks under speech by sidechain compression, and reframing crops or pads to 16:9, 9:16, 1:1 or 4:5 without ever upscaling. The vocabulary on this page is the vocabulary you can type at it.

Edit Your First Video Free

50 credits on signup, no card required. Describe the edit in the words above and watch the agent perform it.

Start free →
See pricing →

Related Articles

Agentic Video Editor
The category page: what it means for the agent, rather than you, to be the one holding the tool.
Getting Started
Upload, describe the edit, judge the preview, export. The vocabulary above, used in order.
All Editing Tools
Every editing job Valmera can do, one page each — the practical counterpart to these definitions.
Captions
Presets, fonts, animations and the editable transcript, and why the captions are burned in.
Music & Audio
Ducking, per-layer gain, voiceover and one-request mastering to a loudness target.
AI Video Editing: The Complete Guide
The longer read that puts this vocabulary to work on an actual edit, start to finish.