J-Cut and L-Cut
A J-cut and an L-cut are the two forms of a split edit — an edit where the picture and the sound change at different moments instead of together. In a J-cut the sound of the incoming shot starts first and the picture catches up. In an L-cut the picture cuts first and the sound of the outgoing shot carries on underneath the new image. The letters describe the shape one clip makes on a timeline that draws video above audio: the audio block sticks out to the left, hooking under the previous shot (J), or to the right, running under the next one (L). Both exist for the same reason — a cut in the picture is perceptually cheap, but a hard cut in continuous sound is heard as a bump, so letting one track run across the other's cut hides the join.
A split edit where sound leads the picture (J-cut) or trails it (L-cut), so the audio runs continuously across a visual cut and hides it.
The Four Edit Points
Every cut between two clips has four edit points, not two: a video out and an audio out on the outgoing clip, and a video in and an audio in on the incoming one. A straight cut aligns all four at the same frame. A split edit unaligns them, and there are exactly two ways to do that — sound early, or sound late. Those two cases are the J-cut and the L-cut. There is no third case, which is why the vocabulary has only two letters in it.
| J-cut | L-cut | |
|---|---|---|
| What crosses the cut | The incoming shot's sound, early | The outgoing shot's sound, late |
| Which changes first | Sound | Picture |
| Which clip draws the letter | The incoming clip | The outgoing clip |
| Where the audio block sticks out | Left — under the previous shot | Right — under the next shot |
| Sound-department name | Pre-lap | Hangover / overlap |
| Typical job | Announce a place or a speaker before you see them | Stay on a listener after the line has ended |
The letters are geometry, not metaphor. Draw the incoming clip of a J-cut on a timeline with video on top: the audio begins at the earlier time and the video at the later one, giving a tall block on the right with a foot hooking left along the bottom. Draw the outgoing clip of an L-cut the same way: video ends first, audio runs on, giving a tall block on the left with a foot extending right along the bottom. Both letters have their foot at the bottom for one boring reason — audio tracks are drawn under video tracks.
How It Actually Works
Why sound is the track you protect
The two tracks are not equally tolerant of being cut. A picture cut is something the visual system handles constantly; a change of viewpoint costs almost nothing to accept. A splice in continuous sound is different: the noise floor, the room's reverb tail and the spectrum of the ambience all change instantaneously, and that discontinuity is heard as a bump or a click even when neither side has dialogue in it. So the working rule is asymmetric — move the picture cut freely, and try never to cut the sound in the same place.
How much overlap
Conventional ranges, not rules. Six to twelve frames (roughly 0.2–0.4 seconds) for a quick hand-off inside a conversation. Half a second to two seconds — 12 to 48 frames at 24 fps, 15 to 60 at 30 fps — for a normal dialogue split. Several seconds for a scene-transition pre-lap, where the next location's sound establishes it before the picture arrives. Under about four frames the offset does not register as a split edit; it just reads as a slightly soft cut. The ceiling is attention, not arithmetic: the overlap has to end before the viewer starts searching the frame for whoever is talking.
Handles, and why an L-cut can be impossible
An L-cut needs audio that exists past the picture out-point. If the take was trimmed to the exact frame of the picture cut there is nothing left to hang over, and the edit cannot be made without going back to the source media. This is why assistant editors ship media with handles — conventionally one to two seconds beyond each end — and why a project cut down from pre-trimmed clips is harder to smooth afterwards than one cut from full takes.
Room tone is the failure mode
Carrying shot A's audio under shot B's picture puts two different acoustic spaces next to each other. If they were recorded at different times, in different positions, or with different gain, the split edit does not remove the bump — it moves it, from the picture cut to the audio cut, where it is more audible. The standard fix is a continuous bed of room tone running under both shots so the noise floor never steps. A short audio crossfade at the audio edit point, typically two to six frames, kills a click but does not fix a mismatched room.
The sync constraint that makes it legal
A half-second of audio ahead of its picture would be a gross sync error if you could see the speaker. Published tolerances are tiny: ITU-R BT.1359-1 places the detection threshold at about 45 ms of audio leading the picture and 125 ms of audio lagging it, and ATSC IS-191 tightens the delivery spec to 15 ms lead and 45 ms lag. A one-second J-cut is more than twenty times past those numbers. It works only because there is nothing on screen to contradict it — during the overlap you are on the listener, a cutaway, an empty room or a wide shot with no readable mouth. The instant the speaking face is in frame, the overlap must already be finished.
Where they come from
The technique is old and the names are new. Overlapping sound was standard practice long before non-linear editing — a picture cut and a sound cut have never had to fall on the same frame — but the letters describe a software timeline that did not exist yet, which is why no one working on film called them J-cuts. The move also got easier: on a digital timeline, offsetting one edge of an edit is a drag rather than a re-conform, and split edits are correspondingly ordinary in contemporary cutting.
Why It Matters: One Real Edit
A two-person interview, two cameras, one continuous audio recording. The guest finishes an answer at 04:12.4. The host's next question starts at 04:13.1. Cut both tracks at 04:13.1 and the exchange works, but every hand-off in the episode now has the identical rhythm: answer ends, picture changes, question begins. A viewer learns that pattern quickly and starts predicting the next cut instead of listening to what is being said.
Now move only the picture cut, to 04:12.6 — half a second early, with the audio untouched. You are on the host while the guest's last word lands, and you are already watching a listener when the question starts. That single offset is an L-cut on the guest's audio and a J-cut on the host's at the same time. It costs one drag, it changes nothing anyone said, and it is the difference between a conversation and a tennis match.
The single-camera version of the same move is more common online and more useful. Cut a filler word or a long pause out of a talking head and the picture jumps — the body teleports a few centimetres. Hold the audio, replace the picture across the splice with b-roll, and the jump is gone; you have entered on an L-cut and left on a J-cut. That is the standard repair, and it is the reason filler-word removal and cutaways are usually the same job rather than two.
Common Mistakes
“The letter tells you which way the audio moves.”
It does not. The letter is the outline of one clip on a timeline, and the two letters are drawn by different clips: the J by the incoming clip, the L by the outgoing one. If your editor is configured with audio above video the shapes invert and the names stay the same, which is a good sign the mnemonic is about the interface rather than about the sound.
“A split edit is a kind of transition.”
There is no dissolve involved. Both the picture edit and the audio edit are hard cuts; they simply happen at different frames. Nothing fades, nothing blends, and no transition effect is applied at either point. The only fade that legitimately appears near a split edit is a two-to-six-frame audio crossfade to suppress a click, and that is a repair, not the technique.
“Unlink the clip, then slide the audio.”
Sliding a detached audio clip is how footage ends up permanently out of sync, because the offset survives every later move. The correct operation trims one edge of the edit — the audio out point or the video in point — while both halves stay locked to their own source timecode. Every major editor supports that without detaching anything.
“Longer overlaps are smoother.”
Past a couple of seconds an overlap stops smoothing the join and starts withholding information. If the audience can hear someone speaking and has no reason to be looking at what they are looking at, the cutaway is no longer hiding a cut; it is competing with the line.
“J-cuts are a dialogue technique.”
The most common J-cut in modern online video has no dialogue in it. It is music or ambience from the next section entering under the tail of the previous shot, so the sound has already moved before the picture does. Same mechanism, no speech required.
How This Works in Valmera
Valmera is an agentic video editor: you describe the edit and an AI agent performs it against an edit decision list. That data model decides exactly which split edits are available, so it is worth being specific about both halves.
What it does
- The cutaway form, in one sentence of English. Ask for a full-frame cutaway — “show the laptop clip while I keep talking” — and the picture switches to another clip while the program's audio, speech and music both, keeps playing underneath. That is an L-cut into the cutaway and a J-cut back out of it, and it is the split edit most online video actually needs.
- Pre-laps on music and voiceover. Music, sound effects and voiceover are positioned in output-timeline seconds, independent of the picture, so “bring the track in two seconds before the intro ends” is a J-cut on the score. Music ducks under speech while it overlaps.
- Placement by description, not by scrubbing. Every upload is indexed into a word-level transcript, detected silences and shot boundaries, so “cut away right after he says margin” resolves to a real frame instead of an estimate, and the natural pauses to hide a picture cut in are already located.
What it does not do
- No true split edit on the main footage's own tracks. The edit decision list keeps one source range per kept segment, so the main video's picture and its own sound come from the same range and move together. You cannot slide the main footage's audio half a second ahead of its own picture. Split edits are available through the cutaway, not through the main timeline.
- A spliced insert is not a split edit. Inserting a clip adds time and pauses the talk; muting that insert makes the block silent rather than letting the speech continue under it. The full-frame cutaway is the one that keeps the audio running — they are different operations and worth asking for by effect.
- A cutaway clip's own audio never plays. Video overlays are silent, so you cannot J-cut into the sound of the b-roll itself.
- No dissolve, and no audio crossfade. The transition library is dip to black, dip to white, whip left, whip right, zoom punch, glitch and flash — there is no true crossfade between two shots, one style applies to every junction in scope rather than being picked per cut, and the two-to-six-frame audio crossfade a conventional editor uses to clean up a split point has no equivalent here.
- Captions do not step out of the way by themselves. During a full-frame cutaway the burned-in captions keep rendering over it. You can ask for them to be suppressed for that window, but that is a second request; the cutaway does not make the decision for you.
Related Terms
- B-roll — The cutaway footage a picture-only split edit usually cuts to.
- Jump cut — The visible discontinuity a split edit is most often deployed to hide.
- Edit decision list — The data structure that decides whether a split edit is even expressible.
- Word-level timestamps — Why “cut away right after he says margin” resolves to a real frame.
- Silence detection — Finds the pauses that are the natural place to put the picture cut.
- Shot detection — Finds the existing picture cuts you might want to split.
- Sidechain ducking — Keeps music under a voice that is now running across a picture cut.
- Text-based video editing — Editing by transcript, which moves picture and sound as one unit.
Browse the full video editing glossary, or see the timeline and transcript documentation for how cuts and cutaways are represented in the editor.
Frequently Asked Questions
Try a Cutaway on Your Own Footage
Upload your footage, describe the cutaway, and the agent holds the audio while the picture changes. 50 free credits, no card.
Start free →