← Home
GLOSSARY

Published · Updated

J-Cut and L-Cut

A J-cut and an L-cut are the two forms of a split edit — an edit where the picture and the sound change at different moments instead of together. In a J-cut the sound of the incoming shot starts first and the picture catches up. In an L-cut the picture cuts first and the sound of the outgoing shot carries on underneath the new image. The letters describe the shape one clip makes on a timeline that draws video above audio: the audio block sticks out to the left, hooking under the previous shot (J), or to the right, running under the next one (L). Both exist for the same reason — a cut in the picture is perceptually cheap, but a hard cut in continuous sound is heard as a bump, so letting one track run across the other's cut hides the join.

IN ONE SENTENCE

A split edit where sound leads the picture (J-cut) or trails it (L-cut), so the audio runs continuously across a visual cut and hides it.

The Four Edit Points

Every cut between two clips has four edit points, not two: a video out and an audio out on the outgoing clip, and a video in and an audio in on the incoming one. A straight cut aligns all four at the same frame. A split edit unaligns them, and there are exactly two ways to do that — sound early, or sound late. Those two cases are the J-cut and the L-cut. There is no third case, which is why the vocabulary has only two letters in it.

J-cutL-cut
What crosses the cutThe incoming shot's sound, earlyThe outgoing shot's sound, late
Which changes firstSoundPicture
Which clip draws the letterThe incoming clipThe outgoing clip
Where the audio block sticks outLeft — under the previous shotRight — under the next shot
Sound-department namePre-lapHangover / overlap
Typical jobAnnounce a place or a speaker before you see themStay on a listener after the line has ended

The letters are geometry, not metaphor. Draw the incoming clip of a J-cut on a timeline with video on top: the audio begins at the earlier time and the video at the later one, giving a tall block on the right with a foot hooking left along the bottom. Draw the outgoing clip of an L-cut the same way: video ends first, audio runs on, giving a tall block on the left with a foot extending right along the bottom. Both letters have their foot at the bottom for one boring reason — audio tracks are drawn under video tracks.

How It Actually Works

Why sound is the track you protect

The two tracks are not equally tolerant of being cut. A picture cut is something the visual system handles constantly; a change of viewpoint costs almost nothing to accept. A splice in continuous sound is different: the noise floor, the room's reverb tail and the spectrum of the ambience all change instantaneously, and that discontinuity is heard as a bump or a click even when neither side has dialogue in it. So the working rule is asymmetric — move the picture cut freely, and try never to cut the sound in the same place.

How much overlap

Conventional ranges, not rules. Six to twelve frames (roughly 0.2–0.4 seconds) for a quick hand-off inside a conversation. Half a second to two seconds — 12 to 48 frames at 24 fps, 15 to 60 at 30 fps — for a normal dialogue split. Several seconds for a scene-transition pre-lap, where the next location's sound establishes it before the picture arrives. Under about four frames the offset does not register as a split edit; it just reads as a slightly soft cut. The ceiling is attention, not arithmetic: the overlap has to end before the viewer starts searching the frame for whoever is talking.

Handles, and why an L-cut can be impossible

An L-cut needs audio that exists past the picture out-point. If the take was trimmed to the exact frame of the picture cut there is nothing left to hang over, and the edit cannot be made without going back to the source media. This is why assistant editors ship media with handles — conventionally one to two seconds beyond each end — and why a project cut down from pre-trimmed clips is harder to smooth afterwards than one cut from full takes.

Room tone is the failure mode

Carrying shot A's audio under shot B's picture puts two different acoustic spaces next to each other. If they were recorded at different times, in different positions, or with different gain, the split edit does not remove the bump — it moves it, from the picture cut to the audio cut, where it is more audible. The standard fix is a continuous bed of room tone running under both shots so the noise floor never steps. A short audio crossfade at the audio edit point, typically two to six frames, kills a click but does not fix a mismatched room.

The sync constraint that makes it legal

A half-second of audio ahead of its picture would be a gross sync error if you could see the speaker. Published tolerances are tiny: ITU-R BT.1359-1 places the detection threshold at about 45 ms of audio leading the picture and 125 ms of audio lagging it, and ATSC IS-191 tightens the delivery spec to 15 ms lead and 45 ms lag. A one-second J-cut is more than twenty times past those numbers. It works only because there is nothing on screen to contradict it — during the overlap you are on the listener, a cutaway, an empty room or a wide shot with no readable mouth. The instant the speaking face is in frame, the overlap must already be finished.

Where they come from

The technique is old and the names are new. Overlapping sound was standard practice long before non-linear editing — a picture cut and a sound cut have never had to fall on the same frame — but the letters describe a software timeline that did not exist yet, which is why no one working on film called them J-cuts. The move also got easier: on a digital timeline, offsetting one edge of an edit is a drag rather than a re-conform, and split edits are correspondingly ordinary in contemporary cutting.

Why It Matters: One Real Edit

A two-person interview, two cameras, one continuous audio recording. The guest finishes an answer at 04:12.4. The host's next question starts at 04:13.1. Cut both tracks at 04:13.1 and the exchange works, but every hand-off in the episode now has the identical rhythm: answer ends, picture changes, question begins. A viewer learns that pattern quickly and starts predicting the next cut instead of listening to what is being said.

Now move only the picture cut, to 04:12.6 — half a second early, with the audio untouched. You are on the host while the guest's last word lands, and you are already watching a listener when the question starts. That single offset is an L-cut on the guest's audio and a J-cut on the host's at the same time. It costs one drag, it changes nothing anyone said, and it is the difference between a conversation and a tennis match.

The single-camera version of the same move is more common online and more useful. Cut a filler word or a long pause out of a talking head and the picture jumps — the body teleports a few centimetres. Hold the audio, replace the picture across the splice with b-roll, and the jump is gone; you have entered on an L-cut and left on a J-cut. That is the standard repair, and it is the reason filler-word removal and cutaways are usually the same job rather than two.

Common Mistakes

“The letter tells you which way the audio moves.”

It does not. The letter is the outline of one clip on a timeline, and the two letters are drawn by different clips: the J by the incoming clip, the L by the outgoing one. If your editor is configured with audio above video the shapes invert and the names stay the same, which is a good sign the mnemonic is about the interface rather than about the sound.

“A split edit is a kind of transition.”

There is no dissolve involved. Both the picture edit and the audio edit are hard cuts; they simply happen at different frames. Nothing fades, nothing blends, and no transition effect is applied at either point. The only fade that legitimately appears near a split edit is a two-to-six-frame audio crossfade to suppress a click, and that is a repair, not the technique.

“Unlink the clip, then slide the audio.”

Sliding a detached audio clip is how footage ends up permanently out of sync, because the offset survives every later move. The correct operation trims one edge of the edit — the audio out point or the video in point — while both halves stay locked to their own source timecode. Every major editor supports that without detaching anything.

“Longer overlaps are smoother.”

Past a couple of seconds an overlap stops smoothing the join and starts withholding information. If the audience can hear someone speaking and has no reason to be looking at what they are looking at, the cutaway is no longer hiding a cut; it is competing with the line.

“J-cuts are a dialogue technique.”

The most common J-cut in modern online video has no dialogue in it. It is music or ambience from the next section entering under the tail of the previous shot, so the sound has already moved before the picture does. Same mechanism, no speech required.

How This Works in Valmera

Valmera is an agentic video editor: you describe the edit and an AI agent performs it against an edit decision list. That data model decides exactly which split edits are available, so it is worth being specific about both halves.

What it does

  • The cutaway form, in one sentence of English. Ask for a full-frame cutaway — “show the laptop clip while I keep talking” — and the picture switches to another clip while the program's audio, speech and music both, keeps playing underneath. That is an L-cut into the cutaway and a J-cut back out of it, and it is the split edit most online video actually needs.
  • Pre-laps on music and voiceover. Music, sound effects and voiceover are positioned in output-timeline seconds, independent of the picture, so “bring the track in two seconds before the intro ends” is a J-cut on the score. Music ducks under speech while it overlaps.
  • Placement by description, not by scrubbing. Every upload is indexed into a word-level transcript, detected silences and shot boundaries, so “cut away right after he says margin” resolves to a real frame instead of an estimate, and the natural pauses to hide a picture cut in are already located.

What it does not do

  • No true split edit on the main footage's own tracks. The edit decision list keeps one source range per kept segment, so the main video's picture and its own sound come from the same range and move together. You cannot slide the main footage's audio half a second ahead of its own picture. Split edits are available through the cutaway, not through the main timeline.
  • A spliced insert is not a split edit. Inserting a clip adds time and pauses the talk; muting that insert makes the block silent rather than letting the speech continue under it. The full-frame cutaway is the one that keeps the audio running — they are different operations and worth asking for by effect.
  • A cutaway clip's own audio never plays. Video overlays are silent, so you cannot J-cut into the sound of the b-roll itself.
  • No dissolve, and no audio crossfade. The transition library is dip to black, dip to white, whip left, whip right, zoom punch, glitch and flash — there is no true crossfade between two shots, one style applies to every junction in scope rather than being picked per cut, and the two-to-six-frame audio crossfade a conventional editor uses to clean up a split point has no equivalent here.
  • Captions do not step out of the way by themselves. During a full-frame cutaway the burned-in captions keep rendering over it. You can ask for them to be suppressed for that window, but that is a second request; the cutaway does not make the decision for you.

Related Terms

  • B-roll The cutaway footage a picture-only split edit usually cuts to.
  • Jump cut The visible discontinuity a split edit is most often deployed to hide.
  • Edit decision list The data structure that decides whether a split edit is even expressible.
  • Word-level timestamps Why “cut away right after he says margin” resolves to a real frame.
  • Silence detection Finds the pauses that are the natural place to put the picture cut.
  • Shot detection Finds the existing picture cuts you might want to split.
  • Sidechain ducking Keeps music under a voice that is now running across a picture cut.
  • Text-based video editing Editing by transcript, which moves picture and sound as one unit.

Browse the full video editing glossary, or see the timeline and transcript documentation for how cuts and cutaways are represented in the editor.

Frequently Asked Questions

Both are split edits — edits where the audio and the picture cut at different times. The difference is direction. In a J-cut the sound arrives before the picture: you hear the next shot while you are still looking at the current one. In an L-cut the picture arrives before the sound is finished: you cut to the next shot while the previous shot's audio keeps running underneath. One offset often creates both at once — moving a picture cut half a second earlier is an L-cut on the outgoing speaker's audio and a J-cut on the incoming shot's picture, made in a single move.
The names describe the shape a clip makes on a non-linear editing timeline that draws video on top and audio underneath. In a J-cut the incoming clip's audio starts earlier than its picture, so the clip is a tall block on the right with a foot hooking back to the left along the bottom — the letter J. In an L-cut the outgoing clip's audio ends later than its picture, so it is a tall block on the left with a foot extending right along the bottom — the letter L. Both letters have their foot at the bottom because audio sits below video in every conventional timeline layout. The names are relatively recent and come from the software interface; the technique itself predates non-linear editing by decades and was called overlapping sound.
There is no standard, but conventional ranges exist. A quick conversational hand-off between two people is commonly 6 to 12 frames, which is roughly 0.2 to 0.4 seconds. A normal dialogue split runs about half a second to two seconds — 12 to 48 frames at 24 fps. A scene-transition pre-lap, where the next location's line or ambience establishes the new scene before you see it, can run several seconds. Below about four frames the offset stops reading as a split edit at all. The upper bound is set by attention rather than by a number: once the viewer starts hunting the frame for the source of the voice they are no longer following the content.
Effectively yes, from two different departments. Pre-lap is the screenwriting and sound term for audio from the next scene beginning before the picture cuts to it; J-cut is the editing-room term for the same event, named after its timeline shape. Pre-lap is usually used for the larger, story-level version — a line or a sound establishing the next scene — while J-cut covers everything down to a few frames of a reply arriving early inside one conversation. The reverse case has no widely used name; editors say L-cut, hangover, or just overlapping sound.
No, and the distinction is precise. Lip-sync tolerance is measured in milliseconds — ITU-R BT.1359-1 puts the detection threshold at roughly 45 ms of audio leading picture and 125 ms of audio lagging it, and ATSC IS-191 tightens the delivery spec further to 15 ms lead and 45 ms lag. A split edit is one to two orders of magnitude larger than that. It is not a sync error because there are no visible lips to contradict during the overlap: whoever is speaking across the join is off screen, or you are on a listener, a cutaway or an empty room. The moment the speaking mouth is in frame, the overlap has to be over.
No. With two cameras on a continuous audio recording you get L-cuts for free — the audio never cut in the first place, so every angle change is already a split edit, which is why multicam interviews sound smooth without any deliberate work. On a single camera the equivalent move is a cutaway: hold the speech and change the picture to b-roll, a product shot or a screen recording. That is the most common split edit in online video, and it is also the standard way to hide the picture jump left behind when you cut a filler word or a pause out of a single-camera talking head.
Valmera can make the cutaway form, which is the one most creators actually need: ask for a full-frame cutaway and the picture switches to another clip while the program's audio — speech and music — keeps playing. That is an L-cut into the cutaway and a J-cut back out of it. Music and voiceover are placed in output-timeline seconds, so a pre-lap like "bring the track in two seconds before the intro ends" is a J-cut on the score. What Valmera cannot do is a true split edit on the main video's own two tracks: its edit decision list keeps one source range per kept segment, so the main footage's picture and its own sound move together and cannot be offset from each other.

Try a Cutaway on Your Own Footage

Upload your footage, describe the cutaway, and the agent holds the audio while the picture changes. 50 free credits, no card.

Start free →
See pricing →

Related Articles

Text & Overlays
How full-frame cutaways and picture-in-picture overlays work, including the audio rules.
How to Edit Talking-Head Videos
Where cutaways go in a single-camera piece, and what to do with the jumps left by cleanup cuts.
Music & Audio
Music placement in output seconds, ducking under speech, voiceover and loudness mastering.
Edit a Podcast with AI
The multicam interview workflow, where every angle change is already a split edit.