← Home
TOOL

Published · Updated

Put Text Behind the Subject in a Video

A title painted on the wall behind you, and you walk in front of the letters. It is one sentence in Valmera, and it is a real depth composite — not a fade, not a transparency.

To put text behind the subject of a video, upload the footage to Valmera and type "put my name behind me as I walk across at 0:12". The AI agent cuts the person out of every frame in that window with a video-matting model, draws the words onto the picture, and lays the cut-out person back over the letters. The person occludes the text the way they would occlude anything genuinely behind them. There is no rotoscoping, no green screen, and no mask to keyframe.

Try the Depth Text Effect Free

50 credits on signup, no credit card. Upload footage, type one sentence, watch the preview.

Put text behind you →
See pricing →

How to Put Text Behind a Subject in a Video

  1. 1
    Upload footage with a person in it
    Drag in any MP4, MOV, MKV or WebM up to 14 GB or 3 hours. Valmera runs a one-time analysis: a word-level transcript, shot detection, silence detection, and labeled frame tiles the agent actually looks at — so it already knows which moments have someone moving across the frame.
  2. 2
    Say what the words are and where they go
    Type the literal sentence: "put VALMERA behind me as I walk across at 0:12, for about 3 seconds". You can also say it by content instead of by clock — "put my name behind me on the shot where I step away from the desk" — because the agent can see the frames.
  3. 3
    Read the measurement in the reply
    The agent reports what it actually measured: how much of the frame the subject covers, and how much of the title's width they interrupt at their most. If the subject never crosses the words, it says so rather than telling you an effect is there that you cannot see.
  4. 4
    Preview, adjust the type, export
    Watch the preview and refine in plain English — "make it bigger", "move it left so I cross the middle of the word", "use Anton instead". When it reads right, export a full-quality H.264 MP4 rendered from your original file.

The subject is measured frame by frame over the window you asked for, so a 3-second placement is quick and a 15-second one is the longest allowed.

EXAMPLE PROMPTS
  • "put my name behind me as I walk across at 0:12"
  • "big title behind the subject on the opening shot, 4 seconds"
  • "the word LAUNCH behind me in Anton, white, for 3 seconds at 0:45"
  • "move that text left so I cross the middle of it"
  • "make the behind text bigger — the letters are getting swallowed"

What "Behind" Actually Means Here

A renderer draws text on top of the picture. There is no "behind" until something knows which pixels are the person. That is the entire job, and it is why this effect is rare outside of desktop compositing apps: everything else about it — the font, the timing, the animation — already existed.

How text is composited behind a subjectFour stages. One: the original frame, a person in a shot with no cut and no speed ramp. Two: the words are drawn onto the picture, on top of the person. Three: a grayscale mask, white where the subject is and black everywhere else, measured frame by frame. Four: the masked subject is laid back over the words, so the letters pass behind the person.1 · THE SHOTOne take. No cut,no speed ramp.2 · THE WORDSBEHINDDrawn ON the picture,over everything.3 · THE MASKGrayscale silhouette,measured every frame.4 · THE DEPTHBEHINDOpaque over the words.Depth, not a fade.The subject’s pixels in step 4 are the render’s own — the mask only says WHERE.That is why a mask measured on a 540p proxy composites cleanly into a 4K export.
The four stages of a behind-subject title. Steps 2 and 4 are the same picture in a different order.

How the Matte Is Actually Built

When you ask for text behind the subject, Valmera does not guess. It measures a window of your footage and stores a grayscale mask clip — white where the subject is, black where they are not, one frame for every frame of the window. The renderer then draws the text onto a copy of its own picture and alpha-merges that mask back over the words.

The mask comes from a video-matting network that carries recurrent state from one frame into the next. That detail is the whole reason the effect holds together. A per-frame segmentation model makes an independent decision every frame, and independent decisions strobe at the boundary — the letters flicker, a chair toggles between masked and not, half-transparent limbs let the words bite into the person. A model that has seen the previous frames answers consistently with them, so the silhouette stays put whether the subject is walking or standing still.

Two more things follow from "the mask says WHERE, not WHAT". It is measured on a small proxy, which is fast, and it still composites correctly into a full-resolution export, because the subject's pixels in the final frame are the export's own — already graded, at source quality. And the mask is cached against a fingerprint of the exact footage and geometry it was built from, so re-wording the title over the same moment, or undoing and re-adding it, costs nothing to measure again.

If the matting model is unreachable, Valmera does not silently produce something worse. It falls back to a photometric method — the per-pixel median of the window is the background, and each frame's distance from it is the subject — and the reply says which one it used, because that fallback needs a locked-off camera and can miss a dark subject on a dark wall.

Big Type Is the Look — and Why Shrinking It Makes Things Worse

The single most common way this effect goes wrong is a title that is too small. The reason is geometric. A person crossing the frame occupies a band roughly from their head to their feet, and a small line of text sits entirely inside their torso — so whole words vanish as they pass, and what the viewer reads is broken text, not depth.

Large type is what fixes it. With tall glyphs the subject crosses the middle of the letters while their tops and bottoms stay visible, and the eye completes the words on its own. That is the mechanism behind every version of this shot you have seen work. A behind-subject title in Valmera therefore defaults to a much larger size than an ordinary one, and if you ask the agent to fix swallowed letters it will enlarge the type or shorten the phrase — never shrink it.

Placement matters just as much. The agent reports two numbers after every placement: how much of the title's width the subject interrupts at their most, and how much of the title's area they cover on average. Width is the honest one, because interrupted width maps to interrupted letters. If the subject crosses only a sliver, you will see an ordinary title and the reply tells you so — move the text, or move the window to when they actually cross. If they interrupt nearly half the line and then keep standing there, that is a different problem, and the answer is bigger type or fewer words. It reads well when the subject sweeps past and the words re-emerge behind them.

Everything else is ordinary text styling: the seven text templates, the bundled font families — Anton, Bebas Neue, Archivo Black, Playfair Display, Montserrat and the rest — any colors, position anywhere in the frame, and animated entrances and exits (typewriter is entrance-only). See add text to video for the full styling surface.

When It Goes Wrong, and What to Do

Valmera refuses this effect rather than shipping a bad version of it, and it refuses with the number that caused the refusal so you know what to change. Here is the honest list.

Nobody is in the shot
If under 0.4% of the frame reads as subject, there is nothing for the words to go behind. Point it at a moment where someone is actually in front of the camera — or take a plain title.
The subject fills the frame
Past 55% average coverage the words would barely ever be visible. A close-up is the wrong shot for this; use a wider one where the person crosses the frame.
There is a cut in the window
The mask belongs to one continuous piece of one take. Shorten the duration or move the words fully inside a single shot.
A speed ramp covers it
The mask is frame-for-frame with the source, so sped or slowed footage would drift the silhouette off the person. Remove the ramp there, or place the words elsewhere.
The window is over 15 seconds
Measurement is per frame and past 15s it will not finish inside one edit turn. Split it into shorter placements.
It lands on a spliced-in clip
B-roll and generated clips are not your indexed footage, so there is no subject measured there. Aim at a moment where your own video is playing.

Beyond the refusals there are three failure modes that are about the picture rather than the rules. Fast motion is the hardest: a limb travelling far between two frames is motion-blurred into the background, and no matte has a crisp edge to find — the arm briefly gets thin or the letters clip through it. Slower movement across the words looks better than a sprint. Hair is the classic matting problem: fine strands are semi-transparent at the pixel level, so a flyaway edge can look slightly cut out rather than soft. Tied-back hair against a plain background is the easy case; loose curls against a busy one is the hard case. Low contrast between subject and background — dark clothes on a dark wall — costs the model confidence at the edges, which shows up as the words bleeding faintly into the silhouette.

The general remedy is the same for all three: put the words where the subject crosses cleanly, keep them big, and keep the window short. And render a preview before you believe any of it — after the render the agent looks at the frames it produced and checks its own claim, and every reply is verified server-side against the edits actually recorded, so it cannot tell you the text is behind you when it is not.

People Occlude the Words. Objects Do Not.

This is worth stating plainly because it is the first thing people test. The matte is a person matte. Words go behind people and behind whatever they are carrying. Static objects — furniture, a wall, a parked car — do not hide the letters; over those the words render as an ordinary title.

That is a deliberate trade, not an omission. An earlier version tried to decide per-frame whether a chair was foreground, and the chair flipped fully on and off twenty times in one window while the walker lost pixels wherever he lingered near it. Defining the foreground as "a person" rather than "whatever seems to be in front" is what makes the occlusion stop flickering. If you ask for text behind a chair, the agent will tell you this instead of quietly producing a title that sits on top of it.

The Text Is Anchored to the Footage, Not to the Timeline

A behind-subject title owns pixels of a particular second of your video, which means it cannot be pinned to a position in the edit. If you cut something earlier, the words travel with their own shot and stay on the same subject. If the shot itself is cut out of the edit, the text goes with it — and the agent names it in the reply rather than leaving a blank title behind.

The same logic is why you should not put a zoom or a speed change over the window: the mask is frame-for-frame with the source, so re-timing or re-framing that footage after the fact moves the picture out from under the silhouette. A speed ramp is refused outright; a zoom is something the agent will warn you about.

How People Do This Without Valmera

All of these work, and some of them will be the right answer for you. It is worth knowing what each one actually asks of you.

After Effects (Roto Brush) is the reference implementation. You paint a stroke on the subject, After Effects propagates the selection across frames, and you correct it wherever it drifts. It gives you frame-level control that nothing else here matches, including the ability to matte an object rather than a person. The cost is that propagation is not free: on difficult footage you are refining strokes shot by shot, and the whole workflow assumes you already own and know After Effects.

DaVinci Resolve (Magic Mask) is the closest desktop equivalent. You draw inside the subject, Resolve tracks the mask forward and backward, and you connect the mask's alpha output so the person sits on a layer above your text. One thing to check before you download it: Magic Mask runs on Resolve's Neural Engine and is a Studio feature, so the free build leaves you drawing Power Windows and tracking them by hand. Studio is a one-off licence rather than a subscription, which makes it good value if you do this often. It is still a node-and-timeline application: three layers, a tracker to babysit, and a colour page mental model to acquire first.

CapCut made the mobile version of this popular. You duplicate the clip, run auto background removal on the top copy, and sandwich a text layer between the two. It is fast and it is on your phone. The trade-offs are that you are managing three tracks by hand, the cutout quality depends on a single contrasting subject, and you re-do the whole sandwich for every placement — and CapCut's own help pages list background removal among the AI features that sit behind its Pro tier rather than the free editor.

ffmpeg cannot do this alone, and it is worth saying so clearly, because plenty of tutorials imply otherwise. ffmpeg will happily composite a mask over a video — that part is a one-line filter graph. What it has no way to produce is the mask itself. You would run a segmentation model over your frames in Python, write out an alpha sequence, and only then hand ffmpeg the compositing. That pipeline is exactly what Valmera runs; the work is in the middle step.

Browser one-shot tools generally do the still-image version of this — one photo, one subject, one word — which is a different and much easier problem, because there is no temporal coherence to hold and no motion blur to matte through.

When the Effect Works, and When It Looks Cheap

It is a trend, and trends have a cheap version. Being specific about the difference is more useful than either hyping it or dismissing it.

It works when the words are a real piece of information — a name, a chapter title, a place, one word that the shot is about — placed where the subject sweeps through them once and the letters re-emerge. It works when the type is large enough to survive being crossed, when the shot has room around the subject, and when it happens once or twice in a video rather than on every cut. The effect reads as production value precisely because it looks like the title was physically in the room.

It looks cheap when the words are decoration — when the same title could have been anywhere and the depth adds nothing. It looks cheap when the type is small enough to be eaten, when the subject stands still in front of the words for the whole window so the title is simply unreadable, and when it is applied to every title in a video, at which point the viewer stops noticing depth and starts noticing a template. It also looks cheap on the wrong shot: a close-up where the person fills the frame has nowhere for the words to live, which is exactly why Valmera refuses above 55% coverage rather than rendering it.

Honest Limits

  • People only. Furniture, walls and vehicles do not occlude the letters.
  • 15 seconds maximum per placement, and a minimum window of 0.4 seconds.
  • The window must sit inside one continuous take, on your own footage, with no speed ramp over it.
  • Refused above 55% average subject coverage, and below 0.4%.
  • Fast motion, fine hair and low subject-to-background contrast all degrade the edge. Preview before you commit.
  • Do not zoom or re-time that footage afterwards — the mask is tied to the source frames.
  • Requires a main video: it is not available on a project built only from images and clips, where a plain title is the right tool.

Valmera is an agentic editor, not a generator — it edits footage you actually shot, and it will not invent a subject that is not in the frame.

Where It Fits in the Rest of the Edit

A behind-subject title is usually the punctuation on a shot that has already been cut. The same chat can remove the silences, add word-accurate captions, grade the picture, score it with music that ducks under speech, and reframe it for Shorts or Reels — in the same conversation, on the same upload, with every cut restorable.

Browse everything in the AI video editing tools hub, read the text & overlays docs, or drive the whole editor from Claude over the Valmera MCP server.

Put Your Name Behind You

One sentence, a measured matte, and a preview you can check. Free to try.

Start free →
See pricing →

Frequently Asked Questions

Upload the footage to Valmera and type where you want it, in plain English: "put my name behind me as I walk across at 0:12". The agent finds that moment, cuts the person out of every frame in the window with a video-matting model, draws the words onto the picture, and lays the cut-out person back over the letters. The person passes in front of the text. There is no rotoscoping, no green screen, and no mask to keyframe.
Yes, when the person-matting model is in play — it finds the subject in each frame on its own, so a moving camera, a pan, or handheld wobble are all fine. Only the fallback path (used if the model host is unreachable, and the reply says so when that happens) needs a locked-off camera, because it works by photographing the still background out of the shot itself.
No, and this is a deliberate trade. The matte is a person matte: it hides the words behind people and whatever they are carrying. Static objects — furniture, walls, parked cars — do not occlude the letters, so over those the words read as an ordinary title. That restriction is what keeps the occlusion rock-steady frame to frame instead of flickering on and off around a chair.
It refuses with the measurement that caused it, so you know what to change. The common ones: nobody is visible in that window (under 0.4% of the frame reads as subject); the subject fills more than 55% of the frame, so words behind them would barely ever be visible; there is a cut inside the window; a speed ramp covers that footage; the window is longer than 15 seconds; or the moment lands inside a spliced-in clip instead of your own footage. In every case nothing is written to the edit and a plain title is offered instead.
Up to 15 seconds per placement. The subject is measured frame by frame, and past 15 seconds that measurement does not finish inside a single edit turn. Longer sequences are built as several placements over separate windows — each one still has to sit inside a single take with no cut in it.
Because big type is what makes the effect read as depth. A title placed behind a subject defaults to a large size on purpose: with tall glyphs the person crosses the middle of the letters while their tops and bottoms stay visible, and your eye completes the words. A small line sits entirely inside the torso band, so whole words disappear as the person passes and it reads as broken text, not depth. If letters are getting lost, the fix is bigger type or a shorter phrase — never smaller type.
Yes. The text is anchored to the footage it was measured on, not to a position in the timeline, so if you cut something earlier the words travel with their own shot and stay on the same subject. If you cut that shot out of the edit entirely, the text is removed with it and the agent tells you so by name in its reply.
No. A green screen gives you a key; this gives you a matte from ordinary footage. The model was trained to separate people from arbitrary backgrounds, so a room, a street, or a wall all work. What matters far more than the background is contrast at the subject's edges and how fast they are moving.
No. The mask is measured on a small proxy for speed, but what gets stored is a grayscale silhouette, not a cut-out of you. The final render draws the words onto its own full-resolution picture and alpha-merges that silhouette, so the subject's pixels are the export's own — already graded, at source quality. Scaling a mask up softens an edge by about a pixel; scaling a cut-out subject up would put a blurry patch in a sharp frame.

Text Behind the Subject, from One Sentence

The agentic AI video editor — describe the edit, review the preview, download the export. 50 free credits, no card.

Start free →
See pricing →

Related Articles

Add Text to Video
Titles, lower thirds, callouts and quotes — placed and animated from a sentence.
Add Zoom Effects
Punch-ins and travelling zooms aimed at any point in the frame.
Blur Part of a Video
Blur, pixelate or black out a region that follows the footage.
Docs: Text & Overlays
How text templates, overlays and picture-in-picture work.
All AI Video Editing Tools
Every tool in the Valmera editor, organized by job.