Put Text Behind the Subject in a Video
A title painted on the wall behind you, and you walk in front of the letters. It is one sentence in Valmera, and it is a real depth composite — not a fade, not a transparency.
To put text behind the subject of a video, upload the footage to Valmera and type "put my name behind me as I walk across at 0:12". The AI agent cuts the person out of every frame in that window with a video-matting model, draws the words onto the picture, and lays the cut-out person back over the letters. The person occludes the text the way they would occlude anything genuinely behind them. There is no rotoscoping, no green screen, and no mask to keyframe.
Try the Depth Text Effect Free
50 credits on signup, no credit card. Upload footage, type one sentence, watch the preview.
Put text behind you →How to Put Text Behind a Subject in a Video
- 1Upload footage with a person in itDrag in any MP4, MOV, MKV or WebM up to 14 GB or 3 hours. Valmera runs a one-time analysis: a word-level transcript, shot detection, silence detection, and labeled frame tiles the agent actually looks at — so it already knows which moments have someone moving across the frame.
- 2Say what the words are and where they goType the literal sentence: "put VALMERA behind me as I walk across at 0:12, for about 3 seconds". You can also say it by content instead of by clock — "put my name behind me on the shot where I step away from the desk" — because the agent can see the frames.
- 3Read the measurement in the replyThe agent reports what it actually measured: how much of the frame the subject covers, and how much of the title's width they interrupt at their most. If the subject never crosses the words, it says so rather than telling you an effect is there that you cannot see.
- 4Preview, adjust the type, exportWatch the preview and refine in plain English — "make it bigger", "move it left so I cross the middle of the word", "use Anton instead". When it reads right, export a full-quality H.264 MP4 rendered from your original file.
The subject is measured frame by frame over the window you asked for, so a 3-second placement is quick and a 15-second one is the longest allowed.
- "put my name behind me as I walk across at 0:12"
- "big title behind the subject on the opening shot, 4 seconds"
- "the word LAUNCH behind me in Anton, white, for 3 seconds at 0:45"
- "move that text left so I cross the middle of it"
- "make the behind text bigger — the letters are getting swallowed"
What "Behind" Actually Means Here
A renderer draws text on top of the picture. There is no "behind" until something knows which pixels are the person. That is the entire job, and it is why this effect is rare outside of desktop compositing apps: everything else about it — the font, the timing, the animation — already existed.
How the Matte Is Actually Built
When you ask for text behind the subject, Valmera does not guess. It measures a window of your footage and stores a grayscale mask clip — white where the subject is, black where they are not, one frame for every frame of the window. The renderer then draws the text onto a copy of its own picture and alpha-merges that mask back over the words.
The mask comes from a video-matting network that carries recurrent state from one frame into the next. That detail is the whole reason the effect holds together. A per-frame segmentation model makes an independent decision every frame, and independent decisions strobe at the boundary — the letters flicker, a chair toggles between masked and not, half-transparent limbs let the words bite into the person. A model that has seen the previous frames answers consistently with them, so the silhouette stays put whether the subject is walking or standing still.
Two more things follow from "the mask says WHERE, not WHAT". It is measured on a small proxy, which is fast, and it still composites correctly into a full-resolution export, because the subject's pixels in the final frame are the export's own — already graded, at source quality. And the mask is cached against a fingerprint of the exact footage and geometry it was built from, so re-wording the title over the same moment, or undoing and re-adding it, costs nothing to measure again.
If the matting model is unreachable, Valmera does not silently produce something worse. It falls back to a photometric method — the per-pixel median of the window is the background, and each frame's distance from it is the subject — and the reply says which one it used, because that fallback needs a locked-off camera and can miss a dark subject on a dark wall.
Big Type Is the Look — and Why Shrinking It Makes Things Worse
The single most common way this effect goes wrong is a title that is too small. The reason is geometric. A person crossing the frame occupies a band roughly from their head to their feet, and a small line of text sits entirely inside their torso — so whole words vanish as they pass, and what the viewer reads is broken text, not depth.
Large type is what fixes it. With tall glyphs the subject crosses the middle of the letters while their tops and bottoms stay visible, and the eye completes the words on its own. That is the mechanism behind every version of this shot you have seen work. A behind-subject title in Valmera therefore defaults to a much larger size than an ordinary one, and if you ask the agent to fix swallowed letters it will enlarge the type or shorten the phrase — never shrink it.
Placement matters just as much. The agent reports two numbers after every placement: how much of the title's width the subject interrupts at their most, and how much of the title's area they cover on average. Width is the honest one, because interrupted width maps to interrupted letters. If the subject crosses only a sliver, you will see an ordinary title and the reply tells you so — move the text, or move the window to when they actually cross. If they interrupt nearly half the line and then keep standing there, that is a different problem, and the answer is bigger type or fewer words. It reads well when the subject sweeps past and the words re-emerge behind them.
Everything else is ordinary text styling: the seven text templates, the bundled font families — Anton, Bebas Neue, Archivo Black, Playfair Display, Montserrat and the rest — any colors, position anywhere in the frame, and animated entrances and exits (typewriter is entrance-only). See add text to video for the full styling surface.
When It Goes Wrong, and What to Do
Valmera refuses this effect rather than shipping a bad version of it, and it refuses with the number that caused the refusal so you know what to change. Here is the honest list.
Beyond the refusals there are three failure modes that are about the picture rather than the rules. Fast motion is the hardest: a limb travelling far between two frames is motion-blurred into the background, and no matte has a crisp edge to find — the arm briefly gets thin or the letters clip through it. Slower movement across the words looks better than a sprint. Hair is the classic matting problem: fine strands are semi-transparent at the pixel level, so a flyaway edge can look slightly cut out rather than soft. Tied-back hair against a plain background is the easy case; loose curls against a busy one is the hard case. Low contrast between subject and background — dark clothes on a dark wall — costs the model confidence at the edges, which shows up as the words bleeding faintly into the silhouette.
The general remedy is the same for all three: put the words where the subject crosses cleanly, keep them big, and keep the window short. And render a preview before you believe any of it — after the render the agent looks at the frames it produced and checks its own claim, and every reply is verified server-side against the edits actually recorded, so it cannot tell you the text is behind you when it is not.
People Occlude the Words. Objects Do Not.
This is worth stating plainly because it is the first thing people test. The matte is a person matte. Words go behind people and behind whatever they are carrying. Static objects — furniture, a wall, a parked car — do not hide the letters; over those the words render as an ordinary title.
That is a deliberate trade, not an omission. An earlier version tried to decide per-frame whether a chair was foreground, and the chair flipped fully on and off twenty times in one window while the walker lost pixels wherever he lingered near it. Defining the foreground as "a person" rather than "whatever seems to be in front" is what makes the occlusion stop flickering. If you ask for text behind a chair, the agent will tell you this instead of quietly producing a title that sits on top of it.
The Text Is Anchored to the Footage, Not to the Timeline
A behind-subject title owns pixels of a particular second of your video, which means it cannot be pinned to a position in the edit. If you cut something earlier, the words travel with their own shot and stay on the same subject. If the shot itself is cut out of the edit, the text goes with it — and the agent names it in the reply rather than leaving a blank title behind.
The same logic is why you should not put a zoom or a speed change over the window: the mask is frame-for-frame with the source, so re-timing or re-framing that footage after the fact moves the picture out from under the silhouette. A speed ramp is refused outright; a zoom is something the agent will warn you about.
How People Do This Without Valmera
All of these work, and some of them will be the right answer for you. It is worth knowing what each one actually asks of you.
After Effects (Roto Brush) is the reference implementation. You paint a stroke on the subject, After Effects propagates the selection across frames, and you correct it wherever it drifts. It gives you frame-level control that nothing else here matches, including the ability to matte an object rather than a person. The cost is that propagation is not free: on difficult footage you are refining strokes shot by shot, and the whole workflow assumes you already own and know After Effects.
DaVinci Resolve (Magic Mask) is the closest desktop equivalent. You draw inside the subject, Resolve tracks the mask forward and backward, and you connect the mask's alpha output so the person sits on a layer above your text. One thing to check before you download it: Magic Mask runs on Resolve's Neural Engine and is a Studio feature, so the free build leaves you drawing Power Windows and tracking them by hand. Studio is a one-off licence rather than a subscription, which makes it good value if you do this often. It is still a node-and-timeline application: three layers, a tracker to babysit, and a colour page mental model to acquire first.
CapCut made the mobile version of this popular. You duplicate the clip, run auto background removal on the top copy, and sandwich a text layer between the two. It is fast and it is on your phone. The trade-offs are that you are managing three tracks by hand, the cutout quality depends on a single contrasting subject, and you re-do the whole sandwich for every placement — and CapCut's own help pages list background removal among the AI features that sit behind its Pro tier rather than the free editor.
ffmpeg cannot do this alone, and it is worth saying so clearly, because plenty of tutorials imply otherwise. ffmpeg will happily composite a mask over a video — that part is a one-line filter graph. What it has no way to produce is the mask itself. You would run a segmentation model over your frames in Python, write out an alpha sequence, and only then hand ffmpeg the compositing. That pipeline is exactly what Valmera runs; the work is in the middle step.
Browser one-shot tools generally do the still-image version of this — one photo, one subject, one word — which is a different and much easier problem, because there is no temporal coherence to hold and no motion blur to matte through.
When the Effect Works, and When It Looks Cheap
It is a trend, and trends have a cheap version. Being specific about the difference is more useful than either hyping it or dismissing it.
It works when the words are a real piece of information — a name, a chapter title, a place, one word that the shot is about — placed where the subject sweeps through them once and the letters re-emerge. It works when the type is large enough to survive being crossed, when the shot has room around the subject, and when it happens once or twice in a video rather than on every cut. The effect reads as production value precisely because it looks like the title was physically in the room.
It looks cheap when the words are decoration — when the same title could have been anywhere and the depth adds nothing. It looks cheap when the type is small enough to be eaten, when the subject stands still in front of the words for the whole window so the title is simply unreadable, and when it is applied to every title in a video, at which point the viewer stops noticing depth and starts noticing a template. It also looks cheap on the wrong shot: a close-up where the person fills the frame has nowhere for the words to live, which is exactly why Valmera refuses above 55% coverage rather than rendering it.
Honest Limits
- People only. Furniture, walls and vehicles do not occlude the letters.
- 15 seconds maximum per placement, and a minimum window of 0.4 seconds.
- The window must sit inside one continuous take, on your own footage, with no speed ramp over it.
- Refused above 55% average subject coverage, and below 0.4%.
- Fast motion, fine hair and low subject-to-background contrast all degrade the edge. Preview before you commit.
- Do not zoom or re-time that footage afterwards — the mask is tied to the source frames.
- Requires a main video: it is not available on a project built only from images and clips, where a plain title is the right tool.
Valmera is an agentic editor, not a generator — it edits footage you actually shot, and it will not invent a subject that is not in the frame.
Where It Fits in the Rest of the Edit
A behind-subject title is usually the punctuation on a shot that has already been cut. The same chat can remove the silences, add word-accurate captions, grade the picture, score it with music that ducks under speech, and reframe it for Shorts or Reels — in the same conversation, on the same upload, with every cut restorable.
Browse everything in the AI video editing tools hub, read the text & overlays docs, or drive the whole editor from Claude over the Valmera MCP server.
Put Your Name Behind You
One sentence, a measured matte, and a preview you can check. Free to try.
Start free →Frequently Asked Questions
Text Behind the Subject, from One Sentence
The agentic AI video editor — describe the edit, review the preview, download the export. 50 free credits, no card.
Start free →