Remove an Object or a Person from a Video
Valmera repaints the thing out of the picture and reconstructs the background behind it — gone, not covered. It is also honest about the half of the time that does not work.
To remove an object from a video, upload it to Valmera and type "remove the bin in the bottom-right corner". The agent looks at real frames of your footage, reads the object's rectangle off the tenths grid printed on them, and repaints that rectangle out of every frame: the hole is rebuilt from the ring of pixels around it, the local grain is measured and matched so the patch does not read as plastic, and the edge is feathered so no rectangle outline survives. It renders a preview, looks at the frames it produced, and tells you if the patch is visible.
The result depends almost entirely on two things: how much of the frame the rectangle has to cover, and how complicated the background behind it is. A small, stationary object over a plain or softly textured background usually disappears completely. A person walking across a detailed room usually does not. The rest of this page explains exactly why that is, and what to do in the second case — because there are four good answers, and pretending inpainting is one of them helps nobody.
Try It on Your Own Shot
Upload the footage, say what to take out, judge the preview yourself. 50 free credits, no card.
Try it free →How to Remove an Object from a Video
- 1Upload the footageDrag in any MP4, MOV, MKV or WebM up to 14 GB or 3 hours. Valmera runs a one-time analysis — word-level transcript, silence and shot detection, and labeled frame tiles the agent reads directly — so it has already seen what is in your shot before you describe it.
- 2Say what to remove, and roughly where"remove the bin in the bottom-right corner" is enough. Name a time window if the object is only there for part of the video: "erase the person in the doorway between 0:12 and 0:19". Several things at once go in one sentence — "remove the mic stand and the light stand" — and they are repainted in a single pass.
- 3Look at the preview, not at the promiseThe reply carries a measured before/after reading for each rectangle, and the agent looks at the frames it actually produced. Neither is a substitute for your own eyes on the patch at full size — this is the one edit in the product where the honest answer is sometimes "that did not work".
- 4Keep it, or put it back"put the bin back" re-derives the source from your untouched original without that region, and costs nothing to redo later. When the shot is right, export a full-quality H.264 MP4.
The repaint decodes and re-encodes every frame of the full-resolution original, so it is the one edit bounded by video length rather than by your patience: roughly 10 minutes of source, and about 900 megapixel-seconds.
- "remove the bin in the bottom-right corner"
- "erase the person in the doorway between 0:12 and 0:19"
- "take the sticker off the laptop lid for the whole video"
- "remove the mic stand and the light stand — both are in shot"
- "the patch looks worse than the object — put it back and blur it instead"
- "crop the frame so the guy in the background is out of shot"
Repainting Is Not Covering
Three different things get sold as "removing" something from a video. A blur keeps every pixel of the object and low-passes it — unreadable, but plainly still there, and a viewer reads the smeared rectangle as "something was taken out here". A box hides the object and the picture behind it. Only repainting is removal: the pixels are replaced and the background is reconstructed, so nothing marks the spot.
Valmera ships all three and the agent picks from what you ask for. Say "hide his badge" and you get a censor; say "get rid of the bin" and you get a repaint. The censoring tools are not a consolation prize — for a face, a document or a licence plate, a visible blur is the correct answer and a perfect reconstruction would be the wrong one.
How the Object Removal Actually Works
An erase is a rectangle in frame fractions — top-left corner, width, height, all between 0 and 1 — with an optional time window in source seconds. The agent does not estimate those numbers. It looks at real frames, each delivered with a faint tenths grid anchored at the top-left, and reads the rectangle off the grid; a coordinate here is a measurement, not an impression. For burned-in text there is a separate structural scan that measures the rectangle from the pixels directly, which is why subtitle and watermark removal is a different page.
The rectangle is then repainted in one of two ways. Stroke fill repaints only the high-contrast ink inside the rectangle and keeps the picture between the letters — right for captions and handles. Box fill repaints the entire rectangle, and that is what an object gets: a bin, a sticker, a stand, a sign, a person.
Box fill is where the honest engineering starts. For text, the repaint builds a plate — sample frames across the region, and for each pixel take the median of every sample in which that pixel was not covered by ink. Because the words change, that median is a genuine photograph of the background. For an object, the same procedure returns nothing: the object is inside the rectangle in every sample, so there is no clean value to take a median of. And no, restricting the erase to the seconds the object is on screen does not help — those are exactly the seconds it is there.
So the object's hole is reconstructed instead, frame by frame, with the Telea algorithm: colour and structure are propagated inward from the boundary of the hole toward its centre. That is why the rectangle is processed inside a larger band — at least twelve pixels, and normally a tenth of the rectangle's width to either side and just over a third of its height above and below. That margin is the entire raw material. It is not a nicety: an early version cropped exactly to the rectangle, and with no ring to reconstruct from, a solid object came back completely untouched.
Two finishing passes decide whether the result is invisible or merely object-free. The high-frequency noise of the untouched pixels in the same band is measured and matched into the fill, because reconstruction returns perfectly smooth pixels and a smooth patch reads as plastic against real sensor noise. Then the mask is feathered, so no rectangle edge appears in the output. Several rectangles — up to eight — go in one call and one pass, which matters more than it sounds: every erase re-derives the whole source, so five separate requests repaint the video five times over, each pass redoing all the earlier rectangles again.
What comes out is a full-resolution drop-in replacement for your source with identical duration, frame rate and audio. That is deliberate: your transcript, shot list, cuts, captions and every timestamp in the edit were measured against the original and all keep pointing at the same moment. Erasing something in the middle of a session never invalidates the editing you already did.
The Measurement, and What It Can and Cannot Tell You
After the repaint, Valmera measures the contrast structure inside each rectangle again — on the file that will actually be rendered — and compares it against the reading taken before. The agent sees something like ink 38 → 2 — gone, and it is explicitly told not to claim a removal until that measures clean. Valmera's replies are verified server-side against the edits actually recorded, so the agent cannot claim work it did not do; for the erase, the verification goes down to the pixels.
Here is the honest caveat, because it belongs on this page rather than the text one. That number measures whether high-contrast structure is still present in the rectangle. For text that is exactly the right question. For an object it is a weaker proxy: a soft, smeared patch scores as clean whether or not it looks acceptable, because a smear has no structure in it. So the numeric check catches the failure where the object survives, and cannot catch the failure where the object is gone but the replacement is obviously wrong. That second one is caught by the agent looking at the frames it rendered — and by you, at full size, before you export.
When It Goes Wrong, and What to Do Instead
The rule that saves the most time: if the repaint looks worse than the object, put it back. A restore re-derives the source from your untouched original, and four other tools in the same chat solve this problem without reconstructing anything.
1. Cut around it. If the object is only in shot for a few seconds, cut those seconds. There is nothing to reconstruct in frames that are not in the video, cuts snap to word boundaries so speech survives, and a clean cut is less noticeable than a good patch, let alone a bad one.
2. Crop it out of frame. Something at the edge of a 16:9 shot is often simply outside a 9:16 or 1:1 crop. Reframing takes a focus point, so the crop window can be aimed at your subject and away from the thing you want gone — and you were probably making a vertical version anyway. Details are in the reframing docs.
3. Push in past it. A zoom across the seconds the object is visible magnifies the frame around a point you choose, which pushes edge content off screen entirely. On a talking-head shot this reads as emphasis rather than as a fix, which is the best kind of workaround.
4. Cover it. A blur, mosaic or black box is honest and instant, and the region follows the footage through later cuts and reframes. Or drop a full-frame b-roll cutaway over the moment: the picture switches to another clip or a still while your audio keeps playing, so the timing never moves. Overlays render above the footage and below captions, and they do not track objects — for a fixed cover, that does not matter.
How People Do This Without Valmera
After Effects Content-Aware Fill. The professional answer and, for a single important shot, a better one than this. You mask the object, track the mask through the shot, and After Effects generates a fill layer using pixels drawn from other frames of the same shot — temporal synthesis, which is precisely the step Valmera does not perform on a solid object. The costs are a Creative Cloud subscription, manual roto per shot, and a generate pass that is slow and memory-hungry. It has failure cases of its own, and they rhyme with ours: shifting light and complex texture.
DaVinci Resolve Studio. Three relevant tools, all part of the paid Studio build rather than the free one: Object Removal for large moving objects and Magic Mask for isolating and tracking a subject — both Neural Engine features Blackmagic lists as Studio-only — plus the Patch Replacer effect for small static blemishes. Studio is a one-off licence rather than a subscription, and if you already own it, it is a strong mainstream route to taking a person out of a shot.
ProPainter and friends, open source. ProPainter (ICCV 2023) propagates information across frames with flow-guided propagation and fills what remains with transformer attention, which handles large masks, camera motion and long occlusions far better than any classical method. Some projects pair a video inpainter like it with a segmentation model so you can point at an object rather than mask it by hand. The price is a Python environment, model weights, a CUDA GPU, and frame-wise masks — the repository asks for the video plus a mask per frame. If that is your idea of a good evening, it is the strongest free result available.
ffmpeg delogo. Free, scriptable, one command. Be clear about what it does: it interpolates the box from the pixels immediately outside it — a controlled smear, not a reconstruction — and it has no idea where anything is, so you supply x, y, w and h yourself. For a small static logo over a soft background it is genuinely fine and costs nothing.
The online object removers. A crowded category of browser tools, mostly brush-or-box selection plus a hosted inpainting model, typically rationed by credits and capped on upload size. Some are very good on short clips. Judge them the way you should judge this one: export the result and look at the patch at 100%, not at the marketing frame.
Valmera's claim against that field is narrow and specific. You never draw a mask or open a masking interface; the removal happens in the same conversation that cuts, captions and reframes the video; the pixels are measured and reported back rather than declared clean; and it is reversible in one sentence. It is not the best video inpainting available, and this page is not going to pretend otherwise. If one shot decides your project, use After Effects or Resolve Studio.
Honest Limits
Rectangles, not shapes. An erase region is an axis-aligned rectangle. There is no segmentation mask, no brush, no roto — everything inside the rectangle is repainted, so an object with a lot of empty frame around it costs you that empty frame too.
No motion tracking. The rectangle is fixed for the whole time window. A moving subject needs a rectangle covering its entire path.
No generative synthesis. Everything that appears in the hole comes from your own footage — a background plate photographed from other frames, or reconstruction from the surrounding pixels. Nothing is invented by a diffusion model. That is why nothing ever hallucinates a new object into your shot, and equally why a complicated background cannot be rebuilt convincingly.
Length and pixels. The pass decodes and re-encodes every frame at full resolution: roughly 10 minutes of source and about 900 megapixel-seconds — near enough 7 minutes at 1080p, under 2 minutes at 4K. Above that the agent refuses and offers what fits. Passing a short time window does not raise the cap; the whole file is still re-encoded, and the window only narrows which frames get touched. Every other edit still works on the full 14 GB / 3-hour upload.
The export path shifts slightly. Normally Valmera exports from your original file at source quality. Once anything is erased, the render reads a repainted full-resolution copy of that original instead — a high-quality H.264 re-encode. Your upload itself is never modified, and every erase is re-derived from it rather than from a previous repaint.
Picture only. Erasing someone from the frame does not remove them from the soundtrack. That is a cut or a mute, and a separate sentence in the same chat.
What People Use It For
Part of the Same Chat
Taking something out of a shot is rarely the whole job. The same conversation can cut the video, add word-level captions, reframe it for another platform, and censor whatever should stay hidden — and one request can carry several: "remove the bin in the corner, cut the dead air, and give me a 9:16 version". Browse everything in the tools hub, or read how censoring and effects work in the effects docs.
The same erase tools are exposed over Valmera's MCP server, so Claude can place and repaint a region from inside a conversation without you opening the studio at all.
Find Out Which Case Your Shot Is
One sentence, one preview, and a straight answer about whether the patch works.
Start free →Frequently Asked Questions
Remove an Object from Your Video
The agentic AI video editor — describe the edit, review the preview, download the export. 50 free credits, no card.
Start free →