← Home
TOOL

Published · Updated

Remove an Object or a Person from a Video

Valmera repaints the thing out of the picture and reconstructs the background behind it — gone, not covered. It is also honest about the half of the time that does not work.

To remove an object from a video, upload it to Valmera and type "remove the bin in the bottom-right corner". The agent looks at real frames of your footage, reads the object's rectangle off the tenths grid printed on them, and repaints that rectangle out of every frame: the hole is rebuilt from the ring of pixels around it, the local grain is measured and matched so the patch does not read as plastic, and the edge is feathered so no rectangle outline survives. It renders a preview, looks at the frames it produced, and tells you if the patch is visible.

The result depends almost entirely on two things: how much of the frame the rectangle has to cover, and how complicated the background behind it is. A small, stationary object over a plain or softly textured background usually disappears completely. A person walking across a detailed room usually does not. The rest of this page explains exactly why that is, and what to do in the second case — because there are four good answers, and pretending inpainting is one of them helps nobody.

Try It on Your Own Shot

Upload the footage, say what to take out, judge the preview yourself. 50 free credits, no card.

Try it free →
See pricing →

How to Remove an Object from a Video

  1. 1
    Upload the footage
    Drag in any MP4, MOV, MKV or WebM up to 14 GB or 3 hours. Valmera runs a one-time analysis — word-level transcript, silence and shot detection, and labeled frame tiles the agent reads directly — so it has already seen what is in your shot before you describe it.
  2. 2
    Say what to remove, and roughly where
    "remove the bin in the bottom-right corner" is enough. Name a time window if the object is only there for part of the video: "erase the person in the doorway between 0:12 and 0:19". Several things at once go in one sentence — "remove the mic stand and the light stand" — and they are repainted in a single pass.
  3. 3
    Look at the preview, not at the promise
    The reply carries a measured before/after reading for each rectangle, and the agent looks at the frames it actually produced. Neither is a substitute for your own eyes on the patch at full size — this is the one edit in the product where the honest answer is sometimes "that did not work".
  4. 4
    Keep it, or put it back
    "put the bin back" re-derives the source from your untouched original without that region, and costs nothing to redo later. When the shot is right, export a full-quality H.264 MP4.

The repaint decodes and re-encodes every frame of the full-resolution original, so it is the one edit bounded by video length rather than by your patience: roughly 10 minutes of source, and about 900 megapixel-seconds.

EXAMPLE PROMPTS
  • "remove the bin in the bottom-right corner"
  • "erase the person in the doorway between 0:12 and 0:19"
  • "take the sticker off the laptop lid for the whole video"
  • "remove the mic stand and the light stand — both are in shot"
  • "the patch looks worse than the object — put it back and blur it instead"
  • "crop the frame so the guy in the background is out of shot"
Why burned-in text inpaints perfectly and an object does notTop row: three sampled frames of a video with burned-in subtitles. The words differ in each frame, so every pixel of the caption band is uncovered in at least one of them, and a real photograph of the background can be assembled from the samples. Bottom row: three sampled frames containing the same object in the same place. The object is present in every frame, so no frame shows what is behind it, and the hole has to be reconstructed inward from the ring of pixels surrounding the rectangle instead.BURNED-INTEXTthe words move, so every covered pixel is photographed in some other frameANOBJECTthe object is in every frame — no frame ever shows what is behind itso the hole is rebuilt inward from the ring of pixels around itsame tool, two completely different problems
Removing burned-in subtitles and removing an object are the same operation with opposite raw material. One recovers a real photograph; the other has to invent one. Every difference in quality between the two jobs comes from this.

Repainting Is Not Covering

Three different things get sold as "removing" something from a video. A blur keeps every pixel of the object and low-passes it — unreadable, but plainly still there, and a viewer reads the smeared rectangle as "something was taken out here". A box hides the object and the picture behind it. Only repainting is removal: the pixels are replaced and the background is reconstructed, so nothing marks the spot.

Valmera ships all three and the agent picks from what you ask for. Say "hide his badge" and you get a censor; say "get rid of the bin" and you get a repaint. The censoring tools are not a consolation prize — for a face, a document or a licence plate, a visible blur is the correct answer and a perfect reconstruction would be the wrong one.

How the Object Removal Actually Works

An erase is a rectangle in frame fractions — top-left corner, width, height, all between 0 and 1 — with an optional time window in source seconds. The agent does not estimate those numbers. It looks at real frames, each delivered with a faint tenths grid anchored at the top-left, and reads the rectangle off the grid; a coordinate here is a measurement, not an impression. For burned-in text there is a separate structural scan that measures the rectangle from the pixels directly, which is why subtitle and watermark removal is a different page.

The rectangle is then repainted in one of two ways. Stroke fill repaints only the high-contrast ink inside the rectangle and keeps the picture between the letters — right for captions and handles. Box fill repaints the entire rectangle, and that is what an object gets: a bin, a sticker, a stand, a sign, a person.

Box fill is where the honest engineering starts. For text, the repaint builds a plate — sample frames across the region, and for each pixel take the median of every sample in which that pixel was not covered by ink. Because the words change, that median is a genuine photograph of the background. For an object, the same procedure returns nothing: the object is inside the rectangle in every sample, so there is no clean value to take a median of. And no, restricting the erase to the seconds the object is on screen does not help — those are exactly the seconds it is there.

So the object's hole is reconstructed instead, frame by frame, with the Telea algorithm: colour and structure are propagated inward from the boundary of the hole toward its centre. That is why the rectangle is processed inside a larger band — at least twelve pixels, and normally a tenth of the rectangle's width to either side and just over a third of its height above and below. That margin is the entire raw material. It is not a nicety: an early version cropped exactly to the rectangle, and with no ring to reconstruct from, a solid object came back completely untouched.

Two finishing passes decide whether the result is invisible or merely object-free. The high-frequency noise of the untouched pixels in the same band is measured and matched into the fill, because reconstruction returns perfectly smooth pixels and a smooth patch reads as plastic against real sensor noise. Then the mask is feathered, so no rectangle edge appears in the output. Several rectangles — up to eight — go in one call and one pass, which matters more than it sounds: every erase re-derives the whole source, so five separate requests repaint the video five times over, each pass redoing all the earlier rectangles again.

What comes out is a full-resolution drop-in replacement for your source with identical duration, frame rate and audio. That is deliberate: your transcript, shot list, cuts, captions and every timestamp in the edit were measured against the original and all keep pointing at the same moment. Erasing something in the middle of a session never invalidates the editing you already did.

The Measurement, and What It Can and Cannot Tell You

After the repaint, Valmera measures the contrast structure inside each rectangle again — on the file that will actually be rendered — and compares it against the reading taken before. The agent sees something like ink 38 → 2 — gone, and it is explicitly told not to claim a removal until that measures clean. Valmera's replies are verified server-side against the edits actually recorded, so the agent cannot claim work it did not do; for the erase, the verification goes down to the pixels.

Here is the honest caveat, because it belongs on this page rather than the text one. That number measures whether high-contrast structure is still present in the rectangle. For text that is exactly the right question. For an object it is a weaker proxy: a soft, smeared patch scores as clean whether or not it looks acceptable, because a smear has no structure in it. So the numeric check catches the failure where the object survives, and cannot catch the failure where the object is gone but the replacement is obviously wrong. That second one is caught by the agent looking at the frames it rendered — and by you, at full size, before you export.

When repainting an object works, plotted against hole size and background detailA two-axis map. The horizontal axis is how complicated the background behind the object is, from plain and uniform on the left to structured and detailed on the right. The vertical axis is how much of the frame the rectangle has to cover, from small and stationary at the bottom to large or moving at the top. The bottom-left quadrant is reliable, the bottom-right leaves a visible seam where straight lines fail to continue, the top-left is often fine over a featureless wall or sky, and the top-right — a person walking through a detailed room — is the quadrant where a different approach is needed.OFTEN FINEa big hole in a flat wallor an empty sky hasnothing to get wrongDO SOMETHING ELSEcut the range, crop theframe, push in, or coverit with a cutawayRELIABLEa sticker, a socket, abin against grass — thepatch is invisibleUSABLE, WITH A SEAMsmall enough to survive,but a brick course or ahorizon will not continuea bin in a locked-off shota person walking pastplain, uniformstructured, detailedTHE BACKGROUND BEHIND ITlargesmallvertical axis: how much of the frame the rectangle has to cover
Both axes fall straight out of the mechanism. Reconstruction propagates inward from the ring, so a bigger hole means more distance from any real information, and a more structured background means the picture it is trying to continue has rules it does not know about. Motion inflates the first axis, because a fixed rectangle has to contain the object's whole route.

When It Goes Wrong, and What to Do Instead

A soft, smeared patch
The hole was too big for the ring around it to explain. Try a tighter rectangle — trim it to the object rather than around it — and a time window covering only the seconds it is on screen. If it still smears, this shot is not an inpainting shot.
Straight lines stop dead
A skirting board, a door frame, a brick course, a horizon. Reconstruction continues colour and gradient inward; it does not know a line is a line and will not carry it across. Cover or crop instead — this one does not improve with a better rectangle.
The object moves
Nothing tracks it, so the rectangle must contain the whole route — repainting far more picture, for far longer, than the object occupies. Cutting the seconds out is nearly always the better edit.
It is a person
People are large, they move, they occlude what is behind them, and viewers look at them harder than at anything else in the frame. Treat person removal as the hard case it is, and expect the answer to be one of the four below.

The rule that saves the most time: if the repaint looks worse than the object, put it back. A restore re-derives the source from your untouched original, and four other tools in the same chat solve this problem without reconstructing anything.

1. Cut around it. If the object is only in shot for a few seconds, cut those seconds. There is nothing to reconstruct in frames that are not in the video, cuts snap to word boundaries so speech survives, and a clean cut is less noticeable than a good patch, let alone a bad one.

2. Crop it out of frame. Something at the edge of a 16:9 shot is often simply outside a 9:16 or 1:1 crop. Reframing takes a focus point, so the crop window can be aimed at your subject and away from the thing you want gone — and you were probably making a vertical version anyway. Details are in the reframing docs.

3. Push in past it. A zoom across the seconds the object is visible magnifies the frame around a point you choose, which pushes edge content off screen entirely. On a talking-head shot this reads as emphasis rather than as a fix, which is the best kind of workaround.

4. Cover it. A blur, mosaic or black box is honest and instant, and the region follows the footage through later cuts and reframes. Or drop a full-frame b-roll cutaway over the moment: the picture switches to another clip or a still while your audio keeps playing, so the timing never moves. Overlays render above the footage and below captions, and they do not track objects — for a fixed cover, that does not matter.

How People Do This Without Valmera

After Effects Content-Aware Fill. The professional answer and, for a single important shot, a better one than this. You mask the object, track the mask through the shot, and After Effects generates a fill layer using pixels drawn from other frames of the same shot — temporal synthesis, which is precisely the step Valmera does not perform on a solid object. The costs are a Creative Cloud subscription, manual roto per shot, and a generate pass that is slow and memory-hungry. It has failure cases of its own, and they rhyme with ours: shifting light and complex texture.

DaVinci Resolve Studio. Three relevant tools, all part of the paid Studio build rather than the free one: Object Removal for large moving objects and Magic Mask for isolating and tracking a subject — both Neural Engine features Blackmagic lists as Studio-only — plus the Patch Replacer effect for small static blemishes. Studio is a one-off licence rather than a subscription, and if you already own it, it is a strong mainstream route to taking a person out of a shot.

ProPainter and friends, open source. ProPainter (ICCV 2023) propagates information across frames with flow-guided propagation and fills what remains with transformer attention, which handles large masks, camera motion and long occlusions far better than any classical method. Some projects pair a video inpainter like it with a segmentation model so you can point at an object rather than mask it by hand. The price is a Python environment, model weights, a CUDA GPU, and frame-wise masks — the repository asks for the video plus a mask per frame. If that is your idea of a good evening, it is the strongest free result available.

ffmpeg delogo. Free, scriptable, one command. Be clear about what it does: it interpolates the box from the pixels immediately outside it — a controlled smear, not a reconstruction — and it has no idea where anything is, so you supply x, y, w and h yourself. For a small static logo over a soft background it is genuinely fine and costs nothing.

The online object removers. A crowded category of browser tools, mostly brush-or-box selection plus a hosted inpainting model, typically rationed by credits and capped on upload size. Some are very good on short clips. Judge them the way you should judge this one: export the result and look at the patch at 100%, not at the marketing frame.

Valmera's claim against that field is narrow and specific. You never draw a mask or open a masking interface; the removal happens in the same conversation that cuts, captions and reframes the video; the pixels are measured and reported back rather than declared clean; and it is reversible in one sentence. It is not the best video inpainting available, and this page is not going to pretend otherwise. If one shot decides your project, use After Effects or Resolve Studio.

Honest Limits

Rectangles, not shapes. An erase region is an axis-aligned rectangle. There is no segmentation mask, no brush, no roto — everything inside the rectangle is repainted, so an object with a lot of empty frame around it costs you that empty frame too.

No motion tracking. The rectangle is fixed for the whole time window. A moving subject needs a rectangle covering its entire path.

No generative synthesis. Everything that appears in the hole comes from your own footage — a background plate photographed from other frames, or reconstruction from the surrounding pixels. Nothing is invented by a diffusion model. That is why nothing ever hallucinates a new object into your shot, and equally why a complicated background cannot be rebuilt convincingly.

Length and pixels. The pass decodes and re-encodes every frame at full resolution: roughly 10 minutes of source and about 900 megapixel-seconds — near enough 7 minutes at 1080p, under 2 minutes at 4K. Above that the agent refuses and offers what fits. Passing a short time window does not raise the cap; the whole file is still re-encoded, and the window only narrows which frames get touched. Every other edit still works on the full 14 GB / 3-hour upload.

The export path shifts slightly. Normally Valmera exports from your original file at source quality. Once anything is erased, the render reads a repainted full-resolution copy of that original instead — a high-quality H.264 re-encode. Your upload itself is never modified, and every erase is re-derived from it rather than from a previous repaint.

Picture only. Erasing someone from the frame does not remove them from the soundtrack. That is a cut or a mute, and a separate sentence in the same chat.

What People Use It For

Clutter in a locked-off shot
A bin, a cable, a light stand, a chair you did not notice until the edit. Static camera, static object, ordinary background — the case this handles best.
Stickers, badges and logos
A brand on a laptop lid or a mug in a product shot. Small, flat, unmoving, on a surface with a simple texture behind it.
Names and numbers in screen recordings
A username, an email, a customer's company in a UI label. Interface backgrounds are flat by design, which makes the reconstruction unusually clean — see removing burned-in text.
A bystander at the edge of frame
Someone standing still in a corner of a static shot, small in frame. If they walk, cut the seconds or crop the frame instead.

Part of the Same Chat

Taking something out of a shot is rarely the whole job. The same conversation can cut the video, add word-level captions, reframe it for another platform, and censor whatever should stay hidden — and one request can carry several: "remove the bin in the corner, cut the dead air, and give me a 9:16 version". Browse everything in the tools hub, or read how censoring and effects work in the effects docs.

The same erase tools are exposed over Valmera's MCP server, so Claude can place and repaint a region from inside a conversation without you opening the studio at all.

Find Out Which Case Your Shot Is

One sentence, one preview, and a straight answer about whether the patch works.

Start free →
See pricing →

Frequently Asked Questions

Yes, and the honest version of the answer matters. Valmera repaints the rectangle you point at out of every frame and rebuilds the picture from the pixels immediately around it, then matches the local grain and feathers the edge so no rectangle outline survives. On a small object over a plain or softly textured background — a bin against grass, a sticker on a laptop lid, a mark on a wall — the result is usually invisible. On a large object over structured detail, or anything that moves through the shot, the reconstruction has to invent picture it has no source for, and it shows. The tool tells you which case you are in instead of promising the first one.
Sometimes. A person standing still at the edge of a locked-off shot, small in frame, against a plain wall or sky is a case this handles. A person walking across a room is not: there is no motion tracking, so the rectangle has to cover the whole route they take, which means repainting a large part of the picture in every frame, and viewers look harder at people than at anything else in a shot. For that case the better answers are in the editor too — cut the seconds they are visible, crop the frame so they fall outside it, push in with a zoom, or cover them with a b-roll cutaway while the audio keeps running.
Yes, in the classical sense of the term. Video inpainting means reconstructing a masked region of the picture from information the video already contains. Valmera does it two ways: for burned-in text it builds a background plate from other frames — the words move, so every covered pixel is genuinely photographed at some other moment — and for an object it reconstructs inward from the ring of pixels around the rectangle, frame by frame, using the Telea algorithm. What it does not do is generative synthesis: no diffusion model invents new content. That is why nothing hallucinates a new object into your frame, and also why a complicated background cannot be rebuilt.
No. You describe the object in plain English and the agent places the rectangle itself — it looks at real frames of your footage, each printed with a faint tenths grid from the top-left corner, and reads the coordinates off that grid rather than estimating them. If it lands in the wrong place you correct it in words: "the rectangle is too small, it left the handle showing" or "that is the wrong chair, I meant the one on the left". There is no masking interface, and equally no brush or roto tool if you wanted one.
Put it back. Ask to restore the region and the source is re-derived from your untouched original without it — derivations are cached, so undoing and redoing the same erase costs nothing the second time. Then reach for the alternatives: cut the range out, crop the frame so the object falls outside it, push in with a zoom over the seconds it is on screen, cover it with a blur or a black box, or drop a full-frame b-roll cutaway over it while your audio keeps playing. A clean cut is almost always less noticeable than a bad patch.
Not by following it. An erase region is a fixed rectangle, optionally limited to a time window in source seconds; nothing tracks the object. A moving subject therefore needs a rectangle large enough to contain the whole path it travels, which repaints far more of the picture than the object occupies and for far longer than it is there. Objects that stay put — which is most of the ones people want gone from a static shot — are the good case.
The repaint decodes and re-encodes every frame of the full-resolution original, so it is capped by total pixel work rather than by file size: roughly 10 minutes of source, and about 900 megapixel-seconds — near enough 7 minutes at 1080p and under 2 minutes at 4K. Above that the agent refuses honestly and offers what does fit. Limiting the erase to a short time window does not raise the cap: the whole file is still re-encoded, the window only narrows which frames get touched. Every other edit — cuts, captions, music, zooms — still works on uploads up to 14 GB or 3 hours.
Never. The repaint writes a separate full-resolution copy that the renderer reads instead, and every erase is re-derived from the untouched original rather than repainted on top of a previous repaint — which is what keeps it undoable and stops reconstruction compounding. The copy keeps the source's exact duration, frame rate and audio, so your transcript, shot list, cuts, captions and every timestamp in the edit still point at the same moment. Erasing something halfway through a session does not invalidate the editing you already did.
No, and it is worth saying out loud because it surprises people. The repaint works on pixels only — someone erased from the frame is still audible on the soundtrack. If you want them gone from the video entirely, cut the range instead, or cut the range and ask for the audio muted across it. Those are separate requests in the same chat.

Remove an Object from Your Video

The agentic AI video editor — describe the edit, review the preview, download the export. 50 free credits, no card.

Start free →
See pricing →

Related Articles

Remove Burned-In Subtitles & Watermarks
The case inpainting is genuinely good at — thin ink whose background other frames already show.
Blur or Pixelate Part of a Video
When covering is the right answer: faces, plates, documents, anything you want visibly hidden.
Resize Video for Any Platform
Crop the frame to 9:16 or 1:1 — often the cheapest way to lose something at the edge.
All AI Video Editing Tools
Every editing job Valmera can do, one page each.