← Home
TOOL

Published · Updated

Remove Burned-In Subtitles, Watermarks or Text from a Video

Text burned into the picture can be taken out, not just hidden. Valmera measures where it is and repaints those pixels, reconstructing what was behind them.

To remove burned-in subtitles or a watermark, upload the video to Valmera and type "remove the burned-in subtitles" or "erase the watermark in the top-right corner". The agent scans the real frames to measure the rectangle the text occupies — it does not estimate one — then repaints it: each covered pixel is replaced with the background that was visible there in other frames, anything no frame ever showed clean is reconstructed from the surrounding picture, and the grain is matched so the patch does not read as plastic. It then re-measures the ink inside that rectangle on the repainted file, so it can only tell you the text is gone when the pixels say it is.

Erase Burned-In Text Free

Upload the footage, say what to remove, watch it disappear from the picture. No credit card required.

Try it free →
See pricing →

How to Remove Burned-In Text from a Video

  1. 1
    Upload the footage
    Drag in any MP4, MOV, MKV or WebM up to 14 GB or 3 hours. Valmera runs a one-time analysis — word-level transcript, silence and shot detection, and labeled frame tiles the agent reads directly, so it can already see the text on screen before you describe it.
  2. 2
    Ask where the text is (optional)
    Type "where is there text burned into this video?". You get back measured rectangles as fractions of the frame, each labeled captions, watermark or text, with the seconds it is visible and how many sampled frames it appeared in. This step is worth taking when there is more than one mark and you only want some of them gone.
  3. 3
    Type the removal
    "remove the burned-in subtitles" clears every caption band in one pass. For one specific mark: "erase the watermark in the top-right corner", or "remove the username at the bottom left between 0:12 and 0:40". You can name a time window; regions apply to the whole video otherwise.
  4. 4
    Check the measurement, then export
    The reply carries an ink reading for each rectangle before and after the repaint, and the agent looks at the frames it produced. If a region still measures dirty it widens it or repaints the whole rectangle instead of just the letter strokes, rather than reporting success. When the preview looks right, export an H.264 MP4.

Detection samples frames by seeking, so it is quick. The repaint itself decodes and re-encodes every frame of the full-resolution original — the one edit in Valmera that is bounded by video length rather than by your patience.

EXAMPLE PROMPTS
  • "where is there text burned into this video?"
  • "remove the burned-in subtitles"
  • "erase the watermark in the top-right corner"
  • "remove the username at the bottom left between 0:12 and 0:40"
  • "erase the old captions, then caption it again in Anton, karaoke style"
  • "put the watermark region back — the patch looks worse than the logo"
Blur, cover and repaint compared on the same frameThree copies of one video frame. In the first, a blur leaves a smeared but visible band where the text was. In the second, a solid box hides the text and the picture behind it. In the third, the text is gone and the picture behind it continues uninterrupted.BLURtext still there,now a smeared bandCOVERtext hidden, and so isthe picture behind itREPAINTtext gone, the picturecontinues through itall three are available — only the third is a removal
Three different things get called "removing text". Blurring and covering are censoring: they are the right tool for a face or a licence plate, and the wrong tool for a subtitle band you want gone.

Blur, Cover, Repaint — Three Different Jobs

Most tools that advertise watermark removal do one of the first two. A blur keeps every pixel of the text in the frame and low-passes it, which is why the result is a conspicuous rectangle of mush that a viewer reads as "something was removed here". A solid box is more honest about itself but destroys the picture behind the mark as well. Both leave the video obviously altered in the exact spot you wanted nobody to look at.

Inpainting is the third option: work out what the picture behind the text was, and put that back. For thin, high-contrast ink over footage that keeps moving — which describes subtitles almost perfectly — this is not a hard problem, because the information is genuinely there. The words change every couple of seconds, so a pixel covered by the stroke of an E at 3.0s is clean background at 3.4s. The reconstruction is a photograph, not a hallucination.

Valmera ships all three. Blur, pixelate and black-box regions exist for censoring; the erase exists for removal. Ask for the outcome and the agent picks — say "hide his name badge" and you get a censor, say "get rid of the subtitles" and you get a repaint.

Finding the Text: Measured, Not Guessed

The first half of the job is knowing exactly where the text is, and getting it wrong is worse than doing nothing — a rectangle in the wrong place censors the picture and leaves the text. So Valmera measures it off the frames rather than asking a vision model to eyeball coordinates.

The scan samples 28 frames spread across the video (never the very first or last, which are the most likely to be a fade or a title card). In each one it runs a morphological stroke response — a top-hat and a black-hat, maximised together, so light text on dark footage and dark text on light footage both register, and white text with a black outline registers twice. Glyphs are merged into lines, and lines are filtered on geometry: a text line is between 1.2% and 28% of the frame height, at least 3.5% of its width, and wider than it is tall. That geometry is what stops foliage, brickwork and a crowd from voting as text. Pixels that fall inside a text line in enough of the samples become a region, padded outward because outlines, drop shadows and antialiasing live just outside the ink and a rectangle that clips them leaves a visible ghost.

Each region is then classified without asking a model anything: the crop is compared between samples. Content that changes is captions — wide, centred, new words every line. Content identical in every frame is a watermark — a handle or a logo. Everything else is text: a headline, a UI label, a lower-third. That is what lets "remove the subtitles but keep the channel logo" work as a sentence.

If the structural scan finds nothing and you insist there is something, the agent looks at four frames itself and names what it sees — and then the box it names is snapped to the measured ink inside it before anything is repainted. The model is allowed to point; only the pixels are allowed to measure. And if that finds nothing either, the agent says so and asks where the mark is instead of inventing a rectangle.

How the background plate is built from other framesThree sampled frames of the same shot, each with a different line of subtitle text in the same band. Because the words differ, every pixel of the band is uncovered in at least one sample. Those uncovered samples are combined into a per-pixel median — a real photograph of the background — which is pasted back where the letters were. Any pixel no sample ever showed clean is reconstructed inward from the surrounding picture instead.SAMPLED FRAMESt = 2.0st = 4.0st = 6.0sper-pixel median of every sample in which that pixel was NOT inkTHE PLATEreal background, no inkEVERY FRAME, REPAINTEDfeathered, grain matched
The words move, so the background does not stay hidden. Pixels that no sample ever caught clean — a mark that never moves over a subject that never moves — fall back to reconstruction from the surrounding picture, which is the case where a soft patch can survive.

How the Repaint Works

The removal runs in two passes over the full-resolution original. The first builds the plate: sample frames across the region's time span, and for each pixel take the median of every sample in which that pixel was not covered by ink. The plate is only trusted where enough clean samples agree and the shot is steady enough for them to line up, measured as the mean absolute deviation between samples — a handheld pan gives a plate that would not register correctly, so it is not used there.

The second pass streams every frame of the video. Where this frame has ink, the plate is pasted in. Pixels the plate never saw fall back to boundary-inward reconstruction, which is excellent on thin strokes and imperfect on busy detail. Then two finishing touches that decide whether the result is invisible or merely text-free: the high-frequency noise of the untouched pixels in the same band is measured and matched into the fill, because a perfectly smooth patch reads as plastic against real sensor noise; and the mask is feathered, so no rectangle edge appears in the output.

Two details handle the awkward styles. The backing bar: news lower-thirds and one popular mobile-editor caption preset sit on a solid strip. Repainting only the letters would lift the text off the bar and leave a grey rectangle floating in the shot — which measures as "no ink" and would otherwise be reported as success. So the region is tested against its own surroundings for colour distance and flatness, and when it is a bar the rectangle grows to the bar's real extent, measured across the whole span and unioned, because a bar measured on the shortest caption leaves both ends of the longest one standing. Strokes versus whole rectangle: the default repaints only the letter strokes and keeps the picture between them, which is right for captions and handles. For a solid object, a sticker or a logo shape, the whole rectangle is repainted instead — a harder problem, because an object sits in every frame and no plate can be built for it. That case has its own page: removing an object or a person from a video.

The repainted file keeps the source's duration, frame rate and audio exactly, which is what makes it a drop-in replacement: your transcript, shot list, cuts, captions and every timestamp in the edit still point at the same moment. Erasing text at any point in a session does not invalidate the editing you already did.

The Part Nobody Else Does: Measuring the Result

After the repaint, Valmera measures the ink response inside each rectangle again — on the file that will actually be rendered — and compares it to the reading taken before. The agent sees something like ink 41 -> 2 — gone, or ink 41 -> 27 — STILL VISIBLE. On the second reading it is told to widen the rectangle or repaint the whole box instead of the strokes, and explicitly told not to claim the text was removed until it measures clean.

This matters more here than anywhere else in the editor. Inpainting is the one operation where a confident-sounding "done — the watermark is gone" is easy to write and expensive to trust, because you find out it was wrong at export. Valmera's replies are verified server-side against the edits actually recorded, so the agent cannot claim work it did not do; for the erase specifically, the verification goes down to the pixels. It also renders a preview and looks at the frames it produced, so if the repaint left a soft patch it is expected to tell you rather than let you discover it.

When It Goes Wrong, and What to Do

A soft patch where the text was
The background moved or was highly detailed, so reconstruction had to invent it. Try limiting the erase to the seconds the text is actually on screen, or accept a blur there instead — a smaller lie than a smear.
Text still faintly visible
Almost always an outline or drop shadow sitting outside the rectangle. Say "the ghost of the letters is still there — widen it". The measurement will have flagged it before you did.
A grey bar left behind
The caption sat on a backing strip and only the letters came off. Ask to repaint the whole rectangle rather than the strokes; on a busy shot expect the patch to be visible and consider covering it.
Nothing was found
Neither the structural scan nor a look at the frames turned anything up. Tell the agent where it is — corner, edge, and roughly which second — and it erases that rectangle directly instead of guessing.

The general rule: if the repaint looks worse than the mark, put it back. Ask to restore the region and the source is re-derived from your untouched original without it, and then reach for a blur or a black box, or crop the band out of frame entirely — reframing a 16:9 video to 9:16 often loses a corner watermark for free.

How People Do This Without Valmera

After Effects Content-Aware Fill. The professional answer, and a good one: mask the object, track it, and After Effects generates replacement pixels using surrounding frames — the same underlying idea as the plate described above, with far more control over the mask and a reference-frame option for hard cases. The costs are a Creative Cloud subscription, a slow and RAM-hungry generate step, and a manual mask per shot.

DaVinci Resolve Studio. Patch Replacer samples a clean area of the frame and blends it over the mark; Fusion's Object Removal does neural reconstruction. Both sit on the paid Studio side rather than the free build, so check which edition you have before planning around them. Studio is a one-off licence rather than a subscription, and if you already own it this is a strong option for a single stubborn shot.

ffmpeg delogo / removelogo. Free, scriptable, one command, no install beyond ffmpeg itself. Be clear about what it does: delogo interpolates the box from the pixels immediately outside it — a controlled smear, not a reconstruction — and it has no idea where the text is, so you supply x, y, w and h yourself. removelogo takes a bitmap mask image marking the logo and fills those pixels in from their neighbours instead of a plain rectangle. On a small static station logo over a soft background this is genuinely fine and costs nothing; over detail, or over a subtitle band, it is visibly wrong.

Open-source subtitle removers. Projects built on video inpainting models can produce very good results on the hardsub case. The price is a Python environment, model weights, a GPU worth using, and drawing the box yourself. If you enjoy that, it is the best free result available.

Online watermark removers. A large category, and mostly blur or smart-looking smudge behind an AI label, with upload caps and — often — a watermark of their own on the output. Check the result at 100% before you trust the marketing.

Valmera's claim against that field is narrow and specific: it is the option where you never draw a box, never open a masking UI, and get told in numbers whether the text actually came out. It is not the option with the most control over a single hero shot — that is still After Effects.

The Line Worth Drawing: Your Text vs Someone Else's

Two requests that look identical to a tool are not the same act.

Removing burned-in captions from your own footage, taking a client's old lower-third out at their request, clearing a caption band so you can re-caption in a different font, stripping a handle from a video you own after rebranding — routine post-production, and the reason this feature exists at all.

Removing someone else's watermark so their footage can be reused is a different thing. The mark identifies the owner and exists precisely to prevent that reuse. Stock previews from libraries like Getty and Shutterstock are licensed as previews; stripping the mark breaks the licence, whatever the footage is then used for. And in the United States, removing copyright management information is prohibited under DMCA section 1202 as a separate wrong from infringement itself — so "my use was fair" is not automatically an answer to "you removed the notice".

Valmera does not inspect your footage and has no way to tell whose it is. The tool will run; the responsibility for what you run it on is yours. None of this is legal advice, and the rules differ by country — if a project turns on the question, ask a lawyer rather than a video editor.

Honest Limits

Length. The repaint decodes and re-encodes every frame at full resolution, so it is capped by total pixel work: about 10 minutes of source, and roughly 900 megapixel-seconds — near enough 7 minutes at 1080p and under 2 minutes at 4K. Above that the agent refuses honestly and offers what fits rather than starting something that cannot finish. Every other edit still works on the full 14 GB / 3-hour upload.

No motion tracking. An erase region is a fixed rectangle, optionally limited to a time window. A watermark that drifts around the frame needs a rectangle covering its whole route, which repaints much more picture than the mark occupies. Corner marks — the overwhelming majority — are the good case.

Large translucent overlays. A big diagonal watermark spread across the whole picture is not thin ink on a background; there is no clean background to recover, and repainting the frame is repainting the video. Inpainting does not save that footage, and no honest tool will tell you otherwise.

The export path changes slightly. Normally Valmera exports from your original file at source quality. Once anything is erased, the render reads a repainted full-resolution copy of that original instead — a high-quality H.264 re-encode. Your upload itself is never modified, every erase is re-derived from it rather than from a previous repaint, and undoing one restores the untouched pixels.

No subtitle extraction. Burned-in subtitles cannot be turned into an SRT here — Valmera has no SRT or VTT import or export at all. It can transcribe the speech itself and burn a fresh caption track once the old one is gone.

And our own mark. On a page about removing watermarks it would be odd not to say it: free-plan exports carry a small Valmera mark in the top-left corner, and every export closes with a brief end card after your video. Creator, Pro and Frontier exports carry no watermark at all, and upgrading re-renders an already-marked export clean. Previews are never marked on any plan.

What People Use It For

Re-captioning in a house style
Footage arrives with captions burned in from another editor. Erase them, then caption it again in your own font and preset — the frame is clear, so nothing stacks.
Repurposing across languages
A cut delivered with hardcoded English subtitles cannot be subtitled into another language on top. Removing the band makes the master reusable.
Rebrands and old handles
An archive of your own videos carrying a handle you no longer use. A watermark that never moves and never changes is the easiest case there is.
Screen recordings with names on them
A username, an email, a customer's company in a UI label. Erase reconstructs the interface behind it — often cleaner than the blur that would otherwise sit there.

Part of the Same Chat

Erasing text is rarely the whole job. The same conversation can cut the video, add word-level captions, score it with music that ducks under speech, reframe it for another platform, and grade it — one request can carry several: "erase the old subtitles, caption it in Bebas Neue, and make me a 9:16 version". Browse everything in the tools hub, or read how censoring and effects work in the effects docs.

The same erase tools are exposed over Valmera's MCP server, so Claude can measure and repaint burned-in text from inside a conversation without you opening the studio at all.

Take the Text Out of the Picture

Measured, repainted, and checked against the pixels before anyone claims it worked.

Start free →
See pricing →

Frequently Asked Questions

Upload the video to Valmera and type "remove the burned-in subtitles". The agent scans the actual frames to measure exactly where the subtitle band sits, then repaints those pixels: the letter strokes are replaced with the background that was visible behind them in other frames, and anything no frame ever showed clean is reconstructed from the surrounding picture. It then measures the ink inside the rectangle again on the repainted file and reports the before/after numbers, so a removal is only claimed when the pixels agree.
Really removing it. A blur leaves the text in the picture — unreadable, but plainly still there as a smeared patch. A black or white box covers the text and the picture behind it. Valmera's erase repaints: it builds a per-pixel median of the frames in which each pixel was not covered by ink, which is a genuine photograph of the background, pastes that where the letters were, fills anything left over by reconstructing inward from the surrounding pixels, matches the film grain of the untouched area and feathers the edge. Blur and box are still available if you want them — they are the right answer for censoring a face or a licence plate.
Technically the tool will do it; whether you may is a separate question, and the two cases are not the same. Removing captions or a logo you burned into your own footage (or a client's, at their instruction) is ordinary post-production. Removing someone else's watermark so their footage can be reused is not — the mark is there to identify the owner, stock previews are licensed as previews, and in the United States removing copyright management information is prohibited separately from infringement itself under DMCA section 1202. Valmera does not inspect whose footage you uploaded and cannot tell; the responsibility is yours. This is not legal advice.
On a static or slow shot with thin text, almost always — the background is a real photograph taken from other frames, not a guess. It degrades in three specific cases: a busy background that moves under the text, a mark that never moves over a subject that never moves (no frame ever showed what is behind it), and a caption sitting on a solid backing bar, where the whole bar has to be repainted rather than the letters. In those cases the result can carry a soft patch. The agent renders a preview and looks at the frames it produced, so you are told about it rather than finding it after export.
Yes, and that is the most common reason people erase them. Ask for "erase the burned-in captions, then caption it again in Anton, karaoke style". The erase clears the frame first so new captions cannot stack on top of the old ones. Valmera's own captions come from a word-level transcript, with 11 presets, 12 bundled fonts, 9 entrance animations and per-word emphasis.
The repaint decodes and re-encodes every frame of the full-resolution original, so it is capped by total pixel work: roughly 10 minutes of source, and about 900 megapixel-seconds — near enough 7 minutes at 1080p, under 2 minutes at 4K. Past that the agent refuses honestly and offers what does fit, such as covering the area with a blur or cropping it out of frame with a reframe. Every other edit — cuts, captions, music, zooms — still works on uploads up to 14 GB or 3 hours.
Only by covering the whole path it travels. An erase region is a fixed rectangle on the frame, optionally limited to a time window; there is no motion tracking, so a mark that drifts around the picture needs a rectangle big enough to contain its route, which means repainting far more of the image than the mark occupies. A watermark that stays in one corner — which is most of them — is the case this handles well.
Your upload is never modified. The repaint produces a separate full-resolution copy that the renderer reads instead, and every erase re-derives that copy from the untouched original rather than repainting a repaint, which is what keeps it undoable and stops reconstruction compounding. Ask to put a region back and it is re-derived without it. Derivations are cached by fingerprint, so undoing and redoing the same erase costs nothing the second time.
No. Valmera has no SRT or VTT export — captions here are burned into the picture. What it can do is transcribe the speech with word-level timing and burn a fresh, styled caption track after the old one is erased. If you specifically need a sidecar subtitle file, use an OCR-based subtitle extractor instead.

Remove Burned-In Text from Your Video

The agentic AI video editor — describe the edit, review the preview, download the export. 50 free credits, no card.

Start free →
See pricing →

Related Articles

Blur or Pixelate Part of a Video
When you want the area hidden rather than reconstructed — faces, plates, usernames.
Add Captions to Video
Burn fresh word-level captions after the old ones are erased.
Docs: Effects & Censoring
Grades, stylize, looks, and how region censoring works.
All AI Video Editing Tools
Every editing job Valmera can do, one page each.