ImageTextRemoverA PictureEditor.com tool
Pictures stay on your machine — only the model travels.

Take a subtitle strip off a frame

A burned-in caption sits in the same place in every still, which is why one box does the work of a hundred careful outlines. This page opens with that box already across the lower third — drag it where your bar actually is.

Drop in a picture with words burned into it, and you will get it back without them.A still, a screenshot of a player, or a set of them at once. PNG keeps the most detail if you have the choice.

Nothing to hand?

The controls arrive with the frame, and nothing is downloaded until you press something that names the size.

Several files at once opens the queue instead, where one box set covers every frame of the same size.

Why a caption band is one box, never eleven

A subtitle is a line of type, but the thing you want gone is almost never just the type. Most players and most burn-in tools lay a translucent band behind the words so they stay readable over a bright shot, and that band has tinted every pixel underneath it. Clear the glyphs alone and you are left with a darkened rectangle with clean letters missing from it, which draws the eye harder than the subtitle ever did.

So box the band. The edges of your box will look for a strong horizontal line within six pixels when you let go, and on a caption band those two lines are usually the clearest edges in the lower third of the frame.

If you cannot see where the band ends, hold Space. The frame as it arrived comes back for as long as you hold, which is the quickest way to find an edge that a fill has already softened.

One pass or four, and what each costs

One pass, softer

The whole bar is squeezed into the 512 square in one go. It is quick — a single pass at about two and a half seconds — and on a 1080p frame it costs roughly three and a half times the detail along the length of the bar. On a soft or blurred background that is invisible. Against sharp rows of detail immediately above and below, it is not.

Sectioned, sharper

The bar is split into overlapping 512 windows with 64 pixels of overlap and blended across the joins. Each window sees the frame at close to its own scale, so the strip comes back as sharp as its surroundings. Four sections is about ten seconds, and the panel prints the count before you commit to it.

The cost is the joins. Where two sections disagree about a smooth gradient — a clear evening sky is the worst case — the blend can read as a faint vertical ripple. Over texture it is invisible. Neither of these routes is the right answer for every frame, which is exactly why you are shown both.

Questions about caption bars

Should I box the letters or the whole bar?
The whole bar, every time. A translucent caption band is two problems stacked — the tint of the band and the glyphs on top of it — and clearing only the glyphs leaves a rectangle of altered brightness exactly where the band was. That reads worse than the subtitle did, because a viewer sees a deliberate patch rather than a caption.
Why does the strip come back softer than the rest of the frame?
Because a 1,800 by 90 bar does not fit a 512-pixel square without being squeezed about three and a half times along its length. That is what the sectioned route is for: overlapping windows along the bar, each close to the frame's own scale, blended across the joins. It costs four passes instead of one and gives the sharpness back.
The bar sits over a face in one shot. What happens?
You get plausible nonsense, and you will know it when you see it. A face carries structure no fill can guess from its surroundings, so over a mouth or an eye the trained fill invents something confident and wrong. Where a bar crosses a face, cropping the frame or picking a different still is usually the better answer, and this site will not pretend otherwise.
Can I clear the same bar off a whole set of stills?
Yes. Drop them together and frames of matching dimensions are grouped; you draw the bar once on the first one and the queue runs it across the group. Stills from a different source at a different size are set aside rather than cleared at coordinates that belong to another frame.

Read on, or go wider