ImageTextRemoverA PictureEditor.com tool

Caption bars, and the two ways to clear one

A subtitle burned into a picture is two things stacked: a band that has tinted every pixel beneath it, and the letters sitting on the band. Which of those you box decides almost everything about how the result looks.

The band is the part that gives it away

Burn-in tools put a translucent band behind subtitles for a good reason: white type over a bright shot is unreadable, and a band at forty per cent black fixes that everywhere at once. The band is why your caption is legible, and it is also why clearing the caption is harder than it looks.

Every pixel under that band has been darkened. If you box only the glyphs, the fill rebuilds them from their immediate surroundings — which are darkened band — so the letters vanish into a rectangle of shadow that is still sitting there. You have traded something a viewer reads as a subtitle for something they read as damage, which is a worse outcome even though strictly less of the frame has changed.

Box the band, not the words. It is more pixels, it is one drag instead of several, and it is the only version of this job that leaves nothing behind.

Where the 512 comes in

The trained fill on this site accepts a 512-pixel square and no other size. Your picture never reaches it: a region is cut around your box with context on every side, resampled into that square, run once, and put back at the frame’s own resolution inside the box.

For a date stamp or a small label that is a fine arrangement, because the region is not much bigger than 512 to begin with. For a caption bar it is the whole problem. A bar 1,800 pixels wide on a 1080p frame gives a region several times wider than the square, so along the length of the bar one model pixel has to stand for three and a half of yours. The band comes back convincing in shape and soft in detail, and on a sharp frame you can see the difference between the strip and the rows around it.

Sections, and what they cost

The other route splits the bar into overlapping 512 windows — 64 pixels of overlap, feather-blended where they meet. Each window now sees the frame at close to its own scale, so the strip comes back as sharp as its surroundings.

The cost is the joins. Two neighbouring windows are two separate guesses, and they do not always agree about a gradient that runs through both of them. On a clear evening sky the disagreement can show as a faint vertical ripple where the blend sits. On grass, gravel or foliage nobody will ever find it.

A 1,800 by 90 bar on a 1080p frame
RoutePassesRoughly
One pass, softer12.4 seconds
Sectioned, sharper49.7 seconds

Those numbers come from the model registry’s median of 2,420 milliseconds per pass on single-threaded WebAssembly, measured in Chromium on a desktop. A mid-range phone will be slower, and the panel shows the arithmetic rather than a promise.

When neither route is the answer

A bar over a face comes back as plausible nonsense. So does a bar lying across a keyboard, a bookshelf, a car grille or a second line of smaller type. In every one of those cases something structured crosses the box, and structure is precisely what a fill cannot guess from the outside — it will produce something crisp and wrong.

When that happens you have three honest options: crop the bar out of the frame, choose a different still where the bar sits over something plainer, or accept the visible smear of the diffusion fill, which at least announces itself as a repair rather than passing as a photograph.

Questions people bring to this one

How do I tell whether my bar is translucent?
Look at the same part of the frame just outside the band. If the picture continues at a different brightness the moment the band ends, the band is tinting what is under it, and your box wants the whole band rather than the letters.
Is four sections always better than one pass?
No. Sections are sharper along the bar and can leave a faint join where the ground is a smooth gradient. One pass is softer and never leaves a join. On a bar over clean sky the single pass is often the better-looking answer even though it is technically coarser.
What if there are two lines of subtitle?
One box around both, if they share a band. Two boxes if they sit on separate bands with picture between them. The thing to avoid is a box per word, which leaves the fill guessing at a row of holes with slivers of glyph between them.

Back to the picture