Every few months a redacted document leaks its own secrets. A court filing, a government report, a settlement agreement — published with the sensitive parts covered in black, and within hours someone has posted the hidden text. The mistake is almost always the same, and it is worth understanding, because the tool that makes it is the one most people reach for first.
This guide covers what a PDF actually stores, why a black rectangle does not remove anything, how to redact properly, and the parts of a file that people forget to check.
A PDF is not a picture of a page
It is tempting to think of a PDF as a flat image, like a photograph of a printed sheet. It is not. A PDF page is a list of drawing instructions — a program, really — that a viewer executes to produce what you see. A typical instruction says something close to: set the font to Helvetica at 11 points, move to coordinates (72, 640), and show the string "Account 4417-2290".
That string is stored as text. It sits in the file as data, independent of what appears on screen. When you select text in a PDF and copy it, you are not running character recognition on pixels — you are reading the strings that were there all along.
Why the black box does not work
When you open a PDF in a general-purpose editor and draw a filled black rectangle over a name, you have added one more instruction to the end of the list: fill a rectangle at these coordinates with black. Because it comes last, it paints over the text on screen.
The text instruction is still in the file. It was never touched. So:
- Select the area and copy — the hidden text lands on your clipboard.
- Run any text extraction tool over the file — the words come out in full.
- Open it in an editor that lets you delete objects — remove the rectangle and read the page.
- Search the document for a word you know is under the box — the viewer finds and highlights it.
None of this requires special skill or software. The information was published; it was merely hidden behind a shape. The same applies to the other improvised methods: changing the text colour to white, drawing a white box, setting a highlight to opaque black, or covering the area with an image. Each one changes what is drawn, not what is stored.
What real redaction does
Genuine redaction is destructive by design. The text has to be removed from the content stream and the page rebuilt without it. There is no "undo" inside the file because there is nothing left to undo — the glyphs, and the coordinates that placed them, are gone.
Practically, a redaction tool has to do three things:
- Remove the marked content. Text-showing operators that fall inside the redaction area are deleted from the page's instruction list, not covered.
- Handle images separately. If the sensitive part is inside a scanned image or a photograph, the pixels themselves must be overwritten. Deleting a drawing instruction is not enough when the secret is in the picture.
- Draw the marker afterwards. The familiar black box is added at the end, as a visual sign that something was removed. It is the receipt, not the mechanism.
Redact PDF works this way: you drag boxes over the parts you want gone, and the text under them is stripped from the page before the file is rebuilt. Run a text search on the result and there is nothing to find, because there is nothing there.
The parts people forget
Even a correct redaction leaves gaps if you only think about the visible page. Before a document goes out, check these too.
Document metadata
Title, author, subject, keywords, and the name of the application that produced the file are
stored separately from the page. A report redacted perfectly can still carry an author field with
the name of the person who wrote it, or a title like
Draft-settlement-Henderson-confidential.docx. Clear these with
Edit PDF Metadata.
The filename itself
Obvious once said, routinely forgotten. The filename travels with the document and is not covered by anything you do inside it.
Form fields and annotations
Comments, sticky notes, and filled-in form values live outside the page content stream. A viewer may not show them by default, but they are in the file. Running Flatten PDF before redacting merges annotations and form values into the page, so the redaction step can actually see and remove them.
Bookmarks and links
A bookmark outline can spell out section names you meant to hide, and a hyperlink's target URL can leak a customer ID or an internal server name.
Inference from what remains
This one is not technical. If you redact a name but leave "the Chief Financial Officer of Acme Ltd resigned on 3 March", you have redacted nothing. Redaction protects the string; it does not protect the fact.
A workable order of operations
- Flatten the document, so annotations and form values become part of the page.
- Redact the sensitive areas, removing rather than covering.
- Strip the metadata.
- Rename the file to something neutral.
- Verify — see below. Do not skip this.
How to verify, in under a minute
Open the finished file and try to break it:
- Select all and copy. Paste into a plain text editor. Read what comes out. If a redacted word appears, the redaction failed.
- Use the viewer's search. Search for a term you know was under a box. Zero results is the answer you want.
- Check the file size. A correctly redacted file is usually slightly smaller than the original, because content was removed. If it grew noticeably, you may have added shapes rather than deleted text.
That copy-and-paste test takes fifteen seconds and would have prevented nearly every un-redaction story you have read about.
Why doing this in the browser matters here
Redaction is, by definition, applied to the documents you least want to hand to a stranger — medical records, legal filings, contracts, anything with an account number in it. Uploading such a file to a website to have the secrets removed involves first giving the secrets away.
Browser-based tools avoid that trade. The file is opened by JavaScript running in your own tab, edited in your device's memory, and written back out as a download. Nothing is transmitted. You can confirm it by opening the network panel in your browser's developer tools and watching for requests while you work, or by disconnecting from the internet after the page loads and redacting anyway — it still works, because there is no server involved.
In short
A black rectangle is a drawing. Redaction is a deletion. The difference does not show on screen, which is exactly why it keeps catching people out — the file looks finished either way. Remove the text, clear the metadata, then copy-and-paste the result to prove it worked.