Every PDF carries a small dossier about itself. Not the visible content — a separate block of information describing who made the file, with what, and when. Most people never look at it, which is precisely why it keeps causing problems.
It is not sinister. It is useful plumbing that makes search and document management work. It is just that it was designed for files living inside one organisation, and it travels perfectly happily outside one.
What is actually in there
Two places store it. The older document information dictionary holds a handful of fields, and the newer XMP block holds a richer, extensible record. Most files have both, and they do not always agree — which is its own small trap, because clearing one and not the other leaves the information in place.
The fields you will meet:
- Title — often not the filename. Frequently the original document's title from the program that made it, which can be a draft name nobody intended to publish.
- Author — usually filled in automatically from the account name on the computer that created the file. This is the field that most often surprises people.
- Subject and Keywords — free text, sometimes containing internal classifications or project codenames.
- Creator — the application the content was authored in, such as a word processor.
- Producer — the library or driver that wrote the PDF, typically including a precise version number.
- CreationDate and ModDate — timestamps, usually with a timezone offset attached.
Why any of this matters
Taken one at a time these are harmless. Taken together, and attached to a document you are sending outside your organisation, they can say more than you intended.
- Anonymity fails. A document submitted anonymously — a complaint, a tender, a peer review, a tip to a journalist — with your account name in the Author field is not anonymous. This is the single most common real-world problem.
- Timelines become checkable. The creation timestamp and its timezone say when and roughly where a document was produced. If a file was supposedly written last month, the metadata is the first place anyone will look.
- Titles leak drafts. A polished proposal whose Title field reads
Bid v7 - final - lowball optionis a bad afternoon. - Software versions are a small security signal. Producer strings tell anyone interested exactly which version of which library your organisation runs.
- Templates carry their origins. A document built from a file someone else made often keeps the original author. This is how a proposal arrives at a client naming a competitor.
Looking at your own files
Every reader exposes this somewhere, usually under document properties. Open Edit PDF Metadata and drop a file in to see the fields laid out, then change or clear whichever you choose and export a clean copy.
It is worth doing this once on a file you sent recently. Most people find at least one field they did not know was populated.
What metadata does not cover
Clearing these fields is one step of several. A file can also be identified by things that live elsewhere:
- The filename. Not metadata at all, and it travels with the document.
salary-review-draft-jsmith.pdfdefeats any amount of field clearing. - Annotations and comments. Sticky notes carry their own author names and timestamps, independently of the document fields. Flattening merges them into the page.
- Form field values. Filled-in data sits in the form structure, not the page content, until flattened.
- Bookmarks. An outline can spell out section names removed from the visible text.
- Hidden or covered content. Text under a black box is still text — see our guide to what redaction really means.
- Embedded attachments. A PDF can carry whole files inside it, including the spreadsheet a chart was built from.
A pre-send checklist
- Clear Author, Title, Subject and Keywords unless you want them read.
- Flatten the document if it has annotations or form fields.
- Check the bookmark outline.
- Rename the file to something neutral.
- For genuinely sensitive documents, redact first, then clear metadata — redaction tools sometimes write their own producer string.
When to keep metadata rather than strip it
Not every document should be anonymised. For published reports, whitepapers and anything you want found, well-filled metadata helps: search engines and library systems read Title and Author, and a document with a proper title is easier for a recipient to file and find again. The goal is to control the fields, not to empty them reflexively.
Editing without handing over the document
There is a mild irony in uploading a confidential file to a stranger's server in order to remove the identifying information from it. Editing metadata in the browser avoids that: the file is parsed in your own tab, the fields are rewritten in memory, and the clean copy is written back out as a download without anything being transmitted.
In short
Your PDFs carry your account name, your software versions and your timestamps by default. Look once at a file you have already sent, decide what each field should say rather than letting your tools decide, and remember that the filename, the annotations and the bookmarks need the same attention.