A scanned document looks like a document and behaves like a photograph. You cannot search it, you cannot copy a sentence out of it, a screen reader cannot read it aloud, and no converter can turn it into anything useful. The page is a grid of pixels that happens to resemble writing.
Optical character recognition is the process of looking at those pixels and working out which letters they represent. This guide covers what actually happens during that process, what makes it succeed or fail, and what a "searchable PDF" contains once it is done.
The quick test
Open your PDF and try to select a line of text with the cursor. If a neat blue highlight follows the words, the file already has real text and you do not need OCR. If you get a rectangular selection over the whole page, or nothing at all, it is an image.
A document can be both, incidentally: a born-digital report with a scanned appendix stapled on the end is common, and only the appendix needs processing.
What happens between the picture and the text
Recognition is not one step but a pipeline, and each stage can undo the one before it.
Cleaning up the image
First the page is converted to pure black and white. This sounds trivial and is not: the threshold between "ink" and "paper" has to adapt across the page, because scans are rarely evenly lit. A shadow along the spine of a book, or the grey cast of a photocopy, can turn whole paragraphs into solid black or make thin strokes vanish.
Straightening it
The engine measures the dominant angle of the text lines and rotates the page back to horizontal. A page skewed by more than a couple of degrees confuses line detection, which is why a photo of a document taken at an angle performs so much worse than a flatbed scan of the same page.
Finding the structure
Next the page is divided into regions — blocks of text, images, ruling lines — and each block into lines, and each line into words and characters. Layout analysis is where multi-column documents are won or lost. Get it wrong and the engine reads straight across the gutter, producing sentences that alternate between two unrelated columns.
Recognising the characters
Modern engines process a whole text line at a time with a neural network trained on enormous
quantities of rendered and scanned text, rather than matching each letter against a template. The
line-at-a-time approach matters because context resolves ambiguity: the same shape is an
l in one word and a 1 in another, and only the surroundings say which.
Applying language knowledge
Finally, a dictionary and a character-sequence model nudge uncertain results toward plausible words. This is why telling the engine which language it is reading improves accuracy so noticeably — and why a page in a language you did not select comes back as confident nonsense.
What actually drives accuracy
In rough order of impact:
- Resolution. The number that matters is roughly 300 DPI for ordinary body text. Below about 200, letter strokes are too few pixels wide to be distinguished and accuracy falls off a cliff. Above 400 you gain almost nothing and processing takes longer. Note this is DPI relative to the page, not megapixels — a 12-megapixel phone photo of an A4 sheet is around 300 DPI only if the page fills the frame.
- Contrast and evenness. Crisp black on white beats grey on cream, every time. Faxes, third-generation photocopies and highlighter pen are the classic problem cases.
- Straightness. Flat and square. Curved text near a book's spine is substantially harder than the same text lying flat.
- The typeface. Ordinary serif and sans-serif body text is what these engines are trained on. Decorative fonts, handwriting, condensed all-caps headings and dot-matrix output are all considerably worse.
- Language selection. Free accuracy if you get it right, a real penalty if you do not.
The errors you should expect
OCR mistakes are not random; they cluster around shapes that genuinely look alike. Worth knowing, because it tells you what to proofread:
0andO,1andlandI,5andS,8andBrnread asm,clread asd- Commas and full stops, especially at low resolution
- Hyphens at line ends, which may or may not be rejoined
Notice that most of these involve digits. Prose can absorb the occasional wrong letter because you read the word anyway; a reference number, an amount or a date cannot. If the document is mainly numbers, check them by hand.
What a searchable PDF actually contains
The output of OCR PDF is not a rebuilt document. It is your original scan with a second, invisible layer added — the recognised words, each positioned over the place on the image where it was found, drawn in an invisible text rendering mode.
The consequences of that design are worth spelling out:
- The page looks exactly as it did before. Nothing about the image changes.
- Search works, and matches highlight in the right place, because the invisible words sit on top of the visible ones.
- Copy and paste produce the recognised text — including any errors.
- Screen readers can read the document.
- The file gets a little larger, since you have added text to an unchanged image.
- Recognition errors are invisible. The picture still shows the correct word; only the hidden layer is wrong. This is the one genuine trap — always spot-check by searching for a word you know is on the page.
Getting a better scan in the first place
Ten seconds of care at the scanner is worth more than any setting afterwards: scan at 300 DPI, in greyscale rather than colour for plain documents, with the page flat and square, and with enough light that the paper is white rather than grey. If you are photographing rather than scanning, fill the frame with the page, keep the camera parallel to it, and avoid your own shadow.
Why running it locally is worth something
The documents people scan are passports, bank statements, medical letters and contracts. Recognition needs the whole page, so a server-based tool needs the whole page too. Running the engine in your browser means the image and the recognised text stay on your device. The language model has to be downloaded once, which is why the first run takes longer than the second — after that, it works offline.
In short
OCR reads pixels and writes text, and everything about its accuracy is decided before it starts: resolution, contrast, straightness, language. The result is your original scan with an invisible text layer added, which makes it searchable and copyable without changing how it looks. Check the numbers, because those are the errors you will not see.