Guides

OCR Explained: How a Scan Becomes Searchable Text

A scanned document looks like a document and behaves like a photograph. You cannot search it, you cannot copy a sentence out of it, a screen reader cannot read it aloud, and no converter can turn it into anything useful. The page is a grid of pixels that happens to resemble writing.

Optical character recognition is the process of looking at those pixels and working out which letters they represent. This guide covers what actually happens during that process, what makes it succeed or fail, and what a "searchable PDF" contains once it is done.

The quick test

Open your PDF and try to select a line of text with the cursor. If a neat blue highlight follows the words, the file already has real text and you do not need OCR. If you get a rectangular selection over the whole page, or nothing at all, it is an image.

A document can be both, incidentally: a born-digital report with a scanned appendix stapled on the end is common, and only the appendix needs processing.

What happens between the picture and the text

Recognition is not one step but a pipeline, and each stage can undo the one before it.

Cleaning up the image

First the page is converted to pure black and white. This sounds trivial and is not: the threshold between "ink" and "paper" has to adapt across the page, because scans are rarely evenly lit. A shadow along the spine of a book, or the grey cast of a photocopy, can turn whole paragraphs into solid black or make thin strokes vanish.

Straightening it

The engine measures the dominant angle of the text lines and rotates the page back to horizontal. A page skewed by more than a couple of degrees confuses line detection, which is why a photo of a document taken at an angle performs so much worse than a flatbed scan of the same page.

Finding the structure

Next the page is divided into regions — blocks of text, images, ruling lines — and each block into lines, and each line into words and characters. Layout analysis is where multi-column documents are won or lost. Get it wrong and the engine reads straight across the gutter, producing sentences that alternate between two unrelated columns.

Recognising the characters

Modern engines process a whole text line at a time with a neural network trained on enormous quantities of rendered and scanned text, rather than matching each letter against a template. The line-at-a-time approach matters because context resolves ambiguity: the same shape is an l in one word and a 1 in another, and only the surroundings say which.

Applying language knowledge

Finally, a dictionary and a character-sequence model nudge uncertain results toward plausible words. This is why telling the engine which language it is reading improves accuracy so noticeably — and why a page in a language you did not select comes back as confident nonsense.

What actually drives accuracy

In rough order of impact:

  • Resolution. The number that matters is roughly 300 DPI for ordinary body text. Below about 200, letter strokes are too few pixels wide to be distinguished and accuracy falls off a cliff. Above 400 you gain almost nothing and processing takes longer. Note this is DPI relative to the page, not megapixels — a 12-megapixel phone photo of an A4 sheet is around 300 DPI only if the page fills the frame.
  • Contrast and evenness. Crisp black on white beats grey on cream, every time. Faxes, third-generation photocopies and highlighter pen are the classic problem cases.
  • Straightness. Flat and square. Curved text near a book's spine is substantially harder than the same text lying flat.
  • The typeface. Ordinary serif and sans-serif body text is what these engines are trained on. Decorative fonts, handwriting, condensed all-caps headings and dot-matrix output are all considerably worse.
  • Language selection. Free accuracy if you get it right, a real penalty if you do not.

The errors you should expect

OCR mistakes are not random; they cluster around shapes that genuinely look alike. Worth knowing, because it tells you what to proofread:

  • 0 and O, 1 and l and I, 5 and S, 8 and B
  • rn read as m, cl read as d
  • Commas and full stops, especially at low resolution
  • Hyphens at line ends, which may or may not be rejoined

Notice that most of these involve digits. Prose can absorb the occasional wrong letter because you read the word anyway; a reference number, an amount or a date cannot. If the document is mainly numbers, check them by hand.

What a searchable PDF actually contains

The output of OCR PDF is not a rebuilt document. It is your original scan with a second, invisible layer added — the recognised words, each positioned over the place on the image where it was found, drawn in an invisible text rendering mode.

The consequences of that design are worth spelling out:

  • The page looks exactly as it did before. Nothing about the image changes.
  • Search works, and matches highlight in the right place, because the invisible words sit on top of the visible ones.
  • Copy and paste produce the recognised text — including any errors.
  • Screen readers can read the document.
  • The file gets a little larger, since you have added text to an unchanged image.
  • Recognition errors are invisible. The picture still shows the correct word; only the hidden layer is wrong. This is the one genuine trap — always spot-check by searching for a word you know is on the page.

Getting a better scan in the first place

Ten seconds of care at the scanner is worth more than any setting afterwards: scan at 300 DPI, in greyscale rather than colour for plain documents, with the page flat and square, and with enough light that the paper is white rather than grey. If you are photographing rather than scanning, fill the frame with the page, keep the camera parallel to it, and avoid your own shadow.

Why running it locally is worth something

The documents people scan are passports, bank statements, medical letters and contracts. Recognition needs the whole page, so a server-based tool needs the whole page too. Running the engine in your browser means the image and the recognised text stay on your device. The language model has to be downloaded once, which is why the first run takes longer than the second — after that, it works offline.

In short

OCR reads pixels and writes text, and everything about its accuracy is decided before it starts: resolution, contrast, straightness, language. The result is your original scan with an invisible text layer added, which makes it searchable and copyable without changing how it looks. Check the numbers, because those are the errors you will not see.

Questions about OCR

Around 300 DPI relative to the page is the practical target for ordinary body text. Below roughly 200 DPI letter strokes become too few pixels wide to tell apart and accuracy drops sharply, while going above 400 adds processing time for very little gain.

No. The recognised words are added as an invisible layer positioned over the existing image, so the page renders exactly as before. The file simply becomes searchable, copyable and readable by screen readers.

Recognition errors stay hidden because the visible page is still the original picture — only the invisible layer is wrong. Low resolution, poor contrast, skew, and unusual typefaces are the usual causes, along with the wrong language being selected.

Generally not reliably. These engines are trained mainly on printed type, and handwriting varies far too much in shape and spacing. Neat block capitals sometimes work; ordinary cursive usually does not.

The recognition engine and the language data have to be downloaded to your browser before any processing can start. Once they are cached, later runs begin immediately and work without a connection.

Try it yourself

Free, no account, and your files never leave your device.

OCR PDF

← All posts