OCR PDF

OCR PDF renders each page to an image and runs optical character recognition on it in your browser using Tesseract. You get the recognised text to copy, and optionally a new PDF with an invisible text layer behind the original page so the document becomes searchable. Accuracy depends on scan quality: clean 300 DPI pages read very well, low-resolution or skewed scans much less so.

1 Add your file

Drop your PDF here Scanned pages, photos of documents, anything with text you cannot select. or PDF
    A note on accuracy.

    OCR accuracy depends almost entirely on the scan. A clean, straight 300 DPI page of printed text usually reads at well over 95 per cent accuracy. A phone photo at an angle, a faint fax, a page with heavy background texture, or handwriting will be much worse — handwriting in particular is not something this engine handles. Always read the output before trusting it. Recognition runs in a background worker, so the page stays usable while it works, but a long document genuinely takes a few seconds per page.

    How to OCR PDF

    Add your scanned PDF

    Any PDF whose text you cannot select is a candidate. Photos of pages work too.

    Set the language and resolution

    Pick the language the document is written in — it matters a lot for accuracy. 200 DPI is a sensible default.

    Read the result and download

    The recognised text appears on the page so you can check it before downloading a searchable PDF or a text file.

    How a scan becomes searchable

    Recognition is a pipeline, and each stage can undo the one before it. The page is reduced to black and white with a threshold that adapts across uneven lighting, straightened, divided into regions and lines, read a line at a time by a neural network, and finally nudged toward plausible words by a language model. Telling it the right language is free accuracy; getting it wrong produces confident nonsense.

    The output is not a rebuilt document. It is your original scan with an invisible text layer added, each recognised word positioned over the place on the image where it was found. The page looks exactly as it did; it simply becomes searchable, copyable and readable by screen readers.

    A worked example

    A 30-page contract photographed with a phone comes out poorly: the pages fill about half the frame, so the effective resolution is nearer 150 DPI than 300, and a shadow falls across the gutter. Re-photographing with the page filling the frame, flat and evenly lit, improves the result more than any setting available afterwards. Ten seconds at the scanner beats ten minutes of correction.

    Limitations worth knowing

    • Around 300 DPI relative to the page is the practical target. Below roughly 200 DPI, letter strokes are too few pixels wide and accuracy falls sharply.
    • Recognition errors are invisible, because the visible page is still the original picture — only the hidden layer is wrong. Spot-check by searching for a word you know is there.
    • Digits suffer most: 0 and O, 1 and l and I, 5 and S, 8 and B. Check reference numbers, amounts and dates by hand.
    • Handwriting is generally not readable. These engines are trained on printed type.
    • The engine and language data download on first use, which is why the first run is slow and later ones are not.

    Further reading

    OCR Explained: How a Scan Becomes Searchable Text

    Related tools

    OCR PDF — frequently asked questions

    Yes. Tesseract, the same engine most open-source OCR is built on, is compiled to WebAssembly and runs in a background worker inside this tab. Your scan is never uploaded. The only thing downloaded is the language data file, once.

    Because it is using your device rather than a rack of servers. A page takes a few seconds. The trade-off is that your document never leaves your hands, which for contracts, medical letters and ID scans is usually worth the wait.

    No. This engine is trained on printed text. Handwriting produces nonsense, and we would rather say that plainly than let you waste ten minutes finding out.

    Your original page images are kept exactly as they are, and the recognised words are written behind them as invisible text positioned where each word sits. The page looks identical, but Ctrl+F finds things and you can select and copy.

    Ten of the most common are offered in the dropdown. Tesseract supports over a hundred; if you need one that is not listed, get in touch and we will add it.

    Rescan at 300 DPI if you can, in black and white rather than colour, and make sure the page is straight. If rescanning is not an option, raising the resolution slider sometimes helps, though it cannot invent detail that is not in the image.