PDF to Text
PDF to Text extracts the text layer of a PDF as plain .txt. Choose one combined file or one file per page, with optional page markers. If the PDF is a scan with no text layer, nothing comes out - run OCR PDF on it first.
1 Add your file
Enter the open password to use it here. It is checked in your browser and never sent anywhere.
2 Choose your options
How to PDF to Text
Add your PDF
The text is extracted immediately and shown in the box so you can check it.
Choose how you want it
One combined file or one per page, with flowing paragraphs or the original line breaks preserved.
Copy or download
Copy straight to the clipboard, or save as .txt.
The conversion that cannot get layout wrong
Text extraction pulls the strings out of the page's content streams and assembles them into lines. It deliberately discards formatting, which sounds like a limitation and is often the point: with no layout to reconstruct, there is no layout to reconstruct incorrectly.
For feeding text into another system, searching a bundle of documents, or pasting content into your own template, this is frequently the fastest and most reliable route — more so than a formatted conversion you then have to repair.
A worked example
Forty supplier contracts need checking for a particular clause. Extracting each to text and searching the results takes minutes and is exact. Opening forty PDFs and reading them, or converting them all to Word and hoping the formatting holds, takes far longer and is less reliable.
Limitations worth knowing
- Scanned pages contain no text at all. Run OCR PDF first.
- Reading order follows how the glyphs were drawn, so multi-column layouts can interleave.
- Tables lose their grid entirely — for data, PDF to Excel is the better destination.
- Word boundaries are inferred from spacing, so unusual typesetting occasionally joins or splits words.
- Hyphenation at line ends may or may not be rejoined, depending on how the original was set.
Further reading
Related tools
PDF to Text — frequently asked questions
Because the PDF has no text layer — it is a picture of a page, which is what every scan is. Run it through OCR PDF first to recognise the text, then come back here.
A PDF stores each fragment of text at a position; it has no idea what a paragraph is. Flowing joins fragments that sit on the same line and merges lines into paragraphs, which reads better. Keep line breaks preserves the visual line structure, which matters for poetry, code and tables.
It follows the order the text is stored in the file, which for a well-made single-column document is the reading order. Multi-column layouts often interleave, because the stored order need not match what your eye does.
Yes, if you know the password. Enter it when prompted and extraction proceeds normally.
No. Extraction is fast because no rendering is involved — a few hundred pages take a couple of seconds.