PDF to Excel
PDF to Excel looks at where each piece of text sits on the page, clusters it into rows and columns, and writes the result as a spreadsheet. It works best on ruled, evenly spaced tables. Merged cells, nested headers and tables that wrap across pages need manual tidying afterwards - this is a head start, not a finished import.
1 Add your file
Enter the open password to use it here. It is checked in your browser and never sent anywhere.
2 Choose your options
Table detection works by looking at where each piece of text sits on the page and clustering it into rows and columns. On a clean, evenly spaced table it is accurate. On merged cells, nested headers, tables that wrap across pages, or columns separated by only a thin gap, it will get some cells wrong. Treat the output as a head start that saves you retyping, not as a finished import — always check the result against the original before you rely on the numbers.
How to PDF to Excel
Add your PDF
Text positions are read from the file — no rendering, so it is fast.
Tune the sensitivity
The preview updates as you adjust the column and row sliders, so you can see the grid being found.
Download the spreadsheet
Export as .xlsx with one sheet per page, or as a single CSV.
Finding a table that was never marked as one
Most PDF tables are not tables. They are text positioned in aligned columns, with no structure saying where a cell begins. Extraction therefore means detecting the grid: looking for ruling lines where they exist, and otherwise clustering text by its horizontal position to work out where the column boundaries must be.
Ruled tables with clear borders detect reliably. Tables held together only by whitespace depend on consistent alignment — which is why a column of right-aligned numbers next to a column of ragged text can confuse the detection.
A worked example
A bank statement extracts cleanly except that the description column occasionally wraps onto a second line, and those continuation lines arrive as separate rows. The quickest repair is in the spreadsheet rather than the conversion: sort or filter on the date column to bring the orphan lines together, then merge them into the row above.
Limitations worth knowing
- Merged cells, nested headers and cells spanning several columns rarely survive as structure.
- Cells whose text wraps across lines may arrive as separate rows.
- Numbers are extracted as they are printed, so currency symbols, thousands separators and trailing minus signs may need reformatting as numbers.
- Scanned tables need OCR first, and recognition errors in digits are both likely and easy to miss.
- Always reconcile a total against the source. A silently dropped row is the failure mode that costs the most.
Further reading
Related tools
PDF to Excel — frequently asked questions
Adjust the column sensitivity slider. A lower value splits on smaller gaps, which helps when columns are close together; a higher value stops a wide gap inside one cell from being read as a column break. The preview shows the effect immediately.
No, because there is no text to position. Run OCR PDF first to create a text layer, then extract from the result. Accuracy will depend on how well the OCR read the page.
Each page is treated separately, so a table running over three pages gives you three blocks. With one sheet per page you can stack them yourself, which is usually a single copy and paste.
Not reliably. A merged cell appears as content in one column with the others left empty, which is often close enough to fix by hand but is not a faithful reproduction.
Because this one never sees your figures. Everything happens in your browser, which matters when the table is a payroll, a patient list or an unreleased forecast.