Extracting Tables From PDFs: What Works and What Does Not
Last updated 19 August 2026 · about 5 minutes to read
A PDF describes a page, not a table. What that means for extraction, why scanned documents need OCR, and how to tell which kind you have.
A PDF does not contain a table. It contains instructions for drawing glyphs at coordinates. The table you see is a pattern your eye finds in that arrangement, and extraction is the work of finding the same pattern mechanically.
The two kinds of PDF
A text-based PDF holds real characters with real positions. Select a word in a viewer and it highlights; that means the text is there and extraction is possible.
A scanned PDF holds a photograph of a page. Nothing highlights when you drag across it because there is nothing there but pixels. Getting data out needs optical character recognition, which is a different kind of software with a different failure mode: it does not fail to find the table, it finds the table and misreads some of the characters.
Test which you have before anything else. Open the PDF, try to select a line of the table, and try to search for a word you can see. If both work, you have text. If neither does, you have an image.
How extraction actually works
Every character in a text PDF has an x and y coordinate. Characters that share a y value are on the same line. Columns are found by looking at the x values across all lines and finding the gaps that persist: if no character ever starts between x=180 and x=240, that gap is a column boundary.
This is why extraction works well on some documents and badly on others. It depends entirely on how the document was laid out, which is a decision made by whoever generated the PDF and is invisible in the result.
What makes it fail
- Cells whose contents wrap onto a second line. The wrapped part has a different y value, so it looks like a new row with most columns empty.
- Merged cells and spanning headers. There is no gap where a boundary is supposed to be.
- Columns separated by ruled lines rather than by whitespace. Lines are drawing operations, not text, so a purely text-based extractor does not see them.
- Right-aligned numeric columns next to left-aligned text ones. The gap moves from row to row, so no consistent boundary exists.
- Multi-column page layouts, where the extractor reads across the columns instead of down them.
What to do when it goes wrong
Fix it in a grid rather than hunting for a better extractor. Every extractor is doing the same coordinate clustering and they fail on the same documents. Twenty minutes fixing cells is usually faster than an afternoon comparing tools.
If the PDF was generated from something else, go back to the source. A PDF is the last step in a pipeline, and whatever produced it almost certainly still has the data in a form that does not need to be reconstructed.
If you have a scan and no source, run OCR and then check every number by eye. OCR confuses 0 with O, 1 with l, 5 with S and 8 with B, and a financial table with one substituted character is worse than no table at all because the error is invisible.
Where Tablizer is at
PDF output works: a table here can be drawn into a paginated PDF with a repeating header. PDF input is not built yet. Rather than ship a page whose converter cannot do the conversion, the format is left out of the input list until the column clustering is done and tested.