What PDF is
PDF describes a page: glyphs at coordinates, not rows and columns. A table in a PDF is a visual arrangement, not a data structure.
Reports, invoices, statements and anything sent to be read rather than processed.
What CSV is
CSV is a text file where each line is a row and commas separate the fields. There is no type system, no formatting and no second sheet. Every value is text until something downstream decides otherwise.
Almost every database, analytics tool and spreadsheet can read and write it, which is why exports default to it.
What changes when you convert PDF to CSV
Every character in a text-based PDF carries its position on the page. Characters sharing a baseline are one row. Columns are found from the vertical strips no character ever occupies, which is why extraction works cleanly on a document laid out with spacing and badly on one laid out with ruled lines. A header repeated at the top of each page is detected and dropped rather than landing in the middle of the data.
Any value containing the delimiter, a quote or a line break gets wrapped in quotes, and embedded quotes are doubled. That is the rule in RFC 4180 and it is what spreadsheet software expects.
A PDF has no table in it. It has characters at coordinates, and the table is a pattern your eye finds in the arrangement. Extraction means finding the same pattern mechanically: characters sharing a baseline are a row, and the columns come from the vertical strips that no character ever occupies.
That is why it works cleanly on some documents and badly on others, and why swapping extractors rarely helps. They are all doing the same coordinate clustering. What decides the outcome is how the PDF was laid out, which is a choice made by whatever generated it and invisible in the result.
A header repeated at the top of every page is detected and dropped, so a forty-page table does not arrive with thirty-nine copies of the header row scattered through it.
Check the output before you use it. Extraction is inference, not reading, and a cell that landed in the wrong column looks exactly like a cell that landed in the right one.
What carries over from PDF to CSV
One property of the PDF has no home in CSV, and it is worth knowing which before you convert.
Anything visual in the PDF - weight, alignment, colour, column widths - has no counterpart in CSV and is dropped. The values are what survives.
The result can be read a record at a time, and appended to by adding to the end of it. That is worth having for a table too large to hold in memory, and it means a CSV file that is cut off part way through still gives you every record before the cut.
This is the direction that recovers structure: PDF has no table in it to read, only an arrangement that looks like one, so the rows and columns are inferred rather than read.
When it will not work
A scanned PDF has no text at all, only a picture of a page. Nothing here can read it, and the honest answer is that you need optical character recognition, which is different software with a different failure rate. You can tell in two seconds: try to select a line of the table in any PDF viewer. If nothing highlights, it is a scan.
Tables ruled with lines instead of separated by whitespace also fail, because the lines are drawing operations rather than text and a text-based extractor does not see them. So do cells whose contents wrap onto a second line, which arrive as an extra row with most columns empty.