Convert PDF to CSV Online

Drop a PDF and get the table out as CSV. The extraction runs in your browser, so the statement or report you would rather not upload to a converter stays on your machine.

From

The only part of this site that uses a server. The address is sent to it so the page can be fetched, because a browser is not allowed to fetch someone else's page itself. Your own files are never involved.

Nothing here is uploaded. The parsing happens on this page.
Edit
Paste or upload something above and it will appear here.
ToDataMarkupDocumentsCodeSchemasImages

The output will appear here once there is a table to convert.

What PDF is

PDF describes a page: glyphs at coordinates, not rows and columns. A table in a PDF is a visual arrangement, not a data structure.

Reports, invoices, statements and anything sent to be read rather than processed.

What CSV is

CSV is a text file where each line is a row and commas separate the fields. There is no type system, no formatting and no second sheet. Every value is text until something downstream decides otherwise.

Almost every database, analytics tool and spreadsheet can read and write it, which is why exports default to it.

What changes when you convert PDF to CSV

Every character in a text-based PDF carries its position on the page. Characters sharing a baseline are one row. Columns are found from the vertical strips no character ever occupies, which is why extraction works cleanly on a document laid out with spacing and badly on one laid out with ruled lines. A header repeated at the top of each page is detected and dropped rather than landing in the middle of the data.

Any value containing the delimiter, a quote or a line break gets wrapped in quotes, and embedded quotes are doubled. That is the rule in RFC 4180 and it is what spreadsheet software expects.

A PDF has no table in it. It has characters at coordinates, and the table is a pattern your eye finds in the arrangement. Extraction means finding the same pattern mechanically: characters sharing a baseline are a row, and the columns come from the vertical strips that no character ever occupies.

That is why it works cleanly on some documents and badly on others, and why swapping extractors rarely helps. They are all doing the same coordinate clustering. What decides the outcome is how the PDF was laid out, which is a choice made by whatever generated it and invisible in the result.

A header repeated at the top of every page is detected and dropped, so a forty-page table does not arrive with thirty-nine copies of the header row scattered through it.

Check the output before you use it. Extraction is inference, not reading, and a cell that landed in the wrong column looks exactly like a cell that landed in the right one.

What carries over from PDF to CSV

One property of the PDF has no home in CSV, and it is worth knowing which before you convert.

Anything visual in the PDF - weight, alignment, colour, column widths - has no counterpart in CSV and is dropped. The values are what survives.

The result can be read a record at a time, and appended to by adding to the end of it. That is worth having for a table too large to hold in memory, and it means a CSV file that is cut off part way through still gives you every record before the cut.

This is the direction that recovers structure: PDF has no table in it to read, only an arrangement that looks like one, so the rows and columns are inferred rather than read.

When it will not work

A scanned PDF has no text at all, only a picture of a page. Nothing here can read it, and the honest answer is that you need optical character recognition, which is different software with a different failure rate. You can tell in two seconds: try to select a line of the table in any PDF viewer. If nothing highlights, it is a scan.

Tables ruled with lines instead of separated by whitespace also fail, because the lines are drawing operations rather than text and a text-based extractor does not see them. So do cells whose contents wrap onto a second line, which arrive as an extra row with most columns empty.

A worked example

Three rows of staff data, with an identifier that has a leading zero, a value containing a comma, a value containing quotes and one blank cell. Those are the four places formats disagree, so they are the four places to look.

PDF in
id   name              role           started     hours
007  Halima Yusuf      Analyst, data  2024-03-15  38.5
012  Jonah "Jo" Pryce  Engineer       2025-11-02
104  Wei Chen          Manager        2023-06-30  40
CSV out
id,name,role,started,hours
007,Halima Yusuf,"Analyst, data",2024-03-15,38.5
012,"Jonah ""Jo"" Pryce",Engineer,2025-11-02,
104,Wei Chen,Manager,2023-06-30,40

Questions

How do I convert PDF to CSV?

Paste your PDF into the box above or drop the file onto it. Check the table in the grid, then copy or download the CSV from the output panel. It takes one step and the data never leaves your browser.

Can I get my table back out of the PDF?

Out of one made here, yes, because the text is real text. Out of a scanned document, no. A scan is an image of a page, and pulling data from it needs optical character recognition, which is a different job with a different failure rate.

Why does my CSV open with all the data in one column?

The file uses a delimiter your spreadsheet did not expect. Exports from European locales often use semicolons because the comma is the decimal separator there. Change the delimiter in the export options and the columns split correctly.

My PDF gives no output at all. Why?

It is almost certainly a scan. Open it in any viewer and try to select a word. If nothing highlights and search finds nothing, the page is an image and there is no text to extract.

The columns came out wrong. Can I fix it?

Yes, in the grid above the output. That is faster than looking for a different tool, because every extractor uses the same approach and fails on the same documents. If the PDF was generated from a spreadsheet or a database, going back to that source beats reconstructing the table.

Is there a limit on file size?

No. The work happens on your own machine, so the limit is your machine's memory rather than an upload cap. A file with tens of thousands of rows converts in a second or two.

Is my data uploaded anywhere?

No. The parsing and generating are done by JavaScript running on this page. You can watch the network tab while you convert and see that nothing is sent.

Related conversions