What PDF is
PDF describes a page: glyphs at coordinates, not rows and columns. A table in a PDF is a visual arrangement, not a data structure.
Reports, invoices, statements and anything sent to be read rather than processed.
What XML is
XML marks up data with nested named tags and optional attributes. It is verbose, strict about being well-formed, and can carry a schema that says what shape a document must have.
Enterprise integrations, government data feeds, RSS, SOAP services and older document formats.
What changes when you convert PDF to XML
Every character in a text-based PDF carries its position on the page. Characters sharing a baseline are one row. Columns are found from the vertical strips no character ever occupies, which is why extraction works cleanly on a document laid out with spacing and badly on one laid out with ruled lines. A header repeated at the top of each page is detected and dropped rather than landing in the middle of the data.
Each row becomes an element with one child per column. Header names are rewritten where they have to be, because an XML tag cannot start with a digit or contain a space, so a column called 2024 Sales becomes _2024_Sales.
What carries over from PDF to XML
One property of the PDF has no home in XML, and it is worth knowing which before you convert.
Anything visual in the PDF - weight, alignment, colour, column widths - has no counterpart in XML and is dropped. The values are what survives.
This is the direction that recovers structure: PDF has no table in it to read, only an arrangement that looks like one, so the rows and columns are inferred rather than read.