What PDF is
PDF describes a page: glyphs at coordinates, not rows and columns. A table in a PDF is a visual arrangement, not a data structure.
Reports, invoices, statements and anything sent to be read rather than processed.
What Markdown is
A Markdown table is rows of pipe-separated cells with a dashed rule under the header. Alignment is set by placing colons in that rule.
README files, GitHub issues, documentation sites, Obsidian and Notion.
What changes when you convert PDF to Markdown
Every character in a text-based PDF carries its position on the page. Characters sharing a baseline are one row. Columns are found from the vertical strips no character ever occupies, which is why extraction works cleanly on a document laid out with spacing and badly on one laid out with ruled lines. A header repeated at the top of each page is detected and dropped rather than landing in the middle of the data.
Cells are padded so the source lines up in a plain text editor, pipes inside values are escaped, and line breaks inside a cell become <br> because a Markdown table row cannot span lines.
What carries over from PDF to Markdown
PDF records 2 things about a table that Markdown has no way to hold.
A line break inside a cell ends the row in Markdown, so line breaks are replaced rather than carried through.
Markdown marks its header cells differently from its data cells, so the first row is written with that marker.
Anything visual in the PDF - weight, alignment, colour, column widths - has no counterpart in Markdown and is dropped. The values are what survives.
This is the direction that recovers structure: PDF has no table in it to read, only an arrangement that looks like one, so the rows and columns are inferred rather than read.