What PDF is
PDF describes a page: glyphs at coordinates, not rows and columns. A table in a PDF is a visual arrangement, not a data structure.
Reports, invoices, statements and anything sent to be read rather than processed.
What HTML is
An HTML table is a table element containing thead and tbody, with th cells for headers and td cells for data.
Web pages, email templates and anything pasted into a CMS.
What changes when you convert PDF to HTML
Every character in a text-based PDF carries its position on the page. Characters sharing a baseline are one row. Columns are found from the vertical strips no character ever occupies, which is why extraction works cleanly on a document laid out with spacing and badly on one laid out with ruled lines. A header repeated at the top of each page is detected and dropped rather than landing in the middle of the data.
Semantic markup: a real thead, a real tbody, th for the header row. Angle brackets and ampersands in your data are escaped, so a value like <script> is written safely and shows as text on the page.
What carries over from PDF to HTML
HTML can express everything PDF holds about a table, so this conversion is about shape rather than loss.
HTML marks its header cells differently from its data cells, so the first row is written with that marker.
This is the direction that recovers structure: PDF has no table in it to read, only an arrangement that looks like one, so the rows and columns are inferred rather than read.