What HTML is
An HTML table is a table element containing thead and tbody, with th cells for headers and td cells for data.
Web pages, email templates and anything pasted into a CMS.
What Pandas DataFrame is
Python that builds a pandas DataFrame from a dictionary of columns.
Data analysis and notebooks.
What changes when you convert HTML to Pandas DataFrame
The first table in the markup is used. Cell contents are stripped of inline tags and entities are decoded, so <td><b>Total</b> </td> arrives as Total. Nested tables inside a cell are not descended into.
The import line is included so the snippet runs as is. Turn type inference on and numeric columns arrive as numbers rather than strings, which is usually what you want in a DataFrame.
What carries over from HTML to Pandas DataFrame
HTML records 2 things about a table that pandas has no way to hold.
pandas wants a type for each column, and HTML does not record one, so each column is typed from what its values look like. A column of digits that should stay text - a zip code, a phone number, a leading-zero id - is the usual thing to check afterwards.
HTML can hold a value that is itself a list or an object. pandas has only flat cells, so nested values are flattened into one cell rather than being spread across columns.
Anything visual in the HTML - weight, alignment, colour, column widths - has no counterpart in pandas and is dropped. The values are what survives.
The identifier 007 comes out of the Pandas DataFrame as 007. Reading it as a number would have made it 7, and an id that changes value is worse than one that stays text.
The role "Analyst, data" survives with its comma, in one cell rather than split across two. That is the first thing to check in any converted table, and the usual place a Pandas DataFrame file goes wrong.
The quotes around Jonah "Jo" Pryce are escaped with a backslash in the Pandas DataFrame.