What XML is
XML marks up data with nested named tags and optional attributes. It is verbose, strict about being well-formed, and can carry a schema that says what shape a document must have.
Enterprise integrations, government data feeds, RSS, SOAP services and older document formats.
What Pandas DataFrame is
Python that builds a pandas DataFrame from a dictionary of columns.
Data analysis and notebooks.
What changes when you convert XML to Pandas DataFrame
The first repeating element is treated as the rows. Attributes are read alongside child elements and prefixed so they do not collide with a same-named tag. Nesting one level deep becomes dotted column names; deeper structures are kept as text.
The import line is included so the snippet runs as is. Turn type inference on and numeric columns arrive as numbers rather than strings, which is usually what you want in a DataFrame.
What carries over from XML to Pandas DataFrame
One property of the XML has no home in pandas, and it is worth knowing which before you convert.
pandas wants a type for each column, and XML does not record one, so each column is typed from what its values look like. A column of digits that should stay text - a zip code, a phone number, a leading-zero id - is the usual thing to check afterwards.
XML can hold a value that is itself a list or an object. pandas has only flat cells, so nested values are flattened into one cell rather than being spread across columns.
pandas marks header cells differently from data cells, so the keys are lifted out and written once as a marked header row rather than repeated on every row.
The identifier 007 comes out of the Pandas DataFrame as 007. Reading it as a number would have made it 7, and an id that changes value is worse than one that stays text.
The role "Analyst, data" survives with its comma, in one cell rather than split across two. That is the first thing to check in any converted table, and the usual place a Pandas DataFrame file goes wrong.
The quotes around Jonah "Jo" Pryce are escaped with a backslash in the Pandas DataFrame.