What HTML is
An HTML table is a table element containing thead and tbody, with th cells for headers and td cells for data.
Web pages, email templates and anything pasted into a CMS.
What Avro Schema is
An Avro schema describes a record: its name and its typed fields. It is the contract, not the data.
Kafka pipelines and anything in the Hadoop family.
What changes when you convert HTML to Avro Schema
The first table in the markup is used. Cell contents are stripped of inline tags and entities are decoded, so <td><b>Total</b> </td> arrives as Total. Nested tables inside a cell are not descended into.
Types are inferred per column from the values present. Any column with a gap in it becomes a union with null, because a reader rejects a record whose non-nullable field is missing.
What carries over from HTML to Avro Schema
HTML records 2 things about a table that Avro has no way to hold.
Avro wants a type for each column, and HTML does not record one, so each column is typed from what its values look like. A column of digits that should stay text - a zip code, a phone number, a leading-zero id - is the usual thing to check afterwards.
Avro has no header row. The column names are repeated as keys on every record instead, which is why the result is bigger on disk than the table it came from.
Anything visual in the HTML - weight, alignment, colour, column widths - has no counterpart in Avro and is dropped. The values are what survives.
The result can be read a record at a time, and appended to by adding to the end of it. That is worth having for a table too large to hold in memory, and it means a Avro file that is cut off part way through still gives you every record before the cut.