What is Avro Schema?

An Avro schema describes a record: its name and its typed fields. It is the contract, not the data.

Who uses it

Kafka pipelines and anything in the Hadoop family.

Writing Avro Schema

Types are inferred per column from the values present. Any column with a gap in it becomes a union with null, because a reader rejects a record whose non-nullable field is missing.

What it can hold

Avro is a data interchange format. Avro carries its own schema, so the column types travel with the data instead of being inferred on the way in. A machine can read the table back out of it without guessing, which is what makes it safe to convert into and out of.

It keeps

It cannot express

What it looks like

Example Avro Schema
{
  "type": "record",
  "name": "staff",
  "fields": [
    {
      "name": "id",
      "type": "long"
    },
    {
      "name": "name",
      "type": "string"
    },
    {
      "name": "role",
      "type": "string"
    },
    {
      "name": "started",
      "type": "string"
    },
    {
      "name": "hours",
      "type": [
        "null",
        "double"
      ]
    }
  ]
}

Questions

How do I open Avro Schema files?

Avro Schema is written by Tablizer rather than read by it. Open the file in whichever tool the format belongs to.

Is Avro Schema a good format for sharing data?

It depends who is receiving it. Kafka pipelines and anything in the Hadoop family. If the recipient is outside that group, a plain CSV travels further than anything else.

Convert something else to Avro Schema