Convert PDF to Avro Schema Online

Paste a PDF table below, edit it if you need to, and get Avro Schema back. The conversion runs in your browser, so nothing is uploaded, there is no size limit and there is no signup.

From

The only part of this site that uses a server. The address is sent to it so the page can be fetched, because a browser is not allowed to fetch someone else's page itself. Your own files are never involved.

Nothing here is uploaded. The parsing happens on this page.
Edit
Paste or upload something above and it will appear here.
ToDataMarkupDocumentsCodeSchemasImages

The output will appear here once there is a table to convert.

What PDF is

PDF describes a page: glyphs at coordinates, not rows and columns. A table in a PDF is a visual arrangement, not a data structure.

Reports, invoices, statements and anything sent to be read rather than processed.

What Avro Schema is

An Avro schema describes a record: its name and its typed fields. It is the contract, not the data.

Kafka pipelines and anything in the Hadoop family.

What changes when you convert PDF to Avro Schema

Every character in a text-based PDF carries its position on the page. Characters sharing a baseline are one row. Columns are found from the vertical strips no character ever occupies, which is why extraction works cleanly on a document laid out with spacing and badly on one laid out with ruled lines. A header repeated at the top of each page is detected and dropped rather than landing in the middle of the data.

Types are inferred per column from the values present. Any column with a gap in it becomes a union with null, because a reader rejects a record whose non-nullable field is missing.

What carries over from PDF to Avro Schema

One property of the PDF has no home in Avro, and it is worth knowing which before you convert.

Avro wants a type for each column, and PDF does not record one, so each column is typed from what its values look like. A column of digits that should stay text - a zip code, a phone number, a leading-zero id - is the usual thing to check afterwards.

Anything visual in the PDF - weight, alignment, colour, column widths - has no counterpart in Avro and is dropped. The values are what survives.

The result can be read a record at a time, and appended to by adding to the end of it. That is worth having for a table too large to hold in memory, and it means a Avro file that is cut off part way through still gives you every record before the cut.

This is the direction that recovers structure: PDF has no table in it to read, only an arrangement that looks like one, so the rows and columns are inferred rather than read.

A worked example

Three rows of staff data, with an identifier that has a leading zero, a value containing a comma, a value containing quotes and one blank cell. Those are the four places formats disagree, so they are the four places to look.

PDF in
id   name              role           started     hours
007  Halima Yusuf      Analyst, data  2024-03-15  38.5
012  Jonah "Jo" Pryce  Engineer       2025-11-02
104  Wei Chen          Manager        2023-06-30  40
Avro Schema out
{
  "type": "record",
  "name": "staff",
  "fields": [
    {
      "name": "id",
      "type": "long"
    },
    {
      "name": "name",
      "type": "string"
    },
    {
      "name": "role",
      "type": "string"
    },
    {
      "name": "started",
      "type": "string"
    },
    {
      "name": "hours",
      "type": [
        "null",
        "double"
      ]
    }
  ]
}

Questions

How do I convert PDF to Avro Schema?

Paste your PDF into the box above or drop the file onto it. Check the table in the grid, then copy or download the Avro Schema from the output panel. It takes one step and the data never leaves your browser.

Can I get my table back out of the PDF?

Out of one made here, yes, because the text is real text. Out of a scanned document, no. A scan is an image of a page, and pulling data from it needs optical character recognition, which is a different job with a different failure rate.

Is there a limit on file size?

No. The work happens on your own machine, so the limit is your machine's memory rather than an upload cap. A file with tens of thousands of rows converts in a second or two.

Is my data uploaded anywhere?

No. The parsing and generating are done by JavaScript running on this page. You can watch the network tab while you convert and see that nothing is sent.

Related conversions