What PDF is
PDF describes a page: glyphs at coordinates, not rows and columns. A table in a PDF is a visual arrangement, not a data structure.
Reports, invoices, statements and anything sent to be read rather than processed.
What YAML is
YAML uses indentation instead of brackets. It is a superset of JSON, supports comments, and is meant to be edited by hand.
Kubernetes manifests, CI pipelines, Ansible playbooks and most modern config files.
What changes when you convert PDF to YAML
Every character in a text-based PDF carries its position on the page. Characters sharing a baseline are one row. Columns are found from the vertical strips no character ever occupies, which is why extraction works cleanly on a document laid out with spacing and badly on one laid out with ruled lines. A header repeated at the top of each page is detected and dropped rather than landing in the middle of the data.
You get a list of maps, one per row, with long lines left unwrapped so values do not get folded across lines.
What carries over from PDF to YAML
One property of the PDF has no home in YAML, and it is worth knowing which before you convert.
YAML wants a type for each column, and PDF does not record one, so each column is typed from what its values look like. A column of digits that should stay text - a zip code, a phone number, a leading-zero id - is the usual thing to check afterwards.
Anything visual in the PDF - weight, alignment, colour, column widths - has no counterpart in YAML and is dropped. The values are what survives.
This is the direction that recovers structure: PDF has no table in it to read, only an arrangement that looks like one, so the rows and columns are inferred rather than read.