How to Convert CSV to JSON, and What Happens to Your Data Types

Last updated 19 August 2026 · about 6 minutes to read

Converting CSV to JSON is easy. Getting the types right is not. Leading zeros, empty cells, nested keys and the choices that decide whether your data survives.

Converting CSV to JSON takes one click. The part that goes wrong is not the conversion, it is the four decisions the conversion has to make on your behalf, because CSV does not record enough information to make them for you.

CSV has no types, and JSON has six

A CSV file is text. The value 42 in a CSV is the characters four and two, not the number forty-two. There is no flag anywhere in the file saying which columns are numeric, and no convention that every tool agrees on.

JSON has strings, numbers, booleans, null, objects and arrays. Something has to decide which of those each cell becomes, and whatever decides is guessing.

Most converters guess aggressively, because the output looks better. Numbers without quotes read as data rather than as text about data. The cost of that guess is paid by every value that looks numeric and is not.

The leading zero problem

A US zip code of 07030 is a five-character identifier for Hoboken, New Jersey. Read it as a number and it becomes 7030, which is four characters and not a zip code at all. The same thing happens to UK phone numbers starting 0, to account numbers, to product codes, and to any ISBN beginning with zero.

This is why type inference is off by default on Tablizer, and why even with it on, a value that starts with a zero followed by another digit is left as text. There is no realistic case where a quantity is written 007 and a large number of cases where an identifier is.

The precision problem

JSON numbers are IEEE 754 doubles in every implementation that matters. That gives about fifteen significant digits. A nineteen-digit identifier, which is what a Twitter status id or a Snowflake key looks like, does not survive: 1234567890123456789 comes back as 1234567890123456800.

The failure is silent. The document parses, the field is present, and the value is wrong in the last four digits. If your identifiers are long, keep them as strings.

Empty cells become one of three things

A CSV has exactly one way to say nothing: an empty field. JSON has three, and they mean different things to whatever reads the file next.

  • An empty string. The key is present, the value is text of zero length. This round-trips: convert back to CSV and you get the file you started with.
  • null. The key is present and explicitly has no value. Most databases map this to NULL, which is what you want if the column is nullable.
  • The key omitted. The record simply lacks the field. Some validators treat a missing optional field differently from a present null one, and some deserialisers will use a default value for a missing key and not for a null.

Pick based on what reads the file, not on which looks tidiest. If you do not know, empty string is the safest because it loses the least.

Headers become keys, and headers are messy

Two CSV columns can share a name. JSON object keys cannot. When that happens the second one has to be renamed, and a converter that silently drops one instead is losing a column.

Column names with spaces, dots or brackets in them are legal JSON keys but awkward to reach in code, because obj.total sales is not valid syntax and obj["total sales"] is. That is a style question rather than a correctness one, and worth fixing in the grid before you export rather than in your code afterwards.

Choosing the JSON shape

Four shapes are common, and they are not interchangeable.

The same three rows in four shapes
// array of objects: the default, and what most code expects
[{ "id": "007", "name": "Halima" }, { "id": "012", "name": "Jonah" }]

// 2D array: smaller, and the consumer has to know the column order
[["id", "name"], ["007", "Halima"], ["012", "Jonah"]]

// column arrays: what charting libraries usually want
{ "id": ["007", "012"], "name": ["Halima", "Jonah"] }

// keyed by first column: for lookups rather than iteration
{ "007": { "name": "Halima" }, "012": { "name": "Jonah" } }

The array of objects is right for almost everything. It repeats the keys on every record, which costs bytes and buys the ability to read any single record without reference to the file around it. That property is worth more than the bytes on all but the largest files.

A checklist before you convert

  • Look at any column that holds an identifier. If it has leading zeros or more than fifteen digits, keep it as a string.
  • Decide what an empty cell should become, and check that the thing reading the JSON agrees.
  • Check for duplicate header names before you export rather than after.
  • If the CSV came from Excel, check the dates. Excel exports them in whatever format the machine was set to, so an export made in London and one made in New York will disagree about 03/04/2026.

Tools mentioned here

Questions

Should I turn on type detection?

Only if you know the columns are quantities rather than identifiers. Numbers you will do arithmetic on should be numbers. Anything that identifies something, even if it is written with digits, should stay a string.

How do I convert CSV to nested JSON?

You cannot, directly. A CSV is flat and nesting is not recorded anywhere in it. Convert to a flat structure and reshape it in code, or use dotted column names and expand them yourself after parsing.

More guides