How to Split a Large CSV File
Last updated 19 August 2026 · about 4 minutes to read
Split a CSV by row count or by column value, keep the header on every part, and know why the obvious command line approach produces broken files.
Splitting a CSV is easy to do badly. The naive approach cuts the file into pieces where only the first one has a header, which means only the first one is a CSV.
Two reasons to split
A limit downstream. An import that caps at ten thousand rows, an API that takes five hundred records a call, a spreadsheet that will not open past a million lines. Here you split by row count and the sizes are arbitrary.
Or the file is several files stacked together. One CSV per region, per month or per customer is more useful than one CSV with a region column, especially if different people own different regions. Here you split by the value in a column and the sizes fall out of the data.
Why the header has to be repeated
A part without a header is not a CSV. It is a fragment that only means something next to the original file, and anything that opens it will treat the first data row as the column names. That row is then both missing from the data and wrong as a header.
Repeating the header costs one line per file. Every part then opens correctly on its own, which is the whole reason for splitting.
The command line version
The Unix split command cuts by line count and knows nothing about headers, so the parts after the first have none. Fixing that means extracting the header first and prepending it to each part afterwards, which is a few lines of shell and easy to get subtly wrong on the last chunk.
It also splits on bytes or lines, neither of which respects a quoted field containing a newline. A CSV with multi-line values can be cut in the middle of a record, which produces two invalid files and no error message.
Things to watch
- Quoted newlines. A value containing a line break spans two physical lines and one logical row. Any splitter that counts lines rather than records will eventually cut through one.
- Encoding. If the source has a byte order mark, only the first part will have it unless the splitter adds one to each.
- File naming when splitting by value. Column values can contain slashes, colons and other characters that are not legal in a filename, so they have to be replaced.
- Very high cardinality. Splitting by a column with fifty thousand distinct values produces fifty thousand files, which is rarely what anyone wanted.
Splitting by column value
Sort by that column first if you are going to inspect the results by hand. It makes no difference to the output and a large one to how easy it is to check.
Check the distinct count before you split. If it is more than a few dozen, filtering to the values you need is usually better than producing a file for every one of them.