Let's See Paul Allen's SIMD CSV Parser
7 hours ago
- The article presents a research CSV parser that leverages SIMD to process 64 bytes at a time, based on techniques from the simdjson paper.
- SIMD enables processing multiple bytes simultaneously and works best with branchless code, avoiding if statements and function calls.
- CSV parsing is broken into three steps: classify structural characters, filter out fake delimiters inside quoted fields, and collect delimiter positions.
- Classification uses vectorized lookup tables with nibbles: each byte is split into high and low nibbles, and ANDing the two table results yields the correct class for commas, quotes, and newlines.
- Classified bytes are compressed into bitmasks, where each bit represents whether a certain character class appears at a given byte position.
- Inside quoted fields, a quote-prefix XOR (running parity) determines whether the parser is inside or outside quotes; escaped quotes cancel out, so the inverted prefix mask removes commas and newlines inside quoted fields.
- Delimiter positions are extracted from the cleaned bitmasks using count-leading-zeros, yielding the offsets needed to split the byte stream into rows and fields.