Hasty Briefsbeta

Bilingual

Let's See Paul Allen's SIMD CSV Parser

7 hours ago
  • The article presents a research CSV parser that leverages SIMD to process 64 bytes at a time, based on techniques from the simdjson paper.
  • SIMD enables processing multiple bytes simultaneously and works best with branchless code, avoiding if statements and function calls.
  • CSV parsing is broken into three steps: classify structural characters, filter out fake delimiters inside quoted fields, and collect delimiter positions.
  • Classification uses vectorized lookup tables with nibbles: each byte is split into high and low nibbles, and ANDing the two table results yields the correct class for commas, quotes, and newlines.
  • Classified bytes are compressed into bitmasks, where each bit represents whether a certain character class appears at a given byte position.
  • Inside quoted fields, a quote-prefix XOR (running parity) determines whether the parser is inside or outside quotes; escaped quotes cancel out, so the inverted prefix mask removes commas and newlines inside quoted fields.
  • Delimiter positions are extracted from the cleaned bitmasks using count-leading-zeros, yielding the offsets needed to split the byte stream into rows and fields.