Hasty Briefsbeta

Bilingual

RIP, Vector Database

19 hours ago
  • turbopuffer is transitioning to a new storage architecture called 'turbopuffer v3' that changes how documents and indexes are laid out, written, compacted, and queried.
  • The new engine aims to support more query plans at greater scale by making ANN 'just another' secondary index instead of the primary index.
  • Originally a serverless vector database focused on cheap vector searches, turbopuffer evolved into a generalized search database with attribute filtering and full-text search, but the storage architecture remained centered around the ANN index.
  • The current architecture causes storage amplification (duplicating document data for multi-vector), write amplification (rebalancing moves entire documents and indexes), and limited vectorization (block size constrained by cluster size).
  • The solution is to stop keying data on ANN addresses; turbopuffer v3 implements this change.
  • All CI tests pass on v3 but there is a significant performance regression, which the team is now tuning.
  • Future updates will include benchmarks and detailed explanations of the new architecture and optimizations.