RIP, Vector Database
19 hours ago
- turbopuffer is transitioning to a new storage architecture called 'turbopuffer v3' that changes how documents and indexes are laid out, written, compacted, and queried.
- The new engine aims to support more query plans at greater scale by making ANN 'just another' secondary index instead of the primary index.
- Originally a serverless vector database focused on cheap vector searches, turbopuffer evolved into a generalized search database with attribute filtering and full-text search, but the storage architecture remained centered around the ANN index.
- The current architecture causes storage amplification (duplicating document data for multi-vector), write amplification (rebalancing moves entire documents and indexes), and limited vectorization (block size constrained by cluster size).
- The solution is to stop keying data on ANN addresses; turbopuffer v3 implements this change.
- All CI tests pass on v3 but there is a significant performance regression, which the team is now tuning.
- Future updates will include benchmarks and detailed explanations of the new architecture and optimizations.