ParadeDB Search Performance Improvements
2 hours ago
- PlanetScale released TIN, a Postgres text search extension with BM25 search and document counts, claiming significant performance advantages.
- ParadeDB initially performed slower but quickly closed the gap through targeted optimizations without changing document identifiers.
- Optimization 1: Storing fieldnorm arrays per term alongside postings lists reduced random page accesses from ~1,500 to 30.
- Optimization 2: Using MAXSCORE instead of WAND for disjunction queries with many terms improved latency up to 8x.
- Benchmark anomalies: ParadeDB syntax caused searches over multiple fields, and TIN's dense-term elision approximates BM25, causing accuracy trade-offs.
- TIN uses Postgres ctid as document identifiers, eliminating DocId–ctid mapping, but ParadeDB argues dense u32 DocId values offer better compression and columnar integration.
- ParadeDB embedded optimizations into Tantivy, staying committed to open-source collaboration, and will release improvements in v0.26.0.
- Part II will cover COUNT performance optimizations.