The Query Transformation Pipeline
4 hours ago
- Readyset replaces traditional pull-based query execution with a dataflow graph that incrementally maintains cached results as data changes, eliminating repeated query execution.
- The query rewrite pipeline transforms arbitrary SQL into a canonical form compilable by the dataflow engine, organized into three blocks: Normalization (desugaring, schema resolution), Deep Rewrites (decorrelation, inlining, join reordering), and Cleanup (removing redundancies, parameterizing literals).
- Key transformations include array constructor rewrite, redundant join elimination, left-spine hoisting for Top-K patterns, subquery decorrelation with three-valued logic handling, derived table inlining, and join reordering for canonical cache sharing.
- The pipeline ensures semantic preservation and satisfies dataflow engine constraints: binary joins with equality predicates, no correlated execution, flat join structures, and explicit GROUP BY/ORDER BY references.
- Future work includes cost-based join reordering, relaxing transformation guardrails, and adding pattern-specific rewrites for common ORM and BI tool query shapes.