Database Paper Browser

Back to papers

Predicate Pushdown for Data Science Pipelines

Summary: MagicPush uses a search-verification approach to predicate pushdown in data science pipelines, discovering input-space predicates and proving pushdown preserves outputs, even with non-relational operators and UDFs. Evaluations on TPC-H and 200 real-world GitHub Notebook pipelines show it beats a strong rule-based baseline, discovers new pushdown opportunities, and yields up to 99% running-time reduction in 42 pipelines while matching baseline opportunities elsewhere. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
6640
Venue
SIGMOD
Year
2023
Pagerank
4.4784651e-05
Overall Rank
8,625 | 40.06%
DOI
10.1145/3589281

Incoming Non-self Citations Over Time

Authors

Incoming Citations (Sorted by Pagerank)

Showing 4 of 4 citing papers.

Rank Citing Paper Year Venue Pagerank
9,765 The UDFBench Benchmark for General-purpose UDF Queries 2025 VLDB 4.2815042e-05
10,152 Data-Semantics-Aware Recommendation of Diverse Pivot Tables 2026 SIGMOD 4.1905499e-05
10,415 Dynamic Pruning for Recursive Joins 2025 SIGMOD 4.1905499e-05
10,858 LiquidCache: Efficient Pushdown Caching for Cloud-Native Data Analytics 2025 VLDB 4.1905499e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 27 of 27 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
66 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00061707583
140 Predicate Migration: Optimizing Queries with Expensive Predicates 1993 SIGMOD 0.00042289025
335 Optimization of Real Conjunctive Queries 1993 PODS 0.00027012705
1,107 Froid: Optimization of Imperative Programs in a Relational Database 2018 VLDB 0.0001397627
1,205 PIVOT and UNPIVOT: Optimization and Execution Strategies in an RDBMS 2004 VLDB 0.00013309393
1,303 Query Optimization by Predicate Move-Around 1994 VLDB 0.00012692678
1,608 Qd-tree: Learning Data Layouts for Big Data Analytics 2020 SIGMOD 0.00011169837
1,884 Tuplex: Data Science in Python at Native Code Speed 2021 SIGMOD 0.00010206514
2,123 SQL-on-Hadoop: Full Circle Back to Shared-Nothing Database Architectures 2014 VLDB 9.4836925e-05
2,595 WeTune: Automatic Discovery and Verification of Query Rewrite Rules 2022 SIGMOD 8.4725961e-05
2,825 Mison: A Fast JSON Parser for Data Analytics 2017 VLDB 8.0575959e-05
2,914 Quantifying TPC-H Choke Points and Their Optimizations 2020 VLDB 7.9197583e-05
2,955 Magpie: Python at Speed and Scale using Cloud Backends 2021 CIDR 7.8188583e-05
3,153 AnalyticDB: Real-time OLAP Database System at Alibaba Cloud 2019 VLDB 7.4706916e-05
3,254 Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science Notebooks 2020 SIGMOD 7.3172324e-05
3,430 Demonstration of the Cosette Automated SQL Prover 2017 SIGMOD 7.0985442e-05
3,903 Automated Verification of Query Equivalence Using Satisfiability Modulo Theories 2019 VLDB 6.6439695e-05
3,923 Pushing Data-Induced Predicates Through Joins in Big-Data Clusters 2020 VLDB 6.6232068e-05
4,645 Aggify: Lifting the Curse of Cursor Loops using Custom Aggregates 2020 SIGMOD 6.0190618e-05
4,672 FlexPushdownDB: Hybrid Pushdown and Caching in a Cloud DBMS 2021 VLDB 6.001444e-05
4,676 Automatically Leveraging MapReduce Frameworks for Data-Intensive Applications 2018 SIGMOD 5.9992323e-05
6,152 Crystal: A Unified Cache Storage System for Analytical Databases 2021 VLDB 5.1802666e-05
6,671 Incorporating Super-Operators in Big-Data Query Optimizers 2020 VLDB 4.9625353e-05
6,703 YeSQL: “You extend SQL” with Rich and Highly Performant User-Defined Functions in Relational Databases 2022 VLDB 4.9514593e-05
7,278 Sia: Optimizing Queries using Learned Predicates 2021 SIGMOD 4.7720613e-05
7,338 Optimizing Recursive Queries with Program Synthesis 2022 SIGMOD 4.7531793e-05
9,818 Generating Application-Specific Data Layouts for In-memory Databases 2019 VLDB 4.2733415e-05
Previous Page 1 / 1 Next

Semantically Similar Papers