DBScholar

Back to papers

Pushing Data-Induced Predicates Through Joins in Big-Data Clusters

Summary: Introduces data-induced predicates that propagate filters across joins, enabling optimizer-only data skipping without execution overhead. Using existing zone maps—and modestly richer statistics—substantially reduces input and roughly doubles median production-cluster query speed. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
12318
Venue
VLDB
Year
2020
Pagerank
7.7204167e-05
Overall Rank
3,137 | 78.48%
DOI
10.14778/3368289.3368292

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{kandula_vldb20,
        title = {{Pushing Data-Induced Predicates Through Joins in Big-Data Clusters}},
        author = {Kandula, Srikanth and Orr, Laurel and Chaudhuri, Surajit},
        journal = {PVLDB},
        series = {{VLDB} '20},
        volume = {13},
        number = {3},
        pages = {252--265},
        doi = {10.14778/3368289.3368292},
        url = {https://doi.org/10.14778/3368289.3368292},
        year = {2020}
}

Incoming Citations (Sorted by Pagerank)

Showing 21 of 21 citing papers.

Rank Citing Paper Year Venue Pagerank
2,202 Quantifying TPC-H Choke Points and Their Optimizations 2020 VLDB 8.9639459e-05
2,865 End-to-end Optimization of Machine Learning Prediction Queries 2022 SIGMOD 8.0180243e-05
3,035 Instance-Optimized Data Layouts for Cloud Analytics Workloads 2021 SIGMOD 7.8297746e-05
4,553 Predicate Transfer: Efficient Pre-Filtering on Multi-Join Queries 2024 CIDR 6.6346951e-05
5,348 Crystal: A Unified Cache Storage System for Analytical Databases 2021 VLDB 6.2562688e-05
6,031 Sia: Optimizing Queries using Learned Predicates 2021 SIGMOD 6.0007422e-05
6,042 Pando: Enhanced Data Skipping with Logical Data Partitioning 2023 VLDB 5.9970052e-05
6,121 The Cosmos Big Data Platform at Microsoft: Over a Decade of Progress and a Decade to Look Forward 2021 VLDB 5.9688569e-05
6,593 Simple Adaptive Query Processing vs. Learned Query Optimizers: Observations and Analysis 2023 VLDB 5.8297039e-05
7,139 Selection Pushdown in Column Stores using Bit Manipulation Instructions 2023 SIGMOD 5.6932481e-05
7,660 Pruning in Snowflake: Working Smarter, Not Harder 2025 SIGMOD 5.5736132e-05
7,760 NOCAP: Near-Optimal Correlation-Aware Partitioning Joins 2023 SIGMOD 5.5505651e-05
8,056 Sieve: A Learned Data-Skipping Index for Data Analytics 2023 VLDB 5.4983582e-05
8,413 Conditional Cuckoo Filters 2021 SIGMOD 5.4306049e-05
8,465 Predicate Pushdown for Data Science Pipelines 2023 SIGMOD 5.4194578e-05
8,721 Accelerate Distributed Joins with Predicate Transfer 2025 SIGMOD 5.3772617e-05
8,927 Hyperspace: The Indexing Subsystem of Azure Synapse 2021 VLDB 5.3483178e-05
9,961 Threshold Queries in Theory and in the Wild 2022 VLDB 5.1879626e-05
10,690 Dynamic Pruning for Recursive Joins 2025 SIGMOD 5.093636e-05
11,166 PLAQUE: Automated Predicate Learning at Query Time 2024 SIGMOD 5.093636e-05
11,413 SH2O: Efficient Data Access for Work-Sharing Databases 2023 SIGMOD 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 41 of 41 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
6 Pig Latin: A Not-So-Foreign Language for Data Processing 2008 SIGMOD 0.0010686205
16 MAGIC SETS AND OTHER STRANGE WAYS TO IMPLEMENT LOGIC PROGRAMS (Extended Abstract) 1986 PODS 0.00060089598
18 How Good Are Query Optimizers, Really? 2016 VLDB 0.00059284255
24 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00054865648
30 SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets 2008 VLDB 0.00051174276
32 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00050111008
87 Automated Selection of Materialized Views and Indexes for SQL Databases 2000 VLDB 0.00035281619
159 CORDS: Automatic Discovery of Correlations and Soft Functional Dependencies 2004 SIGMOD 0.00028129426
186 The Vertica Analytic Database: C-Store 7 Years Later 2012 VLDB 0.00026182534
235 Fast Incremental Maintenance of Approximate Histograms 1997 VLDB 0.00023783792
259 Database Cracking 2007 CIDR 0.00023119313
448 Self-tuning Histograms: Building Histograms Without Looking at Data 1999 SIGMOD 0.00018292618
451 Mergeable Summaries 2012 PODS 0.00018151445
501 Database Tuning Advisor for Microsoft SQL Server 2005 2004 VLDB 0.0001738508
607 F1: A Distributed SQL Database That Scales 2013 VLDB 0.00015800238
736 Materialized View Maintenance and Integrity Constraint Checking: Trading Space for Time 1996 SIGMOD 0.00014508563
880 SuRF: Practical Range Query Filtering with Fast Succinct Tries 2018 SIGMOD 0.00013432693
1,037 Cost-Based Optimization for Magic: Algebra and Implementation 1996 SIGMOD 0.00012494928
1,044 Fine-grained Partitioning for Aggressive Data Skipping 2014 SIGMOD 0.0001244236
1,124 Query Optimization by Predicate Move-Around 1994 VLDB 0.00012087356
1,320 From Theory to Practice: Efficient Join Query Evaluation in a Parallel Database System 2015 SIGMOD 0.00011166426
1,324 Apache Hadoop Goes Realtime at Facebook 2011 SIGMOD 0.00011149314
1,362 Execution Strategies for SQL Subqueries 2007 SIGMOD 0.00011032204
1,648 BHUNT: Automatic Discovery of Fuzzy Algebraic Constraints in Relational Data 2003 VLDB 0.00010120668
2,150 Brighthouse: An Analytic Data Warehouse for Ad-hoc Queries 2008 VLDB 9.081101e-05
2,156 Quickstep: A Data Platform Based on the Scaling-Up Approach 2018 VLDB 9.0635624e-05
2,302 Correlation Maps: A Compressed Access Method for Exploiting Soft Functional Dependencies 2009 VLDB 8.7808696e-05
2,378 The Uncracked Pieces in Database Cracking 2014 VLDB 8.6682285e-05
2,642 CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop 2011 VLDB 8.3059948e-05
2,692 Efficient View Maintenance at Data Warehouses 1997 SIGMOD 8.2491468e-05
2,937 Column Sketches: A Scan Accelerator for Rapid and Robust Predicate Evaluation 2018 SIGMOD 7.9435581e-05
3,106 Skipping-oriented Partitioning for Columnar Layouts 2017 VLDB 7.7515666e-05
3,200 Locality-aware Partitioning in Parallel Database Systems 2015 SIGMOD 7.642132e-05
3,413 Slalom: Coasting Through Raw Data via Adaptive Partitioning and Indexing 2017 VLDB 7.4326381e-05
3,576 Two Birds, One Stone: A Fast, yet Lightweight, Indexing Scheme for Modern Database Systems 2017 VLDB 7.2936598e-05
3,650 Implementation of Magic-sets in a Relational Database System 1994 SIGMOD 7.2249961e-05
4,456 AdaptDB: Adaptive Partitioning for Distributed Joins 2017 VLDB 6.692321e-05
6,907 Statisticum: Data Statistics Management in SAP HANA 2017 VLDB 5.7417459e-05
7,870 Optimizing Iceberg Queries with Complex Joins 2017 SIGMOD 5.5272726e-05
8,996 High Performance Stream Query Processing With Correlation-Aware Partitioning 2014 VLDB 5.3357805e-05
9,953 Amoeba: A Shape changing Storage System for Big Data 2016 VLDB 5.1901412e-05
Previous Page 1 / 1 Next

Semantically Similar Papers