DBScholar

Back to papers

Pushing Data-Induced Predicates Through Joins in Big-Data Clusters

Summary: Introduces data-induced predicates that propagate filters across joins, enabling optimizer-only data skipping without execution overhead. Using existing zone maps—and modestly richer statistics—substantially reduces input and roughly doubles median production-cluster query speed. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
he9c9b372d221fd91
Venue
VLDB
Year
2020
Pagerank
7.6777283e-05
Overall Rank
3,073 | 79.35%
DOI
10.14778/3368289.3368292

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{kandula_vldb20,
        title = {{Pushing Data-Induced Predicates Through Joins in Big-Data Clusters}},
        author = {Kandula, Srikanth and Orr, Laurel and Chaudhuri, Surajit},
        journal = {PVLDB},
        series = {{VLDB} '20},
        volume = {13},
        number = {3},
        pages = {252--265},
        doi = {10.14778/3368289.3368292},
        url = {https://doi.org/10.14778/3368289.3368292},
        year = {2020}
}

Incoming Citations (Sorted by Pagerank)

Showing 22 of 22 citing papers.

Rank Citing Paper Year Venue Pagerank
1,952 Quantifying TPC-H Choke Points and Their Optimizations 2020 VLDB 9.3189525e-05
2,662 End-to-end Optimization of Machine Learning Prediction Queries 2022 SIGMOD 8.1596229e-05
2,765 Instance-Optimized Data Layouts for Cloud Analytics Workloads 2021 SIGMOD 8.0439015e-05
4,369 Predicate Transfer: Efficient Pre-Filtering on Multi-Join Queries 2024 CIDR 6.6315141e-05
5,446 Crystal: A Unified Cache Storage System for Analytical Databases 2021 VLDB 6.1266228e-05
5,854 Pando: Enhanced Data Skipping with Logical Data Partitioning 2023 VLDB 5.9708829e-05
6,147 Sia: Optimizing Queries using Learned Predicates 2021 SIGMOD 5.8710365e-05
6,245 The Cosmos Big Data Platform at Microsoft: Over a Decade of Progress and a Decade to Look Forward 2021 VLDB 5.837329e-05
6,392 Selection Pushdown in Column Stores using Bit Manipulation Instructions 2023 SIGMOD 5.800848e-05
6,710 Simple Adaptive Query Processing vs. Learned Query Optimizers: Observations and Analysis 2023 VLDB 5.7019157e-05
7,245 Pruning in Snowflake: Working Smarter, Not Harder 2025 SIGMOD 5.5761132e-05
7,923 NOCAP: Near-Optimal Correlation-Aware Partitioning Joins 2023 SIGMOD 5.4260253e-05
8,153 Predicate Pushdown for Data Science Pipelines 2023 SIGMOD 5.3891786e-05
8,199 Sieve: A Learned Data-Skipping Index for Data Analytics 2023 VLDB 5.3788727e-05
8,269 Accelerate Distributed Joins with Predicate Transfer 2025 SIGMOD 5.3648571e-05
8,580 Conditional Cuckoo Filters 2021 SIGMOD 5.3100462e-05
9,088 Hyperspace: The Indexing Subsystem of Azure Synapse 2021 VLDB 5.2283159e-05
10,153 Threshold Queries in Theory and in the Wild 2022 VLDB 5.0715586e-05
10,843 I-Rex: An Interactive Debugger for SQL 2026 VLDB 4.9793485e-05
11,126 Dynamic Pruning for Recursive Joins 2025 SIGMOD 4.9793485e-05
11,511 PLAQUE: Automated Predicate Learning at Query Time 2024 SIGMOD 4.9793485e-05
11,727 SH2O: Efficient Data Access for Work-Sharing Databases 2023 SIGMOD 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 41 of 41 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
6 Pig Latin: A Not-So-Foreign Language for Data Processing 2008 SIGMOD 0.001052036
15 How Good Are Query Optimizers, Really? 2016 VLDB 0.00061066921
18 MAGIC SETS AND OTHER STRANGE WAYS TO IMPLEMENT LOGIC PROGRAMS (Extended Abstract) 1986 PODS 0.00059023577
23 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00055406774
30 SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets 2008 VLDB 0.00050495102
31 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00049839909
88 Automated Selection of Materialized Views and Indexes for SQL Databases 2000 VLDB 0.00035351639
160 CORDS: Automatic Discovery of Correlations and Soft Functional Dependencies 2004 SIGMOD 0.00027837289
179 The Vertica Analytic Database: C-Store 7 Years Later 2012 VLDB 0.00026611886
242 Fast Incremental Maintenance of Approximate Histograms 1997 VLDB 0.00023363722
253 Database Cracking 2007 CIDR 0.00023042111
454 Self-tuning Histograms: Building Histograms Without Looking at Data 1999 SIGMOD 0.00017962189
456 Mergeable Summaries 2012 PODS 0.0001791284
491 Database Tuning Advisor for Microsoft SQL Server 2005 2004 VLDB 0.00017413042
617 F1: A Distributed SQL Database That Scales 2013 VLDB 0.00015555815
744 Materialized View Maintenance and Integrity Constraint Checking: Trading Space for Time 1996 SIGMOD 0.00014309622
891 SuRF: Practical Range Query Filtering with Fast Succinct Tries 2018 SIGMOD 0.00013245926
1,032 Cost-Based Optimization for Magic: Algebra and Implementation 1996 SIGMOD 0.00012401489
1,036 Fine-grained Partitioning for Aggressive Data Skipping 2014 SIGMOD 0.00012377471
1,119 Query Optimization by Predicate Move-Around 1994 VLDB 0.00011950395
1,292 From Theory to Practice: Efficient Join Query Evaluation in a Parallel Database System 2015 SIGMOD 0.00011152286
1,302 Apache Hadoop Goes Realtime at Facebook 2011 SIGMOD 0.00011111126
1,378 Execution Strategies for SQL Subqueries 2007 SIGMOD 0.00010864448
1,662 BHUNT: Automatic Discovery of Fuzzy Algebraic Constraints in Relational Data 2003 VLDB 9.9465656e-05
2,137 Quickstep: A Data Platform Based on the Scaling-Up Approach 2018 VLDB 8.9777553e-05
2,155 Brighthouse: An Analytic Data Warehouse for Ad-hoc Queries 2008 VLDB 8.9456729e-05
2,329 Correlation Maps: A Compressed Access Method for Exploiting Soft Functional Dependencies 2009 VLDB 8.6292256e-05
2,383 The Uncracked Pieces in Database Cracking 2014 VLDB 8.5435328e-05
2,687 CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop 2011 VLDB 8.1301312e-05
2,741 Efficient View Maintenance at Data Warehouses 1997 SIGMOD 8.0688837e-05
2,952 Column Sketches: A Scan Accelerator for Rapid and Robust Predicate Evaluation 2018 SIGMOD 7.8153507e-05
3,081 Skipping-oriented Partitioning for Columnar Layouts 2017 VLDB 7.6653727e-05
3,239 Locality-aware Partitioning in Parallel Database Systems 2015 SIGMOD 7.5007569e-05
3,445 Slalom: Coasting Through Raw Data via Adaptive Partitioning and Indexing 2017 VLDB 7.2943981e-05
3,597 Two Birds, One Stone: A Fast, yet Lightweight, Indexing Scheme for Modern Database Systems 2017 VLDB 7.1788912e-05
3,707 Implementation of Magic-sets in a Relational Database System 1994 SIGMOD 7.0820673e-05
4,533 AdaptDB: Adaptive Partitioning for Distributed Joins 2017 VLDB 6.5565658e-05
6,999 Statisticum: Data Statistics Management in SAP HANA 2017 VLDB 5.6252486e-05
8,023 Optimizing Iceberg Queries with Complex Joins 2017 SIGMOD 5.4054487e-05
9,154 High Performance Stream Query Processing With Correlation-Aware Partitioning 2014 VLDB 5.2176811e-05
10,124 Amoeba: A Shape changing Storage System for Big Data 2016 VLDB 5.0753201e-05
Previous Page 1 / 1 Next

Semantically Similar Papers