DBScholar

Back to papers

To Join or Not to Join? Thinking Twice about Joins before Feature Selection

Summary: Safe-join avoidance for feature selection in normalized datasets: many join-derived features can be dropped without hurting ML accuracy. Experiments on real normalized datasets show accurate safety predictions and substantial runtime savings for popular feature selection methods. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h543d7bd254547c72
Venue
SIGMOD
Year
2016
Pagerank
0.00014048128
Overall Rank
779 | 94.77%
DOI
10.1145/2882903.2882952

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{kumar_sigmod16,
        title = {{To Join or Not to Join? Thinking Twice about Joins before Feature Selection}},
        author = {Kumar, Arun and Naughton, Jeffrey and Patel, Jignesh M. and Zhu, Xiaojin},
        series = {{SIGMOD} '16},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/2882903.2882952},
        url = {https://dl.acm.org/doi/10.1145/2882903.2882952},
        year = {2016}
}

Incoming Citations (Sorted by Pagerank)

Showing 41 of 41 citing papers.

Rank Citing Paper Year Venue Pagerank
1,038 ARDA: Automatic Relational Data Augmentation for Machine Learning 2020 VLDB 0.000123653
1,153 Data Management Challenges in Production Machine Learning 2017 SIGMOD 0.00011793347
1,223 Data Management in Machine Learning: Challenges, Techniques, and Systems 2017 SIGMOD 0.00011468426
1,257 Towards Linear Algebra over Normalized Data 2017 VLDB 0.0001130959
1,293 Finding Related Tables in Data Lakes for Interactive Data Science 2020 SIGMOD 0.00011144991
1,568 HELIX: Holistic Optimization for Accelerating Iterative Machine Learning 2019 VLDB 0.00010208225
2,162 DIFF: A Relational Interface for Large-Scale Data Explanation 2019 VLDB 8.9344773e-05
2,201 Enabling and Optimizing Non-linear Feature Interactions in Factorized Linear Algebra 2019 SIGMOD 8.8708356e-05
2,663 A Layered Aggregate Engine for Analytics Workloads 2019 SIGMOD 8.1542952e-05
2,802 Distributed Join Algorithms on Thousands of Cores 2017 VLDB 7.9866934e-05
2,908 AI Meets Database: AI4DB and DB4AI 2021 SIGMOD 7.8716173e-05
2,978 In-Database Learning with Sparse Tensors 2018 PODS 7.7872011e-05
3,016 Data Acquisition for Improving Machine Learning Models 2021 VLDB 7.7492494e-05
3,521 Ember: No-Code Context Enrichment via Similarity-Based Keyless Joins 2022 VLDB 7.2320843e-05
3,617 Leva: Boosting Machine Learning Performance with Relational Embedding Data Augmentation 2022 SIGMOD 7.1540505e-05
3,673 Are Key-Foreign Key Joins Safe to Avoid when Learning High-Capacity Classifiers? 2018 VLDB 7.1075403e-05
4,411 LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data Lakes 2024 VLDB 6.6089506e-05
4,902 Scalable Asynchronous Gradient Descent Optimization for Out-of-Core Models 2017 VLDB 6.3617262e-05
5,078 Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond 2021 SIGMOD 6.2832055e-05
5,279 MATE: Multi-Attribute Table Extraction 2022 VLDB 6.1973533e-05
5,533 Putting Things into Context: Rich Explanations for Query Answers using Join Graphs 2021 SIGMOD 6.0906707e-05
5,748 Optimizing Data Acquisition to Enhance Machine Learning Performance 2024 VLDB 6.0054988e-05
5,877 BlinkML: Efficient Maximum Likelihood Estimation with Probabilistic Guarantees 2019 SIGMOD 5.9599042e-05
6,146 Efficient Construction of Approximate Ad-Hoc ML models Through Materialization and Reuse 2018 VLDB 5.869373e-05
6,594 Tuple-oriented Compression for Large-scale Mini-batch Stochastic Gradient Descent 2019 SIGMOD 5.7394502e-05
6,747 Mitigating the Impedance Mismatch between Prediction Query Execution and Database Engine 2025 SIGMOD 5.6883491e-05
6,775 Causal Feature Selection for Algorithmic Fairness 2022 SIGMOD 5.6838053e-05
7,203 Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning 2023 VLDB 5.5845207e-05
7,872 Mind the Gap: Bridging Multi-Domain Query Workloads with EmptyHeaded 2017 VLDB 5.4339723e-05
8,501 SPRINTER: A Fast n-ary Join Query Processing Method for Complex OLAP Queries 2020 SIGMOD 5.3285575e-05
8,815 ApproxML: Efficient Approximate Ad-Hoc ML Models Through Materialization and Reuse 2019 VLDB 5.2703166e-05
9,261 Leveraging Similarity Joins for Signal Reconstruction 2018 VLDB 5.2032182e-05
10,457 Eliminating Redundant Feature Tests in Decision Tree and Random Forest Inference on SQL Predicates 2026 SIGMOD 4.9769913e-05
10,463 Factorized and Vectorized Execution: Optimizing Analytical and Semantic Queries over Relations 2026 SIGMOD 4.9769913e-05
10,664 InferF: Declarative Factorization of AI/ML Inferences over Joins 2026 SIGMOD 4.9769913e-05
10,749 Database Views as Explanations for Relational Deep Learning 2026 VLDB 4.9769913e-05
11,584 Relational Query Synthesis ⋈ Decision Tree Learning 2024 VLDB 4.9769913e-05
11,596 Enriching Relations with Additional Attributes for ER 2024 VLDB 4.9769913e-05
11,712 Regularized Pairwise Relationship based Analytics for Structured Data 2023 SIGMOD 4.9769913e-05
11,987 Enforcing Constraints for Machine Learning Systems via Declarative Feature Selection: An Experimental Study 2021 SIGMOD 4.9769913e-05
12,253 Learning Efficiently Over Heterogeneous Databases 2018 VLDB 4.9769913e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 8 of 8 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers