DBScholar

Back to papers

To Join or Not to Join? Thinking Twice about Joins before Feature Selection

Summary: Safe-join avoidance for feature selection in normalized datasets: many join-derived features can be dropped without hurting ML accuracy. Experiments on real normalized datasets show accurate safety predictions and substantial runtime savings for popular feature selection methods. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h543d7bd254547c72
Venue
SIGMOD
Year
2016
Pagerank
0.00014054709
Overall Rank
777 | 94.78%
DOI
10.1145/2882903.2882952

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{kumar_sigmod16,
        title = {{To Join or Not to Join? Thinking Twice about Joins before Feature Selection}},
        author = {Kumar, Arun and Naughton, Jeffrey and Patel, Jignesh M. and Zhu, Xiaojin},
        series = {{SIGMOD} '16},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/2882903.2882952},
        url = {https://dl.acm.org/doi/10.1145/2882903.2882952},
        year = {2016}
}

Incoming Citations (Sorted by Pagerank)

Showing 41 of 41 citing papers.

Rank Citing Paper Year Venue Pagerank
1,038 ARDA: Automatic Relational Data Augmentation for Machine Learning 2020 VLDB 0.00012370691
1,153 Data Management Challenges in Production Machine Learning 2017 SIGMOD 0.00011798912
1,255 Data Management in Machine Learning: Challenges, Techniques, and Systems 2017 SIGMOD 0.00011325762
1,256 Towards Linear Algebra over Normalized Data 2017 VLDB 0.00011314687
1,293 Finding Related Tables in Data Lakes for Interactive Data Science 2020 SIGMOD 0.00011149857
1,568 HELIX: Holistic Optimization for Accelerating Iterative Machine Learning 2019 VLDB 0.0001021302
2,160 DIFF: A Relational Interface for Large-Scale Data Explanation 2019 VLDB 8.9364035e-05
2,199 Enabling and Optimizing Non-linear Feature Interactions in Factorized Linear Algebra 2019 SIGMOD 8.8750296e-05
2,663 A Layered Aggregate Engine for Analytics Workloads 2019 SIGMOD 8.1581558e-05
2,802 Distributed Join Algorithms on Thousands of Cores 2017 VLDB 7.9903139e-05
2,908 AI Meets Database: AI4DB and DB4AI 2021 SIGMOD 7.8742664e-05
2,975 In-Database Learning with Sparse Tensors 2018 PODS 7.7907759e-05
3,015 Data Acquisition for Improving Machine Learning Models 2021 VLDB 7.7529194e-05
3,521 Ember: No-Code Context Enrichment via Similarity-Based Keyless Joins 2022 VLDB 7.2351481e-05
3,617 Leva: Boosting Machine Learning Performance with Relational Embedding Data Augmentation 2022 SIGMOD 7.1574349e-05
3,670 Are Key-Foreign Key Joins Safe to Avoid when Learning High-Capacity Classifiers? 2018 VLDB 7.1108704e-05
4,409 LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data Lakes 2024 VLDB 6.6120807e-05
4,901 Scalable Asynchronous Gradient Descent Optimization for Out-of-Core Models 2017 VLDB 6.3647386e-05
5,076 Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond 2021 SIGMOD 6.2860582e-05
5,276 MATE: Multi-Attribute Table Extraction 2022 VLDB 6.2002884e-05
5,529 Putting Things into Context: Rich Explanations for Query Answers using Join Graphs 2021 SIGMOD 6.0935553e-05
5,746 Optimizing Data Acquisition to Enhance Machine Learning Performance 2024 VLDB 6.0083431e-05
5,877 BlinkML: Efficient Maximum Likelihood Estimation with Probabilistic Guarantees 2019 SIGMOD 5.9627218e-05
6,143 Efficient Construction of Approximate Ad-Hoc ML models Through Materialization and Reuse 2018 VLDB 5.8721471e-05
6,592 Tuple-oriented Compression for Large-scale Mini-batch Stochastic Gradient Descent 2019 SIGMOD 5.7421684e-05
6,742 Mitigating the Impedance Mismatch between Prediction Query Execution and Database Engine 2025 SIGMOD 5.6910432e-05
6,771 Causal Feature Selection for Algorithmic Fairness 2022 SIGMOD 5.6864972e-05
7,201 Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning 2023 VLDB 5.5871656e-05
7,867 Mind the Gap: Bridging Multi-Domain Query Workloads with EmptyHeaded 2017 VLDB 5.4365458e-05
8,493 SPRINTER: A Fast n-ary Join Query Processing Method for Complex OLAP Queries 2020 SIGMOD 5.3310013e-05
8,807 ApproxML: Efficient Approximate Ad-Hoc ML Models Through Materialization and Reuse 2019 VLDB 5.2728127e-05
9,251 Leveraging Similarity Joins for Signal Reconstruction 2018 VLDB 5.2056825e-05
10,446 Eliminating Redundant Feature Tests in Decision Tree and Random Forest Inference on SQL Predicates 2026 SIGMOD 4.9793485e-05
10,452 Factorized and Vectorized Execution: Optimizing Analytical and Semantic Queries over Relations 2026 SIGMOD 4.9793485e-05
10,653 InferF: Declarative Factorization of AI/ML Inferences over Joins 2026 SIGMOD 4.9793485e-05
10,739 Database Views as Explanations for Relational Deep Learning 2026 VLDB 4.9793485e-05
11,578 Relational Query Synthesis ⋈ Decision Tree Learning 2024 VLDB 4.9793485e-05
11,590 Enriching Relations with Additional Attributes for ER 2024 VLDB 4.9793485e-05
11,706 Regularized Pairwise Relationship based Analytics for Structured Data 2023 SIGMOD 4.9793485e-05
11,981 Enforcing Constraints for Machine Learning Systems via Declarative Feature Selection: An Experimental Study 2021 SIGMOD 4.9793485e-05
12,247 Learning Efficiently Over Heterogeneous Databases 2018 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 8 of 8 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers