DBScholar

Back to papers

Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning

Summary: Efficiently selects coresets for feature-augmented ML over multi-table 1-to-many/fuzzy joins without materializing the enormous join. Pushes gradient estimation to per-table partial similarities with provable bounds, achieving nearly 100× speedups at comparable accuracy. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
h85ecc67b2c908f2a
Venue
VLDB
Year
2023
Pagerank
5.5845207e-05
Overall Rank
7,203 | 51.59%
DOI
10.14778/3561261.3561267
PDF
Download (CC BY-NC-ND 4.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{wang_vldb23,
        title = {{Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning}},
        author = {Wang, Jiayi and Chai, Chengliang and Tang, Nan and Liu, Jiabin and Li, Guoliang},
        journal = {PVLDB},
        series = {{VLDB} '23},
        volume = {16},
        number = {1},
        pages = {64--76},
        doi = {10.14778/3561261.3561267},
        url = {https://doi.org/10.14778/3561261.3561267},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 8 of 8 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 25 of 25 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
15 How Good Are Query Optimizers, Really? 2016 VLDB 0.00061067652
105 The MADlib Analytics Library or MAD Skills, the SQL 2012 VLDB 0.00033633007
154 MAD Skills: New Analysis Practices for Big Data 2009 VLDB 0.00028568843
361 Bao: Making Learned Query Optimization Practical 2021 SIGMOD 0.00020000855
416 SystemML: Declarative Machine Learning on Spark 2016 VLDB 0.00018650998
510 NeuroCard: One Cardinality Estimator for All Tables 2021 VLDB 0.00017059914
521 Learning Linear Regression Models over Factorized Joins 2016 SIGMOD 0.00016923519
731 Learning Generalized Linear Models Over Normalized Data 2015 SIGMOD 0.00014400356
779 To Join or Not to Join? Thinking Twice about Joins before Feature Selection 2016 SIGMOD 0.00014048128
795 Random Sampling over Joins Revisited 2018 SIGMOD 0.00013934719
851 Scaling Factorization Machines to Relational Data 2013 VLDB 0.00013453278
1,038 ARDA: Automatic Relational Data Augmentation for Machine Learning 2020 VLDB 0.000123653
1,257 Towards Linear Algebra over Normalized Data 2017 VLDB 0.0001130959
2,201 Enabling and Optimizing Non-linear Feature Interactions in Factorized Linear Algebra 2019 SIGMOD 8.8708356e-05
3,373 F: Regression Models over Factorized Views 2016 VLDB 7.3627128e-05
3,673 Are Key-Foreign Key Joins Safe to Avoid when Learning High-Capacity Classifiers? 2018 VLDB 7.1075403e-05
3,742 FACE: A Normalizing Flow based Cardinality Estimator 2022 VLDB 7.0564546e-05
4,625 Cost-Effective Crowdsourced Entity Resolution: A Partial-Order Approach 2016 SIGMOD 6.4948389e-05
4,869 Selective Data Acquisition in the Wild for Model Charging 2022 VLDB 6.3719187e-05
5,683 Demonstration of Santoku: Optimizing Machine Learning over Normalized Data 2015 VLDB 6.0369434e-05
6,159 Automatic Data Acquisition for Deep Learning 2021 VLDB 5.8653048e-05
7,043 Human-in-the-loop Outlier Detection 2020 SIGMOD 5.611642e-05
8,806 Towards A Polyglot Framework for Factorized ML 2021 VLDB 5.2721353e-05
9,591 CDB: Optimizing Queries with Crowd-Based Selections and Joins 2017 SIGMOD 5.154741e-05
12,086 Interactively Discovering and Ranking Desired Tuples without Writing SQL Queries 2020 SIGMOD 4.9769913e-05
Previous Page 1 / 1 Next

Semantically Similar Papers