DBScholar

Back to papers

GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete Data

Summary: GoodCore selects a coreset for incomplete data by modeling missingness as repairs over worlds and optimizing the expected subset without cleaning. It proves NP-hard and offers an approximation with imputation-based variants, enabling data-efficient ML. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
6722
Venue
SIGMOD
Year
2023
Pagerank
7.0460597e-05
Overall Rank
3,886 | 73.34%
DOI
10.1145/3589302

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{chai_sigmod23,
        title = {{GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete Data}},
        author = {Chai, Chengliang and Liu, Jiabin and Tang, Nan and Fan, Ju and Miao, Dongjing and Wang, Jiayi and Luo, Yuyu and Li, Guoliang},
        series = {{SIGMOD} '23},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3589302},
        url = {https://dl.acm.org/doi/10.1145/3589302},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 12 of 12 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 16 of 16 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
33 Consistent Query Answers in Inconsistent Databases 1999 PODS 0.00049907763
549 ERACER: A Database Approach for Statistical Inference and Data Cleaning 2010 SIGMOD 0.00016692839
582 ActiveClean: Interactive Data Cleaning For Statistical Modeling 2016 VLDB 0.00016148948
2,147 Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions 2021 VLDB 9.0831495e-05
2,565 Database Repairs and Consistent Query Answering: Origins and Further Developments 2019 PODS 8.4116912e-05
3,386 Efficient and Effective Data Imputation with Influence Functions 2022 VLDB 7.453827e-05
4,538 Cost-Effective Crowdsourced Entity Resolution: A Partial-Order Approach 2016 SIGMOD 6.6404776e-05
4,808 Selective Data Acquisition in the Wild for Model Charging 2022 VLDB 6.4991553e-05
4,809 Synthesizing Natural Language to Visualization (NL2VIS) Benchmarks from NL2SQL Benchmarks 2021 SIGMOD 6.4984309e-05
4,835 Adaptive Data Augmentation for Supervised Learning over Missing Data 2021 VLDB 6.486592e-05
5,211 CDB: A Crowd-Powered Database System 2018 VLDB 6.3161892e-05
6,074 Automatic Data Acquisition for Deep Learning 2021 VLDB 5.9860133e-05
6,899 Human-in-the-loop Outlier Detection 2020 SIGMOD 5.7431609e-05
7,112 Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning 2023 VLDB 5.6990782e-05
9,298 VisClean: Interactive Cleaning for Progressive Visualization 2020 VLDB 5.2901384e-05
11,778 Interactively Discovering and Ranking Desired Tuples without Writing SQL Queries 2020 SIGMOD 5.093636e-05
Previous Page 1 / 1 Next

Semantically Similar Papers