GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete Data
Summary: GoodCore selects a coreset for incomplete data by modeling missingness as repairs over worlds and optimizing the expected subset without cleaning. It proves NP-hard and offers an approximation with imputation-based variants, enabling data-efficient ML. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Chengliang Chai (Beijing Institute of Technology)
- 2. Jiabin Liu (Beijing Institute of Technology)
- 3. Nan Tang (Hong Kong University of Science and Technology; Qatar Computing Research Institute)
- 4. Ju Fan (Renmin University of China)
- 5. Dongjing Miao (Harbin Engineering University)
- 6. Jiayi Wang (Tsinghua University)
- 7. Yuyu Luo (Tsinghua University)
- 8. Guoliang Li (Tsinghua University)
BibTeX Citation
@inproceedings{chai_sigmod23,
title = {{GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete Data}},
author = {Chai, Chengliang and Liu, Jiabin and Tang, Nan and Fan, Ju and Miao, Dongjing and Wang, Jiayi and Luo, Yuyu and Li, Guoliang},
series = {{SIGMOD} '23},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3589302},
url = {https://dl.acm.org/doi/10.1145/3589302},
year = {2023}
}
Incoming Citations (Sorted by Pagerank)
Showing 12 of 12 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 16 of 16 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 8,838 | The Cost of Representation by Subset Repairs | 2025 | VLDB |
| 2 | 5,167 | Enriching Data Imputation with Extensive Similarity Neighbors | 2015 | VLDB |
| 3 | 7,000 | A Cost-based Optimizer for Gradient Descent Optimization | 2017 | SIGMOD |
| 4 | 2,959 | Composable Core-sets for Diversity and Coverage Maximization | 2014 | PODS |
| 5 | 3,386 | Efficient and Effective Data Imputation with Influence Functions | 2022 | VLDB |
| 6 | 11,258 | Win-Win: On Simultaneous Clustering and Imputing over Incomplete Data | 2024 | VLDB |
| 7 | 2,147 | Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions | 2021 | VLDB |
| 8 | 11,169 | Certain and Approximately Certain Models for Statistical Learning | 2024 | SIGMOD |
| 9 | 11,104 | Datamap-Driven Tabular Coreset Selection for Classifier Training | 2025 | VLDB |
| 10 | 7,112 | Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning | 2023 | VLDB |