Datamap-Driven Tabular Coreset Selection for Classifier Training
Summary: Uses Gradient-Boosted-Tree datamaps to select tabular training coresets in minutes, matching or surpassing full-data and baseline models. Datamap-guided inference enhancement offers guarantees under a stated property, plus explainability and coreset-size optimization. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Aviv Hadar (Tel Aviv University)
- 2. Tova Milo (Tel Aviv University)
- 3. Kathy Razmadze (Tel Aviv University)
BibTeX Citation
@article{hadar_vldb25,
title = {{Datamap-Driven Tabular Coreset Selection for Classifier Training}},
author = {Hadar, Aviv and Milo, Tova and Razmadze, Kathy},
journal = {PVLDB},
series = {{VLDB} '25},
volume = {18},
number = {3},
pages = {876--888},
doi = {10.14778/3712221.3712249},
url = {https://doi.org/10.14778/3712221.3712249},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 11,164 | Sentence to Model: Cost-Effective Data Collection LLM Agent | 2025 | SIGMOD | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 784 | VerdictDB: Universalizing Approximate Query Processing | 2018 | SIGMOD | 0.00014012614 |
| 931 | Dynamic Sample Selection for Approximate Query Processing | 2003 | SIGMOD | 0.00013011667 |
| 1,829 | DBEst: Revisiting Approximate Query Processing Engines with Machine Learning Models | 2019 | SIGMOD | 9.5510333e-05 |
| 2,414 | Structured Search Result Differentiation | 2009 | VLDB | 8.5081277e-05 |
| 7,201 | Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning | 2023 | VLDB | 5.5871656e-05 |
| 7,275 | SubStrat: A Subset-Based Optimization Strategy for Faster AutoML | 2023 | VLDB | 5.5679786e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 11,527 | Settling Time vs. Accuracy Tradeoffs for Clustering Big Data | 2024 | SIGMOD |
| 2 | 2,870 | BOAT—Optimistic Decision Tree Construction | 1999 | SIGMOD |
| 3 | 12,129 | Personal Insights for Altering Decisions of Tree-based Ensembles over Time | 2020 | VLDB |
| 4 | 2,031 | RainForest - A Framework for Fast Decision Tree Construction of Large Datasets | 1998 | VLDB |
| 5 | 13,175 | Decision Tables: Scalable Classification Exploring RDBMS Capabilities | 2000 | VLDB |
| 6 | 5,723 | An Experimental Evaluation of Large Scale GBDT Systems | 2019 | VLDB |
| 7 | 7,524 | DeltaBoost: Gradient Boosting Decision Trees with Efficient Machine Unlearning | 2023 | SIGMOD |
| 8 | 9,846 | DimBoost: Boosting Gradient Boosting Decision Tree to Higher Dimensions | 2018 | SIGMOD |
| 9 | 3,941 | GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete Data | 2023 | SIGMOD |
| 10 | 7,201 | Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning | 2023 | VLDB |