Datamap-Driven Tabular Coreset Selection for Classifier Training
Summary: Uses Gradient-Boosted-Tree datamaps to select tabular training coresets in minutes, matching or surpassing full-data and baseline models. Datamap-guided inference enhancement offers guarantees under a stated property, plus explainability and coreset-size optimization. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Aviv Hadar (Tel Aviv University)
- 2. Tova Milo (Tel Aviv University)
- 3. Kathy Razmadze (Tel Aviv University)
BibTeX Citation
@article{hadar_vldb25,
title = {{Datamap-Driven Tabular Coreset Selection for Classifier Training}},
author = {Hadar, Aviv and Milo, Tova and Razmadze, Kathy},
journal = {PVLDB},
series = {{VLDB} '25},
volume = {18},
number = {3},
pages = {876--888},
doi = {10.14778/3712221.3712249},
url = {https://doi.org/10.14778/3712221.3712249},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 11,173 | Sentence to Model: Cost-Effective Data Collection LLM Agent | 2025 | SIGMOD | 4.9769913e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 772 | VerdictDB: Universalizing Approximate Query Processing | 2018 | SIGMOD | 0.0001409096 |
| 930 | Dynamic Sample Selection for Approximate Query Processing | 2003 | SIGMOD | 0.00013009255 |
| 1,828 | DBEst: Revisiting Approximate Query Processing Engines with Machine Learning Models | 2019 | SIGMOD | 9.547768e-05 |
| 2,415 | Structured Search Result Differentiation | 2009 | VLDB | 8.5041001e-05 |
| 7,203 | Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning | 2023 | VLDB | 5.5845207e-05 |
| 7,278 | SubStrat: A Subset-Based Optimization Strategy for Faster AutoML | 2023 | VLDB | 5.5653428e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 11,533 | Settling Time vs. Accuracy Tradeoffs for Clustering Big Data | 2024 | SIGMOD |
| 2 | 2,871 | BOAT—Optimistic Decision Tree Construction | 1999 | SIGMOD |
| 3 | 12,135 | Personal Insights for Altering Decisions of Tree-based Ensembles over Time | 2020 | VLDB |
| 4 | 2,034 | RainForest - A Framework for Fast Decision Tree Construction of Large Datasets | 1998 | VLDB |
| 5 | 13,181 | Decision Tables: Scalable Classification Exploring RDBMS Capabilities | 2000 | VLDB |
| 6 | 5,724 | An Experimental Evaluation of Large Scale GBDT Systems | 2019 | VLDB |
| 7 | 7,529 | DeltaBoost: Gradient Boosting Decision Trees with Efficient Machine Unlearning | 2023 | SIGMOD |
| 8 | 9,853 | DimBoost: Boosting Gradient Boosting Decision Tree to Higher Dimensions | 2018 | SIGMOD |
| 9 | 3,942 | GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete Data | 2023 | SIGMOD |
| 10 | 7,203 | Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning | 2023 | VLDB |