Datamap-Driven Tabular Coreset Selection for Classifier Training
Summary: Uses Gradient-Boosted-Tree datamaps to select tabular training coresets in minutes, matching or surpassing full-data and baseline models. Datamap-guided inference enhancement offers guarantees under a stated property, plus explainability and coreset-size optimization. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Aviv Hadar (Tel Aviv University)
- 2. Tova Milo (Tel Aviv University)
- 3. Kathy Razmadze (Tel Aviv University)
BibTeX Citation
@article{hadar_vldb25,
title = {{Datamap-Driven Tabular Coreset Selection for Classifier Training}},
author = {Hadar, Aviv and Milo, Tova and Razmadze, Kathy},
journal = {PVLDB},
series = {{VLDB} '25},
volume = {18},
number = {3},
pages = {876--888},
doi = {10.14778/3712221.3712249},
url = {https://doi.org/10.14778/3712221.3712249},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 10,736 | Sentence to Model: Cost-Effective Data Collection LLM Agent | 2025 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 772 | VerdictDB: Universalizing Approximate Query Processing | 2018 | SIGMOD | 0.00014147905 |
| 909 | Dynamic Sample Selection for Approximate Query Processing | 2003 | SIGMOD | 0.00013291205 |
| 1,799 | DBEst: Revisiting Approximate Query Processing Engines with Machine Learning Models | 2019 | SIGMOD | 9.7326398e-05 |
| 2,379 | Structured Search Result Differentiation | 2009 | VLDB | 8.666326e-05 |
| 7,112 | Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning | 2023 | VLDB | 5.6990782e-05 |
| 7,127 | SubStrat: A Subset-Based Optimization Strategy for Faster AutoML | 2023 | VLDB | 5.6957765e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 11,184 | Settling Time vs. Accuracy Tradeoffs for Clustering Big Data | 2024 | SIGMOD |
| 2 | 2,810 | BOAT—Optimistic Decision Tree Construction | 1999 | SIGMOD |
| 3 | 11,828 | Personal Insights for Altering Decisions of Tree-based Ensembles over Time | 2020 | VLDB |
| 4 | 1,994 | RainForest - A Framework for Fast Decision Tree Construction of Large Datasets | 1998 | VLDB |
| 5 | 12,885 | Decision Tables: Scalable Classification Exploring RDBMS Capabilities | 2000 | VLDB |
| 6 | 5,589 | An Experimental Evaluation of Large Scale GBDT Systems | 2019 | VLDB |
| 7 | 7,381 | DeltaBoost: Gradient Boosting Decision Trees with Efficient Machine Unlearning | 2023 | SIGMOD |
| 8 | 9,672 | DimBoost: Boosting Gradient Boosting Decision Tree to Higher Dimensions | 2018 | SIGMOD |
| 9 | 3,886 | GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete Data | 2023 | SIGMOD |
| 10 | 7,112 | Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning | 2023 | VLDB |