DBScholar

Back to papers

How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses

Summary: First systematic empirical study of categorical duplicates (e.g., "CA" vs "California") on ML classification: labeled corpus of 1,262 categorical columns and a 16-dataset benchmark across five classifiers and five encoders. Finds logistic regression and similarity encoding robust to duplicates while one-hot with high-capacity models degrade; provides benchmarks and actionable takeaways for AutoML and data-prep. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
hca6419be9adcb4d3
Venue
VLDB
Year
2024
Pagerank
5.6056639e-05
Overall Rank
7,070 | 52.49%
DOI
10.14778/3648160.3648178
PDF
Download (CC BY-NC-ND 4.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{shah_vldb24,
        title = {{How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses}},
        author = {Shah, Vraj and Parashos, Thomas and Kumar, Arun},
        journal = {PVLDB},
        series = {{VLDB} '24},
        volume = {17},
        number = {6},
        pages = {1391--1404},
        doi = {10.14778/3648160.3648178},
        url = {https://doi.org/10.14778/3648160.3648178},
        year = {2024}
}

Incoming Citations (Sorted by Pagerank)

Showing 2 of 2 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 19 of 19 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
104 HoloClean: Holistic Data Repairs with Probabilistic Inference 2017 VLDB 0.00033676943
134 Deep Entity Matching with Pre-Trained Language Models 2021 VLDB 0.00030042569
158 Deep Learning for Entity Matching: A Design Space Exploration 2018 SIGMOD 0.00028038831
329 Can Foundation Models Wrangle Your Data? 2023 VLDB 0.00020867521
530 Magellan: Toward Building Entity Matching Management Systems 2016 VLDB 0.00016847532
1,044 Data Cleaning: Overview and Emerging Challenges 2016 SIGMOD 0.00012329478
1,493 Synthesizing Entity Matching Rules by Examples 2018 VLDB 0.00010500948
1,993 RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation 2021 VLDB 9.2348951e-05
2,325 ZeroER: Entity Resolution using Zero Labeled Examples 2020 SIGMOD 8.6308459e-05
2,475 A Comprehensive Benchmark Framework for Active Learning Methods in Entity Matching 2020 SIGMOD 8.4074448e-05
2,744 Transform-Data-by-Example (TDE): An Extensible Search Engine for Data Transformations 2018 VLDB 8.064168e-05
3,473 Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data Integration 2023 SIGMOD 7.2697311e-05
3,684 Generating Concise Entity Matching Rules 2017 SIGMOD 7.0983735e-05
4,109 Smurf: Self-Service String Matching Using Random Forests 2019 VLDB 6.7994518e-05
4,972 Auto-FuzzyJoin: Auto-Program Fuzzy Similarity Joins Without Labeled Examples 2021 SIGMOD 6.3300201e-05
5,102 Towards Benchmarking Feature Type Inference for AutoML Platforms 2021 SIGMOD 6.2705617e-05
5,642 DataPrep.EDA: Task-Centric Exploratory Data Analysis for Statistical Modeling in Python 2021 SIGMOD 6.0519777e-05
5,882 ActiveClean: An Interactive Data Cleaning Framework For Modern Machine Learning 2016 SIGMOD 5.9560519e-05
8,030 Foofah: A Programming-By-Example System for Synthesizing Data Transformation Programs 2017 SIGMOD 5.4021584e-05
Previous Page 1 / 1 Next

Semantically Similar Papers