DBScholar

Back to papers

How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses

Summary: First systematic empirical study of categorical duplicates (e.g., "CA" vs "California") on ML classification: labeled corpus of 1,262 categorical columns and a 16-dataset benchmark across five classifiers and five encoders. Finds logistic regression and similarity encoding robust to duplicates while one-hot with high-capacity models degrade; provides benchmarks and actionable takeaways for AutoML and data-prep. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
13571
Venue
VLDB
Year
2024
Pagerank
5.7370426e-05
Overall Rank
6,928 | 52.47%
DOI
10.14778/3648160.3648178

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{shah_vldb24,
        title = {{How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses}},
        author = {Shah, Vraj and Parashos, Thomas and Kumar, Arun},
        journal = {PVLDB},
        series = {{VLDB} '24},
        volume = {17},
        number = {6},
        pages = {1391--1404},
        doi = {10.14778/3648160.3648178},
        url = {https://doi.org/10.14778/3648160.3648178},
        year = {2024}
}

Incoming Citations (Sorted by Pagerank)

Showing 2 of 2 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 19 of 19 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
112 HoloClean: Holistic Data Repairs with Probabilistic Inference 2017 VLDB 0.00032801121
141 Deep Entity Matching with Pre-Trained Language Models 2021 VLDB 0.0002964847
176 Deep Learning for Entity Matching: A Design Space Exploration 2018 SIGMOD 0.00027191081
420 Can Foundation Models Wrangle Your Data? 2023 VLDB 0.00018789852
529 Magellan: Toward Building Entity Matching Management Systems 2016 VLDB 0.00017096361
1,323 Data Cleaning: Overview and Emerging Challenges 2016 SIGMOD 0.00011152602
1,465 Synthesizing Entity Matching Rules by Examples 2018 VLDB 0.00010689571
2,019 RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation 2021 VLDB 9.2983994e-05
2,290 ZeroER: Entity Resolution using Zero Labeled Examples 2020 SIGMOD 8.799251e-05
2,463 A Comprehensive Benchmark Framework for Active Learning Methods in Entity Matching 2020 SIGMOD 8.5486912e-05
2,978 Transform-Data-by-Example (TDE): An Extensible Search Engine for Data Transformations 2018 VLDB 7.9030989e-05
3,436 Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data Integration 2023 SIGMOD 7.4157897e-05
3,713 Generating Concise Entity Matching Rules 2017 SIGMOD 7.176496e-05
4,023 Smurf: Self-Service String Matching Using Random Forests 2019 VLDB 6.949387e-05
5,051 Towards Benchmarking Feature Type Inference for AutoML Platforms 2021 SIGMOD 6.385354e-05
5,134 Auto-FuzzyJoin: Auto-Program Fuzzy Similarity Joins Without Labeled Examples 2021 SIGMOD 6.3532024e-05
5,620 DataPrep.EDA: Task-Centric Exploratory Data Analysis for Statistical Modeling in Python 2021 SIGMOD 6.1482864e-05
5,794 ActiveClean: An Interactive Data Cleaning Framework For Modern Machine Learning 2016 SIGMOD 6.0863045e-05
7,869 Foofah: A Programming-By-Example System for Synthesizing Data Transformation Programs 2017 SIGMOD 5.5273054e-05
Previous Page 1 / 1 Next

Semantically Similar Papers