Towards Benchmarking Feature Type Inference for AutoML Platforms
Summary: First benchmark for ML-driven feature type inference in AutoML; presents a 9,921-sample, 9-class labeled dataset to standardize evaluation. ML-based typing yields 14% avg lift (up to 38%), beats industrial tools on 47/60 downstream models, and the dataset, models, and leaderboards are publicly released. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Vraj Shah (University of California San Diego)
- 2. Jonathan Lacanlale (California State University, Northridge)
- 3. Premanand Kumar (University of California San Diego)
- 4. Kevin Yang (University of California San Diego)
- 5. Arun Kumar (University of California San Diego)
BibTeX Citation
@inproceedings{shah_sigmod21,
title = {{Towards Benchmarking Feature Type Inference for AutoML Platforms}},
author = {Shah, Vraj and Lacanlale, Jonathan and Kumar, Premanand and Yang, Kevin and Kumar, Arun},
series = {{SIGMOD} '21},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3448016.3457274},
url = {https://dl.acm.org/doi/10.1145/3448016.3457274},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 9 of 9 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 4,990 | Explaining Dataset Changes for Semantic Data Versioning with Explain-Da-V | 2023 | VLDB | 6.4105738e-05 |
| 5,466 | SchemaPile: A Large Collection of Relational Database Schemas | 2024 | SIGMOD | 6.2075052e-05 |
| 5,498 | SimpleTS: An Efficient and Universal Model Selection Framework for Time Series Forecasting | 2023 | VLDB | 6.1972571e-05 |
| 6,538 | UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads | 2022 | VLDB | 5.8477764e-05 |
| 6,928 | How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses | 2024 | VLDB | 5.7370426e-05 |
| 7,803 | Pollock: A Data Loading Benchmark | 2023 | VLDB | 5.5418222e-05 |
| 8,373 | Automation of Data Prep, ML, and Data Science: New Cure or Snake Oil? | 2021 | SIGMOD | 5.4399999e-05 |
| 9,436 | GIO: Generating Efficient Matrix and Frame Readers for Custom Data Formats by Example | 2023 | SIGMOD | 5.2692207e-05 |
| 10,882 | CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines | 2025 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 205 | Snorkel: Rapid Training Data Creation with Weak Supervision | 2018 | VLDB | 0.00025235185 |
| 1,094 | Snuba: Automating Weak Supervision to Label Training Data | 2019 | VLDB | 0.00012214617 |
| 1,350 | Automating Large-Scale Data Quality Verification | 2018 | VLDB | 0.00011065626 |
| 2,978 | Transform-Data-by-Example (TDE): An Extensible Search Engine for Data Transformations | 2018 | VLDB | 7.9030989e-05 |
| 3,947 | Overton: A Data System for Monitoring and Improving Machine-Learned Products | 2020 | CIDR | 7.0040437e-05 |
| 4,382 | Slice Tuner: A Selective Data Acquisition Framework for Accurate and Fair Machine Learning Models | 2021 | SIGMOD | 6.7331832e-05 |
| 5,794 | ActiveClean: An Interactive Data Cleaning Framework For Modern Machine Learning | 2016 | SIGMOD | 6.0863045e-05 |
| 6,436 | Synthesizing Type-Detection Logic for Rich Semantic Data Types using Open-source Code | 2018 | SIGMOD | 5.8793793e-05 |
| 7,869 | Foofah: A Programming-By-Example System for Synthesizing Data Transformation Programs | 2017 | SIGMOD | 5.5273054e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 5,804 | A Relational Framework for Classifier Engineering | 2017 | PODS |
| 2 | 10,785 | Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables | 2025 | SIGMOD |
| 3 | 11,406 | Steered Training Data Generation for Learned Semantic Type Detection | 2023 | SIGMOD |
| 4 | 5,901 | A Scalable AutoML Approach Based on Graph Neural Networks | 2022 | VLDB |
| 5 | 9,433 | FEBench: A Benchmark for Real-Time Relational Data Feature Extraction | 2023 | VLDB |
| 6 | 6,436 | Synthesizing Type-Detection Logic for Rich Semantic Data Types using Open-source Code | 2018 | SIGMOD |
| 7 | 10,042 | Scalable and Usable Relational Learning With Automatic Language Bias | 2021 | SIGMOD |
| 8 | 6,928 | How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses | 2024 | VLDB |
| 9 | 7,763 | MLBench: Benchmarking Machine Learning Services Against Human Experts | 2018 | VLDB |
| 10 | 4,515 | ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models | 2024 | VLDB |