Towards Benchmarking Feature Type Inference for AutoML Platforms
Summary: First benchmark for ML-driven feature type inference in AutoML; presents a 9,921-sample, 9-class labeled dataset to standardize evaluation. ML-based typing yields 14% avg lift (up to 38%), beats industrial tools on 47/60 downstream models, and the dataset, models, and leaderboards are publicly released. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Vraj Shah (University of California San Diego)
- 2. Jonathan Lacanlale (California State University, Northridge)
- 3. Premanand Kumar (University of California San Diego)
- 4. Kevin Yang (University of California San Diego)
- 5. Arun Kumar (University of California San Diego)
BibTeX Citation
@inproceedings{shah_sigmod21,
title = {{Towards Benchmarking Feature Type Inference for AutoML Platforms}},
author = {Shah, Vraj and Lacanlale, Jonathan and Kumar, Premanand and Yang, Kevin and Kumar, Arun},
series = {{SIGMOD} '21},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3448016.3457274},
url = {https://dl.acm.org/doi/10.1145/3448016.3457274},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 9 of 9 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 4,534 | SimpleTS: An Efficient and Universal Model Selection Framework for Time Series Forecasting | 2023 | VLDB | 6.5561421e-05 |
| 5,097 | Explaining Dataset Changes for Semantic Data Versioning with Explain-Da-V | 2023 | VLDB | 6.2742361e-05 |
| 5,385 | SchemaPile: A Large Collection of Relational Database Schemas | 2024 | SIGMOD | 6.1523255e-05 |
| 6,662 | UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads | 2022 | VLDB | 5.7171651e-05 |
| 7,068 | How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses | 2024 | VLDB | 5.6083188e-05 |
| 7,910 | Pollock: A Data Loading Benchmark | 2023 | VLDB | 5.4284527e-05 |
| 8,544 | Automation of Data Prep, ML, and Data Science: New Cure or Snake Oil? | 2021 | SIGMOD | 5.3186267e-05 |
| 9,549 | GIO: Generating Efficient Matrix and Frame Readers for Custom Data Formats by Example | 2023 | SIGMOD | 5.1599622e-05 |
| 11,284 | CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines | 2025 | VLDB | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 205 | Snorkel: Rapid Training Data Creation with Weak Supervision | 2018 | VLDB | 0.00025181304 |
| 1,120 | Snuba: Automating Weak Supervision to Label Training Data | 2019 | VLDB | 0.00011946047 |
| 1,308 | Automating Large-Scale Data Quality Verification | 2018 | VLDB | 0.0001107886 |
| 2,745 | Transform-Data-by-Example (TDE): An Extensible Search Engine for Data Transformations | 2018 | VLDB | 8.0668395e-05 |
| 3,997 | Overton: A Data System for Monitoring and Improving Machine-Learned Products | 2020 | CIDR | 6.8655222e-05 |
| 4,030 | Slice Tuner: A Selective Data Acquisition Framework for Accurate and Fair Machine Learning Models | 2021 | SIGMOD | 6.8407272e-05 |
| 5,881 | ActiveClean: An Interactive Data Cleaning Framework For Modern Machine Learning | 2016 | SIGMOD | 5.9588636e-05 |
| 6,488 | Synthesizing Type-Detection Logic for Rich Semantic Data Types using Open-source Code | 2018 | SIGMOD | 5.76799e-05 |
| 8,025 | Foofah: A Programming-By-Example System for Synthesizing Data Transformation Programs | 2017 | SIGMOD | 5.4047151e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 5,921 | A Relational Framework for Classifier Engineering | 2017 | PODS |
| 2 | 11,720 | Steered Training Data Generation for Learned Semantic Type Detection | 2023 | SIGMOD |
| 3 | 9,229 | Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables | 2025 | SIGMOD |
| 4 | 6,028 | A Scalable AutoML Approach Based on Graph Neural Networks | 2022 | VLDB |
| 5 | 9,179 | FEBench: A Benchmark for Real-Time Relational Data Feature Extraction | 2023 | VLDB |
| 6 | 6,488 | Synthesizing Type-Detection Logic for Rich Semantic Data Types using Open-source Code | 2018 | SIGMOD |
| 7 | 10,235 | Scalable and Usable Relational Learning With Automatic Language Bias | 2021 | SIGMOD |
| 8 | 7,068 | How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses | 2024 | VLDB |
| 9 | 7,888 | MLBench: Benchmarking Machine Learning Services Against Human Experts | 2018 | VLDB |
| 10 | 4,589 | ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models | 2024 | VLDB |