Towards Benchmarking Feature Type Inference for AutoML Platforms
Summary: First benchmark for ML-driven feature type inference in AutoML; presents a 9,921-sample, 9-class labeled dataset to standardize evaluation. ML-based typing yields 14% avg lift (up to 38%), beats industrial tools on 47/60 downstream models, and the dataset, models, and leaderboards are publicly released. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Vraj Shah (University of California San Diego)
- 2. Jonathan Lacanlale (California State University, Northridge)
- 3. Premanand Kumar (University of California San Diego)
- 4. Kevin Yang (University of California San Diego)
- 5. Arun Kumar (University of California San Diego)
BibTeX Citation
@inproceedings{shah_sigmod21,
title = {{Towards Benchmarking Feature Type Inference for AutoML Platforms}},
author = {Shah, Vraj and Lacanlale, Jonathan and Kumar, Premanand and Yang, Kevin and Kumar, Arun},
series = {{SIGMOD} '21},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3448016.3457274},
url = {https://dl.acm.org/doi/10.1145/3448016.3457274},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 9 of 9 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 4,535 | SimpleTS: An Efficient and Universal Model Selection Framework for Time Series Forecasting | 2023 | VLDB | 6.5530385e-05 |
| 5,100 | Explaining Dataset Changes for Semantic Data Versioning with Explain-Da-V | 2023 | VLDB | 6.271266e-05 |
| 5,393 | SchemaPile: A Large Collection of Relational Database Schemas | 2024 | SIGMOD | 6.1494131e-05 |
| 6,666 | UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads | 2022 | VLDB | 5.7144587e-05 |
| 7,070 | How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses | 2024 | VLDB | 5.6056639e-05 |
| 7,912 | Pollock: A Data Loading Benchmark | 2023 | VLDB | 5.4260784e-05 |
| 8,551 | Automation of Data Prep, ML, and Data Science: New Cure or Snake Oil? | 2021 | SIGMOD | 5.3161212e-05 |
| 9,545 | GIO: Generating Efficient Matrix and Frame Readers for Custom Data Formats by Example | 2023 | SIGMOD | 5.1600923e-05 |
| 11,292 | CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines | 2025 | VLDB | 4.9769913e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 205 | Snorkel: Rapid Training Data Creation with Weak Supervision | 2018 | VLDB | 0.00025171314 |
| 1,120 | Snuba: Automating Weak Supervision to Label Training Data | 2019 | VLDB | 0.000119406 |
| 1,308 | Automating Large-Scale Data Quality Verification | 2018 | VLDB | 0.00011073863 |
| 2,744 | Transform-Data-by-Example (TDE): An Extensible Search Engine for Data Transformations | 2018 | VLDB | 8.064168e-05 |
| 3,998 | Overton: A Data System for Monitoring and Improving Machine-Learned Products | 2020 | CIDR | 6.862274e-05 |
| 4,031 | Slice Tuner: A Selective Data Acquisition Framework for Accurate and Fair Machine Learning Models | 2021 | SIGMOD | 6.8374909e-05 |
| 5,882 | ActiveClean: An Interactive Data Cleaning Framework For Modern Machine Learning | 2016 | SIGMOD | 5.9560519e-05 |
| 6,490 | Synthesizing Type-Detection Logic for Rich Semantic Data Types using Open-source Code | 2018 | SIGMOD | 5.7653145e-05 |
| 8,030 | Foofah: A Programming-By-Example System for Synthesizing Data Transformation Programs | 2017 | SIGMOD | 5.4021584e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 5,922 | A Relational Framework for Classifier Engineering | 2017 | PODS |
| 2 | 11,726 | Steered Training Data Generation for Learned Semantic Type Detection | 2023 | SIGMOD |
| 3 | 9,239 | Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables | 2025 | SIGMOD |
| 4 | 6,031 | A Scalable AutoML Approach Based on Graph Neural Networks | 2022 | VLDB |
| 5 | 6,490 | Synthesizing Type-Detection Logic for Rich Semantic Data Types using Open-source Code | 2018 | SIGMOD |
| 6 | 9,189 | FEBench: A Benchmark for Real-Time Relational Data Feature Extraction | 2023 | VLDB |
| 7 | 10,241 | Scalable and Usable Relational Learning With Automatic Language Bias | 2021 | SIGMOD |
| 8 | 7,070 | How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses | 2024 | VLDB |
| 9 | 7,894 | MLBench: Benchmarking Machine Learning Services Against Human Experts | 2018 | VLDB |
| 10 | 4,591 | ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models | 2024 | VLDB |