DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular Data
Summary: DiffPrep introduces differentiable bi-level search for preprocessing pipelines, enabling optimization in a continuous space. Continuous relaxation enables descent to find pipelines with a single training, yielding up to 6.6pp gains on 15/18 datasets. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Peng Li (Georgia Institute of Technology)
- 2. Zhiyi Chen (Georgia Institute of Technology)
- 3. Xu Chu (Georgia Institute of Technology)
- 4. Kexin Rong (Georgia Institute of Technology)
BibTeX Citation
@inproceedings{li_sigmod23,
title = {{DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular Data}},
author = {Li, Peng and Chen, Zhiyi and Chu, Xu and Rong, Kexin},
series = {{SIGMOD} '23},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3589328},
url = {https://dl.acm.org/doi/10.1145/3589328},
year = {2023}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 8,914 | CtxPipe: Context-aware Data Preparation Pipeline Construction for Machine Learning | 2024 | SIGMOD | 5.3483178e-05 |
| 10,303 | Understanding the Impact of Data Noise in Federated Learning: [Experiments & Analysis] | 2026 | SIGMOD | 5.093636e-05 |
| 10,614 | LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning | 2026 | VLDB | 5.093636e-05 |
| 10,991 | Stress-Testing ML Pipelines with Adversarial Data Corruption | 2025 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 582 | ActiveClean: Interactive Data Cleaning For Statistical Modeling | 2016 | VLDB | 0.00016148948 |
| 1,004 | Democratizing Data Science through Interactive Curation of ML Pipelines | 2019 | SIGMOD | 0.00012701932 |
| 1,323 | Data Cleaning: Overview and Emerging Challenges | 2016 | SIGMOD | 0.00011152602 |
| 1,351 | Detecting Data Errors: Where are we and what needs to be done? | 2016 | VLDB | 0.00011064851 |
| 1,402 | Creating Embeddings of Heterogeneous Relational Datasets for Data Integration Tasks | 2020 | SIGMOD | 0.00010888094 |
| 2,147 | Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions | 2021 | VLDB | 9.0831495e-05 |
| 4,735 | Leva: Boosting Machine Learning Performance with Relational Embedding Data Augmentation | 2022 | SIGMOD | 6.5315782e-05 |
Previous
Page 1 / 1
Next