DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular Data
Summary: DiffPrep introduces differentiable bi-level search for preprocessing pipelines, enabling optimization in a continuous space. Continuous relaxation enables descent to find pipelines with a single training, yielding up to 6.6pp gains on 15/18 datasets. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Peng Li (Georgia Institute of Technology)
- 2. Zhiyi Chen (Georgia Institute of Technology)
- 3. Xu Chu (Georgia Institute of Technology)
- 4. Kexin Rong (Georgia Institute of Technology)
BibTeX Citation
@inproceedings{li_sigmod23,
title = {{DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular Data}},
author = {Li, Peng and Chen, Zhiyi and Chu, Xu and Rong, Kexin},
series = {{SIGMOD} '23},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3589328},
url = {https://dl.acm.org/doi/10.1145/3589328},
year = {2023}
}
Incoming Citations (Sorted by Pagerank)
Showing 5 of 5 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 6,921 | CtxPipe: Context-aware Data Preparation Pipeline Construction for Machine Learning | 2024 | SIGMOD | 5.6432616e-05 |
| 10,515 | Understanding the Impact of Data Noise in Federated Learning: [Experiments & Analysis] | 2026 | SIGMOD | 4.9793485e-05 |
| 10,990 | MAPPipe: A System for Bridging Efficiency and Quality in Data Preprocessing via Knowledge-Augmented Structural Pruning | 2026 | VLDB | 4.9793485e-05 |
| 11,061 | LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning | 2026 | VLDB | 4.9793485e-05 |
| 11,368 | Stress-Testing ML Pipelines with Adversarial Data Corruption | 2025 | VLDB | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 483 | ActiveClean: Interactive Data Cleaning For Statistical Modeling | 2016 | VLDB | 0.00017590977 |
| 975 | Democratizing Data Science through Interactive Curation of ML Pipelines | 2019 | SIGMOD | 0.00012750518 |
| 1,043 | Data Cleaning: Overview and Emerging Challenges | 2016 | SIGMOD | 0.00012335114 |
| 1,344 | Detecting Data Errors: Where are we and what needs to be done? | 2016 | VLDB | 0.00010956518 |
| 1,391 | Creating Embeddings of Heterogeneous Relational Datasets for Data Integration Tasks | 2020 | SIGMOD | 0.00010816237 |
| 1,847 | Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions | 2021 | VLDB | 9.5120573e-05 |
| 3,617 | Leva: Boosting Machine Learning Performance with Relational Embedding Data Augmentation | 2022 | SIGMOD | 7.1574349e-05 |
Previous
Page 1 / 1
Next