DBScholar

Back to papers

CtxPipe: Context-aware Data Preparation Pipeline Construction for Machine Learning

Summary: CtxPipe automates context-aware data-prep pipeline construction for ML using pretrained embeddings to capture semantics and guide component choice. A deep RL framework searches the pipeline, delivering higher feature quality and faster models. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h0ad1e9cd7e0e1803
Venue
SIGMOD
Year
2024
Pagerank
5.6405901e-05
Overall Rank
6,924 | 53.47%
DOI
10.1145/3698831
PDF
Download (CC BY 4.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{gao_sigmod24,
        title = {{CtxPipe: Context-aware Data Preparation Pipeline Construction for Machine Learning}},
        author = {Gao, Haotian and Cai, Shaofeng and Dinh, Tien Tuan Anh and Huang, Zhiyong and Ooi, Beng Chin},
        series = {{SIGMOD} '24},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3698831},
        url = {https://dl.acm.org/doi/10.1145/3698831},
        year = {2024}
}

Incoming Citations (Sorted by Pagerank)

Showing 2 of 2 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 24 of 24 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
377 TURL: Table Understanding through Representation Learning 2021 VLDB 0.00019564011
483 ActiveClean: Interactive Data Cleaning For Statistical Modeling 2016 VLDB 0.00017584249
728 Functional Dependency Discovery: An Experimental Evaluation of Seven Algorithms 2015 VLDB 0.00014436728
976 Democratizing Data Science through Interactive Curation of ML Pipelines 2019 SIGMOD 0.0001274453
1,044 Data Cleaning: Overview and Emerging Challenges 2016 SIGMOD 0.00012329478
1,344 Detecting Data Errors: Where are we and what needs to be done? 2016 VLDB 0.00010951939
1,620 Efficient Denial Constraint Discovery with Hydra 2018 VLDB 0.00010050865
1,669 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.9324573e-05
1,848 Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions 2021 VLDB 9.5075544e-05
1,993 RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation 2021 VLDB 9.2348951e-05
2,189 Production Machine Learning Pipelines: Empirical Analysis and Optimization Opportunities 2021 SIGMOD 8.8854572e-05
2,445 Approximate Denial Constraints 2020 VLDB 8.4546614e-05
2,687 Data X-Ray: A Diagnostic Tool for Data Errors 2015 SIGMOD 8.1271333e-05
3,419 Data Profiling – A Tutorial 2017 SIGMOD 7.3175997e-05
4,226 Scalable Discovery of Unique Column Combinations 2014 VLDB 6.7151469e-05
4,525 DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular Data 2023 SIGMOD 6.559044e-05
4,845 Pattern Functional Dependencies for Data Cleaning 2020 VLDB 6.3808938e-05
5,574 Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications 2023 SIGMOD 6.0773771e-05
5,851 HAIPipe: Combining Human-generated and Machine-generated Pipelines for Data Preparation 2023 SIGMOD 5.9692535e-05
6,173 Fundamentals of Order Dependencies 2012 VLDB 5.8595496e-05
7,185 DataPrism: Exposing Disconnect between Data and Systems 2022 SIGMOD 5.5890883e-05
7,510 Conformance Constraint Discovery: Measuring Trust in Data-Driven Systems 2021 SIGMOD 5.504811e-05
7,632 BugDoc: Algorithms to Debug Computational Processes 2020 SIGMOD 5.4788866e-05
7,873 WindTunnel: Towards Differentiable ML Pipelines Beyond a Single Model 2022 VLDB 5.4339503e-05
Previous Page 1 / 1 Next

Semantically Similar Papers