DBScholar

Back to papers

CtxPipe: Context-aware Data Preparation Pipeline Construction for Machine Learning

Summary: CtxPipe automates context-aware data-prep pipeline construction for ML using pretrained embeddings to capture semantics and guide component choice. A deep RL framework searches the pipeline, delivering higher feature quality and faster models. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h0ad1e9cd7e0e1803
Venue
SIGMOD
Year
2024
Pagerank
5.6432616e-05
Overall Rank
6,921 | 53.47%
DOI
10.1145/3698831

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{gao_sigmod24,
        title = {{CtxPipe: Context-aware Data Preparation Pipeline Construction for Machine Learning}},
        author = {Gao, Haotian and Cai, Shaofeng and Dinh, Tien Tuan Anh and Huang, Zhiyong and Ooi, Beng Chin},
        series = {{SIGMOD} '24},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3698831},
        url = {https://dl.acm.org/doi/10.1145/3698831},
        year = {2024}
}

Incoming Citations (Sorted by Pagerank)

Showing 2 of 2 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 24 of 24 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
377 TURL: Table Understanding through Representation Learning 2021 VLDB 0.00019570264
483 ActiveClean: Interactive Data Cleaning For Statistical Modeling 2016 VLDB 0.00017590977
727 Functional Dependency Discovery: An Experimental Evaluation of Seven Algorithms 2015 VLDB 0.00014443394
975 Democratizing Data Science through Interactive Curation of ML Pipelines 2019 SIGMOD 0.00012750518
1,043 Data Cleaning: Overview and Emerging Challenges 2016 SIGMOD 0.00012335114
1,344 Detecting Data Errors: Where are we and what needs to be done? 2016 VLDB 0.00010956518
1,619 Efficient Denial Constraint Discovery with Hydra 2018 VLDB 0.00010055625
1,668 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.9371612e-05
1,847 Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions 2021 VLDB 9.5120573e-05
1,991 RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation 2021 VLDB 9.2383849e-05
2,187 Production Machine Learning Pipelines: Empirical Analysis and Optimization Opportunities 2021 SIGMOD 8.8896655e-05
2,443 Approximate Denial Constraints 2020 VLDB 8.4586656e-05
2,686 Data X-Ray: A Diagnostic Tool for Data Errors 2015 SIGMOD 8.1308928e-05
3,418 Data Profiling – A Tutorial 2017 SIGMOD 7.3210362e-05
4,226 Scalable Discovery of Unique Column Combinations 2014 VLDB 6.7183259e-05
4,524 DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular Data 2023 SIGMOD 6.5621504e-05
4,843 Pattern Functional Dependencies for Data Cleaning 2020 VLDB 6.3839159e-05
5,572 Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications 2023 SIGMOD 6.0802555e-05
5,848 HAIPipe: Combining Human-generated and Machine-generated Pipelines for Data Preparation 2023 SIGMOD 5.9720806e-05
6,171 Fundamentals of Order Dependencies 2012 VLDB 5.8623235e-05
7,183 DataPrism: Exposing Disconnect between Data and Systems 2022 SIGMOD 5.5917354e-05
7,505 Conformance Constraint Discovery: Measuring Trust in Data-Driven Systems 2021 SIGMOD 5.5074181e-05
7,626 BugDoc: Algorithms to Debug Computational Processes 2020 SIGMOD 5.4814815e-05
7,868 WindTunnel: Towards Differentiable ML Pipelines Beyond a Single Model 2022 VLDB 5.4365239e-05
Previous Page 1 / 1 Next

Semantically Similar Papers