DBScholar

Back to papers

Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications

Summary: Saga automatically searches top data-cleaning pipelines for ML, combining AutoML, feature selection, and hyper-parameter tuning. Generates hybrid local/distributed runtime plans, extensible to new primitives, with monotonicity pruning and accuracy gains. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
6783
Venue
SIGMOD
Year
2023
Pagerank
5.6659017e-05
Overall Rank
7,232 | 50.39%
DOI
10.1145/3617338

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{siddiqi_sigmod23,
        title = {{Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications}},
        author = {Siddiqi, Shafaq and Kern, Roman and Boehm, Matthias},
        series = {{SIGMOD} '23},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3617338},
        url = {https://dl.acm.org/doi/10.1145/3617338},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 6 of 6 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 43 of 43 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
94 Potter's Wheel: An Interactive Data Cleaning System 2001 VLDB 0.00034616103
112 HoloClean: Holistic Data Repairs with Probabilistic Inference 2017 VLDB 0.00032801121
516 Data Curation at Scale: The Data Tamer System 2013 CIDR 0.00017171198
549 ERACER: A Database Approach for Statistical Inference and Data Cleaning 2010 SIGMOD 0.00016692839
582 ActiveClean: Interactive Data Cleaning For Statistical Modeling 2016 VLDB 0.00016148948
661 Don’t be SCAREd: Use SCalable Automatic REpairing with Maximal Likelihood and Bounded Changes 2013 SIGMOD 0.0001519162
714 Guided Data Repair 2011 VLDB 0.00014662041
725 NADEEF: A Commodity Data Cleaning System 2013 SIGMOD 0.00014617251
762 Model Management 2.0: Manipulating Richer Mappings 2007 SIGMOD 0.00014231163
946 HoloDetect: Few-Shot Learning for Error Detection 2019 SIGMOD 0.00013054126
963 The Data Civilizer System 2017 CIDR 0.00012935145
1,004 Democratizing Data Science through Interactive Curation of ML Pipelines 2019 SIGMOD 0.00012701932
1,079 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012258469
1,147 Data Management Challenges in Production Machine Learning 2017 SIGMOD 0.00011974846
1,157 Cerebro: A Data System for Optimized Deep Learning Model Selection 2020 VLDB 0.00011924049
1,350 Automating Large-Scale Data Quality Verification 2018 VLDB 0.00011065626
1,476 Baran: Effective Error Correction via a Unified Context Representation and Transfer Learning 2020 VLDB 0.00010659277
1,497 Generic Schema Matching, Ten Years Later 2011 VLDB 0.00010568859
1,569 HELIX: Holistic Optimization for Accelerating Iterative Machine Learning 2019 VLDB 0.00010335423
1,756 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.8172465e-05
2,029 Ease.ml: Towards Multi-tenant Resource Sharing for Machine Learning Workloads 2018 VLDB 9.2843642e-05
2,147 Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions 2021 VLDB 9.0831495e-05
2,208 Query Optimization for Dynamic Imputation 2017 VLDB 8.9512455e-05
2,273 SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model Debugging 2021 SIGMOD 8.8230899e-05
2,297 Raha: A Configuration-Free Error Detection System 2019 SIGMOD 8.7897221e-05
2,398 BigDansing: A System for Big Data Cleansing 2015 SIGMOD 8.631172e-05
2,893 Distributed Data Deduplication 2016 VLDB 7.983961e-05
2,995 Time Series Data Cleaning: From Anomaly Detection to Anomaly Repairing 2017 VLDB 7.8750141e-05
3,630 TensorFlow Data Validation: Data Analysis and Validation in Continuous ML Pipelines 2020 SIGMOD 7.2387749e-05
3,670 Learning to Validate the Predictions of Black Box Classifiers on Unseen Data 2020 SIGMOD 7.2118928e-05
3,883 Magellan: Toward Building Entity Matching Management Systems over Data Science Stacks 2016 VLDB 7.0487615e-05
3,905 Automated Feature Engineering for Algorithmic Fairness 2021 VLDB 7.029145e-05
4,240 LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems 2021 SIGMOD 6.809685e-05
4,382 Slice Tuner: A Selective Data Acquisition Framework for Accurate and Fair Machine Learning Models 2021 SIGMOD 6.7331832e-05
4,940 BEER: Blocking for Effective Entity Resolution 2021 SIGMOD 6.4336209e-05
5,319 xPAD: A Platform for Analytic Data Flows 2013 SIGMOD 6.2671857e-05
5,397 KATARA: Reliable Data Cleaning with Knowledge Bases and Crowdsourcing 2015 VLDB 6.232136e-05
5,785 BlinkML: Efficient Maximum Likelihood Estimation with Probabilistic Guarantees 2019 SIGMOD 6.0892672e-05
5,979 QoX-Driven ETL Design: Reducing the Cost of ETL Consulting Engagements 2009 SIGMOD 6.0212043e-05
6,801 Unit Testing Data with Deequ 2019 SIGMOD 5.7690726e-05
7,415 SystemER: A Human-in-the-loop System for Explainable Entity Resolution 2019 VLDB 5.6240334e-05
8,958 The Power of Nested Parallelism in Big Data Processing – Hitting Three Flies with One Slap – 2021 SIGMOD 5.3449654e-05
10,071 AlphaEvolve: A Learning Framework to Discover Novel Alphas in Quantitative Investment 2021 SIGMOD 5.1627323e-05
Previous Page 1 / 1 Next

Semantically Similar Papers