DBScholar

Back to papers

Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications

Summary: Saga automatically searches top data-cleaning pipelines for ML, combining AutoML, feature selection, and hyper-parameter tuning. Generates hybrid local/distributed runtime plans, extensible to new primitives, with monotonicity pruning and accuracy gains. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h08e1ff2061d765e1
Venue
SIGMOD
Year
2023
Pagerank
6.0802555e-05
Overall Rank
5,572 | 62.54%
DOI
10.1145/3617338

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{siddiqi_sigmod23,
        title = {{Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications}},
        author = {Siddiqi, Shafaq and Kern, Roman and Boehm, Matthias},
        series = {{SIGMOD} '23},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3617338},
        url = {https://dl.acm.org/doi/10.1145/3617338},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 8 of 8 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 43 of 43 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
95 Potter's Wheel: An Interactive Data Cleaning System 2001 VLDB 0.00034382643
104 HoloClean: Holistic Data Repairs with Probabilistic Inference 2017 VLDB 0.00033690989
483 ActiveClean: Interactive Data Cleaning For Statistical Modeling 2016 VLDB 0.00017590977
514 Data Curation at Scale: The Data Tamer System 2013 CIDR 0.00017006745
547 ERACER: A Database Approach for Statistical Inference and Data Cleaning 2010 SIGMOD 0.00016578131
652 Don’t be SCAREd: Use SCalable Automatic REpairing with Maximal Likelihood and Bounded Changes 2013 SIGMOD 0.00015121325
697 NADEEF: A Commodity Data Cleaning System 2013 SIGMOD 0.00014694048
716 Guided Data Repair 2011 VLDB 0.00014553463
793 Model Management 2.0: Manipulating Richer Mappings 2007 SIGMOD 0.00013942384
883 HoloDetect: Few-Shot Learning for Error Detection 2019 SIGMOD 0.00013268059
972 The Data Civilizer System 2017 CIDR 0.00012763234
975 Democratizing Data Science through Interactive Curation of ML Pipelines 2019 SIGMOD 0.00012750518
1,081 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012123917
1,152 Cerebro: A Data System for Optimized Deep Learning Model Selection 2020 VLDB 0.00011801961
1,153 Data Management Challenges in Production Machine Learning 2017 SIGMOD 0.00011798912
1,308 Automating Large-Scale Data Quality Verification 2018 VLDB 0.0001107886
1,342 Baran: Effective Error Correction via a Unified Context Representation and Transfer Learning 2020 VLDB 0.00010968223
1,505 Generic Schema Matching, Ten Years Later 2011 VLDB 0.00010454635
1,568 HELIX: Holistic Optimization for Accelerating Iterative Machine Learning 2019 VLDB 0.0001021302
1,668 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.9371612e-05
1,805 Raha: A Configuration-Free Error Detection System 2019 SIGMOD 9.59842e-05
1,847 Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions 2021 VLDB 9.5120573e-05
2,036 Ease.ml: Towards Multi-tenant Resource Sharing for Machine Learning Workloads 2018 VLDB 9.1520279e-05
2,239 Query Optimization for Dynamic Imputation 2017 VLDB 8.7745792e-05
2,326 SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model Debugging 2021 SIGMOD 8.6309237e-05
2,423 BigDansing: A System for Big Data Cleansing 2015 SIGMOD 8.4877894e-05
2,948 Distributed Data Deduplication 2016 VLDB 7.8230494e-05
2,971 Time Series Data Cleaning: From Anomaly Detection to Anomaly Repairing 2017 VLDB 7.7971821e-05
3,706 TensorFlow Data Validation: Data Analysis and Validation in Continuous ML Pipelines 2020 SIGMOD 7.0829596e-05
3,735 Learning to Validate the Predictions of Black Box Classifiers on Unseen Data 2020 SIGMOD 7.0642839e-05
3,745 Automated Feature Engineering for Algorithmic Fairness 2021 VLDB 7.055875e-05
3,843 Magellan: Toward Building Entity Matching Management Systems over Data Science Stacks 2016 VLDB 6.987246e-05
4,030 Slice Tuner: A Selective Data Acquisition Framework for Accurate and Fair Machine Learning Models 2021 SIGMOD 6.8407272e-05
4,334 LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems 2021 SIGMOD 6.6569314e-05
5,049 BEER: Blocking for Effective Entity Resolution 2021 SIGMOD 6.297901e-05
5,427 xPAD: A Platform for Analytic Data Flows 2013 SIGMOD 6.1340137e-05
5,512 KATARA: Reliable Data Cleaning with Knowledge Bases and Crowdsourcing 2015 VLDB 6.0991137e-05
5,877 BlinkML: Efficient Maximum Likelihood Estimation with Probabilistic Guarantees 2019 SIGMOD 5.9627218e-05
6,075 QoX-Driven ETL Design: Reducing the Cost of ETL Consulting Engagements 2009 SIGMOD 5.8944111e-05
6,898 Unit Testing Data with Deequ 2019 SIGMOD 5.6530554e-05
7,516 SystemER: A Human-in-the-loop System for Explainable Entity Resolution 2019 VLDB 5.504757e-05
9,034 The Power of Nested Parallelism in Big Data Processing – Hitting Three Flies with One Slap – 2021 SIGMOD 5.2334993e-05
10,240 AlphaEvolve: A Learning Framework to Discover Novel Alphas in Quantitative Investment 2021 SIGMOD 5.0534979e-05
Previous Page 1 / 1 Next

Semantically Similar Papers