DBScholar

Back to papers

Raha: A Configuration-Free Error Detection System

Summary: Raha is a configuration-free error detection system for data cleaning. It generates a compact set of configurations to form per-tuple feature vectors, then uses sampling and learning to select representative values, leveraging historical data to prune irrelevant detectors and outperform prior work with at most 20 labels. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
he28533df965ac4bb
Venue
SIGMOD
Year
2019
Pagerank
9.59842e-05
Overall Rank
1,805 | 87.87%
DOI
10.1145/3299869.3324956

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{mahdavi_sigmod19,
        title = {{Raha: A Configuration-Free Error Detection System}},
        author = {Mahdavi, Mohammad and Abedjan, Ziawasch and Fernandez, Raul Castro and Madden, Samuel and Ouzzani, Mourad and Stonebraker, Michael and Tang, Nan},
        series = {{SIGMOD} '19},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3299869.3324956},
        url = {https://dl.acm.org/doi/10.1145/3299869.3324956},
        year = {2019}
}

Incoming Citations (Sorted by Pagerank)

Showing 37 of 37 citing papers.

Rank Citing Paper Year Venue Pagerank
1,342 Baran: Effective Error Correction via a Unified Context Representation and Transfer Learning 2020 VLDB 0.00010968223
1,991 RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation 2021 VLDB 9.2383849e-05
2,371 SCODED: Statistical Constraint Oriented Data Error Detection 2020 SIGMOD 8.5615698e-05
2,598 Complaint-driven Training Data Debugging for Query 2.0 2020 SIGMOD 8.2385793e-05
3,263 Automatic Data Repair: Are We Ready to Deploy? 2024 VLDB 7.4809771e-05
4,424 Auto-Transform: Learning-to-Transform by Patterns 2020 VLDB 6.6048065e-05
4,953 Adaptive Data Augmentation for Supervised Learning over Missing Data 2021 VLDB 6.3416889e-05
5,076 Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond 2021 SIGMOD 6.2860582e-05
5,247 Enabling SQL-based Training Data Debugging for Federated Learning 2022 VLDB 6.2121767e-05
5,513 Big Graphs: Challenges and Opportunities 2022 VLDB 6.0986544e-05
5,572 Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications 2023 SIGMOD 6.0802555e-05
5,643 Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial Networks 2023 VLDB 6.0538129e-05
5,954 Semi-Supervised Data Cleaning with Raha and Baran 2021 CIDR 5.9357922e-05
6,525 Parallel Discrepancy Detection and Incremental Detection 2021 VLDB 5.756805e-05
7,526 Auto-Prep: Holistic Prediction of Data Preparation Steps for Self-Service Business Intelligence 2025 VLDB 5.5020723e-05
7,738 Auto-Validate: Unsupervised Data Validation Using Data-Domain Patterns Inferred from Data Lakes 2021 SIGMOD 5.4623434e-05
7,850 CoClean: Collaborative Data Cleaning 2020 SIGMOD 5.4405793e-05
7,872 Rock: Cleaning Data by Embedding ML in Logic Rules 2024 SIGMOD 5.4362062e-05
8,224 Sparcle: Boosting the Accuracy of Data Cleaning Systems through Spatial Awareness 2024 VLDB 5.3747673e-05
8,756 Exploratory Training: When Annotators Learn About Data 2023 SIGMOD 5.2840289e-05
9,229 Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables 2025 SIGMOD 5.2056825e-05
9,311 VerifAI: Verified Generative AI 2024 CIDR 5.1965878e-05
9,608 Discovering Top-k Rules using Subjective and Objective Criteria 2023 SIGMOD 5.1527671e-05
9,616 GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language Models 2024 SIGMOD 5.1510548e-05
9,722 DataVinci: Learning Syntactic and Semantic String Repairs 2025 SIGMOD 5.1349531e-05
9,785 The Battleship Approach to the Low Resource Entity Matching Problem 2023 SIGMOD 5.1283279e-05
9,806 Making It Tractable to Catch Duplicates and Conflicts in Graphs 2023 SIGMOD 5.1257999e-05
10,112 Data Augmentation for ML-driven Data Preparation and Integration 2021 VLDB 5.0789354e-05
10,188 Reptile: Aggregation-level Explanations for Hierarchical Data 2022 SIGMOD 5.0651993e-05
10,529 Minimum Change ≠ Best Cleaning: Parallel and Incremental Error Detection under Integrity Constraints 2026 SIGMOD 4.9793485e-05
10,855 PipeLens: Identifying Interventions for Resolving Malfunctioning Data Science Pipelines 2026 VLDB 4.9793485e-05
11,003 ImputePilot: A Graphical Model Selection Toolkit for Time Series Imputation 2026 VLDB 4.9793485e-05
11,182 Data Enhancement for Binary Classification of Relational Data 2025 SIGMOD 4.9793485e-05
11,353 UniClean: A Scalable Data Cleaning Solution for Mixed Errors based on Unified Cleaners and Optimized Cleaning Workflow 2025 VLDB 4.9793485e-05
11,411 Demonstrating Matelda for Multi-Table Error Detection 2025 VLDB 4.9793485e-05
11,635 Rock: Cleaning Data with both ML and Logic Rules 2024 VLDB 4.9793485e-05
11,738 Splitting Tuples of Mismatched Entities 2023 SIGMOD 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 13 of 13 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers