DBScholar

Back to papers

Raha: A Configuration-Free Error Detection System

Summary: Raha is a configuration-free error detection system for data cleaning. It generates a compact set of configurations to form per-tuple feature vectors, then uses sampling and learning to select representative values, leveraging historical data to prune irrelevant detectors and outperform prior work with at most 20 labels. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
he28533df965ac4bb
Venue
SIGMOD
Year
2019
Pagerank
9.5938877e-05
Overall Rank
1,805 | 87.87%
DOI
10.1145/3299869.3324956

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{mahdavi_sigmod19,
        title = {{Raha: A Configuration-Free Error Detection System}},
        author = {Mahdavi, Mohammad and Abedjan, Ziawasch and Fernandez, Raul Castro and Madden, Samuel and Ouzzani, Mourad and Stonebraker, Michael and Tang, Nan},
        series = {{SIGMOD} '19},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3299869.3324956},
        url = {https://dl.acm.org/doi/10.1145/3299869.3324956},
        year = {2019}
}

Incoming Citations (Sorted by Pagerank)

Showing 37 of 37 citing papers.

Rank Citing Paper Year Venue Pagerank
1,342 Baran: Effective Error Correction via a Unified Context Representation and Transfer Learning 2020 VLDB 0.00010963254
1,993 RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation 2021 VLDB 9.2348951e-05
2,372 SCODED: Statistical Constraint Oriented Data Error Detection 2020 SIGMOD 8.5575299e-05
2,600 Complaint-driven Training Data Debugging for Query 2.0 2020 SIGMOD 8.2346824e-05
3,264 Automatic Data Repair: Are We Ready to Deploy? 2024 VLDB 7.4774357e-05
4,426 Auto-Transform: Learning-to-Transform by Patterns 2020 VLDB 6.601896e-05
4,955 Adaptive Data Augmentation for Supervised Learning over Missing Data 2021 VLDB 6.3386868e-05
5,078 Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond 2021 SIGMOD 6.2832055e-05
5,251 Enabling SQL-based Training Data Debugging for Federated Learning 2022 VLDB 6.2092359e-05
5,516 Big Graphs: Challenges and Opportunities 2022 VLDB 6.0957673e-05
5,574 Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications 2023 SIGMOD 6.0773771e-05
5,645 Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial Networks 2023 VLDB 6.0509471e-05
5,956 Semi-Supervised Data Cleaning with Raha and Baran 2021 CIDR 5.9329823e-05
6,528 Parallel Discrepancy Detection and Incremental Detection 2021 VLDB 5.7540798e-05
7,531 Auto-Prep: Holistic Prediction of Data Preparation Steps for Self-Service Business Intelligence 2025 VLDB 5.4994676e-05
7,742 Auto-Validate: Unsupervised Data Validation Using Data-Domain Patterns Inferred from Data Lakes 2021 SIGMOD 5.4599258e-05
7,854 CoClean: Collaborative Data Cleaning 2020 SIGMOD 5.4380042e-05
7,876 Rock: Cleaning Data by Embedding ML in Logic Rules 2024 SIGMOD 5.4336606e-05
8,230 Sparcle: Boosting the Accuracy of Data Cleaning Systems through Spatial Awareness 2024 VLDB 5.372223e-05
8,764 Exploratory Training: When Annotators Learn About Data 2023 SIGMOD 5.2815275e-05
9,239 Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables 2025 SIGMOD 5.2032182e-05
9,320 VerifAI: Verified Generative AI 2024 CIDR 5.1941278e-05
9,616 Discovering Top-k Rules using Subjective and Objective Criteria 2023 SIGMOD 5.1503279e-05
9,623 GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language Models 2024 SIGMOD 5.1486163e-05
9,727 DataVinci: Learning Syntactic and Semantic String Repairs 2025 SIGMOD 5.1325223e-05
9,791 The Battleship Approach to the Low Resource Entity Matching Problem 2023 SIGMOD 5.1259002e-05
9,813 Making It Tractable to Catch Duplicates and Conflicts in Graphs 2023 SIGMOD 5.1233734e-05
10,116 Data Augmentation for ML-driven Data Preparation and Integration 2021 VLDB 5.0765311e-05
10,191 Reptile: Aggregation-level Explanations for Hierarchical Data 2022 SIGMOD 5.0628015e-05
10,540 Minimum Change ≠ Best Cleaning: Parallel and Incremental Error Detection under Integrity Constraints 2026 SIGMOD 4.9769913e-05
10,864 PipeLens: Identifying Interventions for Resolving Malfunctioning Data Science Pipelines 2026 VLDB 4.9769913e-05
11,012 ImputePilot: A Graphical Model Selection Toolkit for Time Series Imputation 2026 VLDB 4.9769913e-05
11,191 Data Enhancement for Binary Classification of Relational Data 2025 SIGMOD 4.9769913e-05
11,360 UniClean: A Scalable Data Cleaning Solution for Mixed Errors based on Unified Cleaners and Optimized Cleaning Workflow 2025 VLDB 4.9769913e-05
11,417 Demonstrating Matelda for Multi-Table Error Detection 2025 VLDB 4.9769913e-05
11,641 Rock: Cleaning Data with both ML and Logic Rules 2024 VLDB 4.9769913e-05
11,744 Splitting Tuples of Mismatched Entities 2023 SIGMOD 4.9769913e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 13 of 13 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers