KATARA: Reliable Data Cleaning with Knowledge Bases and Crowdsourcing
Summary: Katara: end-to-end data cleaning powered by knowledge bases and crowdsourcing for reliable repairs. Interprets table semantics against a KB, flags correctness, and outputs top-k repairs; adds browser-based setup, pattern validation, data annotation, and repair-status visualization. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Xu Chu
- 2. Mourad Ouzzani
- 3. John Morcos
- 4. Paolo Papotti
- 5. Ihab F. Ilyas
- 6. Nan Tang
- 7. Yin Ye
Incoming Citations (Sorted by Pagerank)
Showing 8 of 8 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 5,150 | Horizon: Scalable Dependency-driven Data Cleaning | 2021 | VLDB | 5.6553571e-05 |
| 6,188 | Semi-Supervised Data Cleaning with Raha and Baran | 2021 | CIDR | 5.1607275e-05 |
| 6,279 | Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial Networks | 2023 | VLDB | 5.1241232e-05 |
| 7,233 | CleanM: An Optimizable Query Language for Unified Scale-Out Data Cleaning | 2017 | VLDB | 4.788267e-05 |
| 8,096 | Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications | 2023 | SIGMOD | 4.583522e-05 |
| 9,849 | Reptile: Aggregation-level Explanations for Hierarchical Data | 2022 | SIGMOD | 4.2680295e-05 |
| 10,826 | Demonstrating Matelda for Multi-Table Error Detection | 2025 | VLDB | 4.1905499e-05 |
| 11,738 | A Demonstration of PERC: Probabilistic Entity Resolution With Crowd Errors | 2018 | VLDB | 4.1905499e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 365 | Annotating and Searching Web Tables Using Entities, Types and Relationships | 2010 | VLDB | 0.00025616694 |
| 656 | ERACER: A Database Approach for Statistical Inference and Data Cleaning | 2010 | SIGMOD | 0.00018590675 |
| 830 | Guided Data Repair | 2011 | VLDB | 0.00016125759 |
| 879 | Don’t be SCAREd: Use SCalable Automatic REpairing with Maximal Likelihood and Bounded Changes | 2013 | SIGMOD | 0.00015649604 |
| 1,160 | Towards Certain Fixes with Editing Rules and Master Data | 2010 | VLDB | 0.0001358129 |
| 1,544 | KATARA: A Data Cleaning System Powered by Knowledge Bases and Crowdsourcing | 2015 | SIGMOD | 0.00011438274 |
| 2,830 | Interaction between Record Matching and Data Repairing | 2011 | SIGMOD | 8.0515409e-05 |
| 2,853 | Building, Maintaining, and Using Knowledge Bases: A Report from the Trenches | 2013 | SIGMOD | 8.0147968e-05 |
| 3,000 | BigDansing: A System for Big Data Cleansing | 2015 | SIGMOD | 7.7447724e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| Overall Rank | Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 198 | Declarative Data Cleaning: Language, Model, and Algorithms | 2001 | VLDB | 0.0003505869 |
| 4,414 | CrowdMatcher: Crowd-Assisted Schema Matching | 2014 | SIGMOD | 6.1975384e-05 |
| 7,565 | PIClean: A Probabilistic and Interactive Data Cleaning System | 2019 | SIGMOD | 4.7048523e-05 |
| 2,895 | Sato: Contextual Semantic Type Detection in Tables | 2020 | VLDB | 7.9539265e-05 |
| 9,283 | Interactive and Deterministic Data Cleaning: A Tossed Stone Raises a Thousand Ripples | 2016 | SIGMOD | 4.3598353e-05 |
| 488 | Data Curation at Scale: The Data Tamer System | 2013 | CIDR | 0.00022029993 |
| 10,521 | Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables | 2025 | SIGMOD | 4.1905499e-05 |
| 6,188 | Semi-Supervised Data Cleaning with Raha and Baran | 2021 | CIDR | 5.1607275e-05 |
| 10,826 | Demonstrating Matelda for Multi-Table Error Detection | 2025 | VLDB | 4.1905499e-05 |
| 1,544 | KATARA: A Data Cleaning System Powered by Knowledge Bases and Crowdsourcing | 2015 | SIGMOD | 0.00011438274 |