DBScholar

Back to papers

Learning Over Dirty Data Without Cleaning

Summary: DLearn learns directly from dirty data without cleaning, bypassing data-repair bottlenecks. It leverages database constraints to infer relational models that summarize patterns across all plausible clean versions; empirical evaluation on large real-world datasets shows accuracy and efficiency. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
5985
Venue
SIGMOD
Year
2020
Pagerank
5.5244204e-05
Overall Rank
7,880 | 45.94%
DOI
10.1145/3318464.3389708

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{picado_sigmod20,
        title = {{Learning Over Dirty Data Without Cleaning}},
        author = {Picado, Jose and Davis, John and Termehchy, Arash and Lee, Ga Young},
        series = {{SIGMOD} '20},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3318464.3389708},
        url = {https://dl.acm.org/doi/10.1145/3318464.3389708},
        year = {2020}
}

Incoming Citations (Sorted by Pagerank)

Showing 4 of 4 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 20 of 20 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
33 Consistent Query Answers in Inconsistent Databases 1999 PODS 0.00049907763
58 The Merge/Purge Problem for Large Databases 1995 SIGMOD 0.00040116748
112 HoloClean: Holistic Data Repairs with Probabilistic Inference 2017 VLDB 0.00032801121
185 A Cost-Based Model and Effective Heuristic for Repairing Constraints by Value Modification 2005 SIGMOD 0.00026231189
201 Declarative Data Cleaning: Language, Model, and Algorithms 2001 VLDB 0.00025558602
494 Dependencies Revisited for Improving Data Quality 2008 PODS 0.00017526549
533 Improving Data Quality: Consistency and Accuracy 2007 VLDB 0.0001705859
582 ActiveClean: Interactive Data Cleaning For Statistical Modeling 2016 VLDB 0.00016148948
619 Reasoning about Record Matching Rules 2009 VLDB 0.00015707247
714 Guided Data Repair 2011 VLDB 0.00014662041
1,033 On Generating Near-Optimal Tableaux for Conditional Functional Dependencies 2008 VLDB 0.00012529852
1,192 Query From Examples: An Iterative, Data-Driven Approach to Query Construction 2015 VLDB 0.00011740132
1,323 Data Cleaning: Overview and Emerging Challenges 2016 SIGMOD 0.00011152602
2,631 FastQRE: Fast Query Reverse Engineering 2018 SIGMOD 8.3231158e-05
2,639 Learning and Verifying Quantified Boolean Queries by Example 2013 PODS 8.3078378e-05
4,190 QuickFOIL: Scalable Inductive Logic Programming 2015 VLDB 6.8431849e-05
5,391 MDedup: Duplicate Detection with Matching Dependencies 2020 VLDB 6.2343031e-05
5,754 Industry-Scale Duplicate Detection 2008 VLDB 6.0977946e-05
7,985 Schema Independent Relational Learning 2017 SIGMOD 5.5125969e-05
8,676 Machine Learning for Data Management: Problems and Solutions 2018 SIGMOD 5.3867577e-05
Previous Page 1 / 1 Next

Semantically Similar Papers