DBScholar

Back to papers

Learning Over Dirty Data Without Cleaning

Summary: DLearn learns directly from dirty data without cleaning, bypassing data-repair bottlenecks. It leverages database constraints to infer relational models that summarize patterns across all plausible clean versions; empirical evaluation on large real-world datasets shows accuracy and efficiency. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h4094dc02503a4092
Venue
SIGMOD
Year
2020
Pagerank
5.4330096e-05
Overall Rank
7,880 | 47.04%
DOI
10.1145/3318464.3389708

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{picado_sigmod20,
        title = {{Learning Over Dirty Data Without Cleaning}},
        author = {Picado, Jose and Davis, John and Termehchy, Arash and Lee, Ga Young},
        series = {{SIGMOD} '20},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3318464.3389708},
        url = {https://dl.acm.org/doi/10.1145/3318464.3389708},
        year = {2020}
}

Incoming Citations (Sorted by Pagerank)

Showing 4 of 4 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 20 of 20 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
33 Consistent Query Answers in Inconsistent Databases 1999 PODS 0.00049326063
60 The Merge/Purge Problem for Large Databases 1995 SIGMOD 0.00039440583
104 HoloClean: Holistic Data Repairs with Probabilistic Inference 2017 VLDB 0.00033676943
188 A Cost-Based Model and Effective Heuristic for Repairing Constraints by Value Modification 2005 SIGMOD 0.00025861123
204 Declarative Data Cleaning: Language, Model, and Algorithms 2001 VLDB 0.00025179068
483 ActiveClean: Interactive Data Cleaning For Statistical Modeling 2016 VLDB 0.00017584249
501 Dependencies Revisited for Improving Data Quality 2008 PODS 0.00017273167
526 Improving Data Quality: Consistency and Accuracy 2007 VLDB 0.00016879
618 Reasoning about Record Matching Rules 2009 VLDB 0.0001553004
716 Guided Data Repair 2011 VLDB 0.00014546968
1,031 On Generating Near-Optimal Tableaux for Conditional Functional Dependencies 2008 VLDB 0.00012399252
1,044 Data Cleaning: Overview and Emerging Challenges 2016 SIGMOD 0.00012329478
1,211 Query From Examples: An Iterative, Data-Driven Approach to Query Construction 2015 VLDB 0.00011522028
2,679 FastQRE: Fast Query Reverse Engineering 2018 SIGMOD 8.1415698e-05
2,685 Learning and Verifying Quantified Boolean Queries by Example 2013 PODS 8.1334309e-05
4,279 QuickFOIL: Scalable Inductive Logic Programming 2015 VLDB 6.6873645e-05
5,525 MDedup: Duplicate Detection with Matching Dependencies 2020 VLDB 6.0927371e-05
5,852 Industry-Scale Duplicate Detection 2008 VLDB 5.9691287e-05
8,157 Schema Independent Relational Learning 2017 SIGMOD 5.387426e-05
8,848 Machine Learning for Data Management: Problems and Solutions 2018 SIGMOD 5.2642362e-05
Previous Page 1 / 1 Next

Semantically Similar Papers