Sampling Dirty Data for Matching Attributes
Summary: Sampling dirty relational data to reveal overlapping string-value sets for joins. Proposes measures blending set-overlap and string-instance similarity, with distributed sampling and comparisons; adds a two-stage filter balancing accuracy and speed. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Henning Köhler (University of Queensland)
- 2. Xiaofang Zhou (National ICT Australia; University of Queensland)
- 3. Shazia Sadiq (University of Queensland)
- 4. Yanfeng Shu (Commonwealth Scientific and Industrial Research Organisation)
- 5. Kerry Taylor (Commonwealth Scientific and Industrial Research Organisation)
BibTeX Citation
@inproceedings{kohler_sigmod10,
title = {{Sampling Dirty Data for Matching Attributes}},
author = {Köhler, Henning and Zhou, Xiaofang and Sadiq, Shazia and Shu, Yanfeng and Taylor, Kerry},
series = {{SIGMOD} '10},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/1807167.1807177},
url = {https://dl.acm.org/doi/10.1145/1807167.1807177},
year = {2010}
}
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 6,349 | Robust Set Reconciliation | 2014 | SIGMOD | 5.8092399e-05 |
| 8,028 | Dscaler: Synthetically Scaling A Given Relational Database | 2016 | VLDB | 5.4041252e-05 |
| 11,408 | Mining Meaningful Keys and Foreign Keys with High Precision and Recall | 2025 | VLDB | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 14 of 14 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 2,921 | Leveraging Set Relations in Exact Set Similarity Join | 2017 | VLDB |
| 2 | 4,849 | Probabilistic String Similarity Joins | 2010 | SIGMOD |
| 3 | 175 | Simple Random Sampling from Relational Databases | 1986 | VLDB |
| 4 | 12,468 | Similarity Joins for Uncertain Strings | 2014 | SIGMOD |
| 5 | 3,593 | Merging the Results of Approximate Match Operations | 2004 | VLDB |
| 6 | 161 | Robust and Efficient Fuzzy Match for Online Data Cleaning | 2003 | SIGMOD |
| 7 | 1,720 | A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data | 2014 | SIGMOD |
| 8 | 9,878 | Towards a Unified Framework for String Similarity Joins | 2019 | VLDB |
| 9 | 2,220 | String Similarity Joins: An Experimental Evaluation | 2014 | VLDB |
| 10 | 4,087 | Flexible String Matching Against Large Databases in Practice | 2004 | VLDB |