Sampling Dirty Data for Matching Attributes
Summary: Sampling dirty relational data to reveal overlapping string-value sets for joins. Proposes measures blending set-overlap and string-instance similarity, with distributed sampling and comparisons; adds a two-stage filter balancing accuracy and speed. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Henning Köhler (University of Queensland)
- 2. Xiaofang Zhou (National ICT Australia; University of Queensland)
- 3. Shazia Sadiq (University of Queensland)
- 4. Yanfeng Shu (Commonwealth Scientific and Industrial Research Organisation)
- 5. Kerry Taylor (Commonwealth Scientific and Industrial Research Organisation)
BibTeX Citation
@inproceedings{kohler_sigmod10,
title = {{Sampling Dirty Data for Matching Attributes}},
author = {Köhler, Henning and Zhou, Xiaofang and Sadiq, Shazia and Shu, Yanfeng and Taylor, Kerry},
series = {{SIGMOD} '10},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/1807167.1807177},
url = {https://dl.acm.org/doi/10.1145/1807167.1807177},
year = {2010}
}
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 6,220 | Robust Set Reconciliation | 2014 | SIGMOD | 5.9425753e-05 |
| 7,866 | Dscaler: Synthetically Scaling A Given Relational Database | 2016 | VLDB | 5.5281143e-05 |
| 11,044 | Mining Meaningful Keys and Foreign Keys with High Precision and Recall | 2025 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 14 of 14 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 3,040 | Leveraging Set Relations in Exact Set Similarity Join | 2017 | VLDB |
| 2 | 4,743 | Probabilistic String Similarity Joins | 2010 | SIGMOD |
| 3 | 173 | Simple Random Sampling from Relational Databases | 1986 | VLDB |
| 4 | 12,177 | Similarity Joins for Uncertain Strings | 2014 | SIGMOD |
| 5 | 3,610 | Merging the Results of Approximate Match Operations | 2004 | VLDB |
| 6 | 158 | Robust and Efficient Fuzzy Match for Online Data Cleaning | 2003 | SIGMOD |
| 7 | 1,736 | A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data | 2014 | SIGMOD |
| 8 | 9,702 | Towards a Unified Framework for String Similarity Joins | 2019 | VLDB |
| 9 | 2,186 | String Similarity Joins: An Experimental Evaluation | 2014 | VLDB |
| 10 | 4,009 | Flexible String Matching Against Large Databases in Practice | 2004 | VLDB |