Robust and Efficient Fuzzy Match for Online Data Cleaning
Summary: Proposes a novel similarity function addressing limitations of fuzzy-match metrics for data cleaning. Develops an efficient fuzzy-match algorithm for real-time validation/cleansing of incoming tuples against reference tables; demonstrated on real datasets. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Surajit Chaudhuri (Microsoft)
- 2. Kris Ganjam (Microsoft)
- 3. Venkatesh Ganti (Microsoft)
- 4. Rajeev Motwani (Stanford University)
BibTeX Citation
@inproceedings{chaudhuri_sigmod03,
title = {{Robust and Efficient Fuzzy Match for Online Data Cleaning}},
author = {Chaudhuri, Surajit and Ganjam, Kris and Ganti, Venkatesh and Motwani, Rajeev},
series = {{SIGMOD} '03},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/872757.872796},
url = {https://dl.acm.org/doi/10.1145/872757.872796},
year = {2003}
}
Incoming Citations (Sorted by Pagerank)
Showing 50 of 57 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 56 | M-tree: An Efficient Access Method for Similarity Search in Metric Spaces | 1997 | VLDB | 0.00040719947 |
| 58 | The Merge/Purge Problem for Large Databases | 1995 | SIGMOD | 0.00040116748 |
| 107 | Approximate String Joins in a Database (Almost) for Free | 2001 | VLDB | 0.00033511706 |
| 160 | Integration of Heterogeneous Databases Without Common Domains Using Queries Based on Textual Similarity | 1998 | SIGMOD | 0.0002802209 |
| 306 | Eliminating Fuzzy Duplicates in Data Warehouses | 2002 | VLDB | 0.00021839661 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 975 | Can We Beat the Prefix Filtering? An Adaptive Framework for Similarity Join and Search | 2012 | SIGMOD |
| 2 | 1,560 | Example-driven Design of Efficient Record Matching Queries | 2007 | VLDB |
| 3 | 11,504 | TokenJoin: Efficient Filtering for Set Similarity Join with Maximum Weighted Bipartite Matching | 2023 | VLDB |
| 4 | 7,719 | Data Cleaning in Microsoft SQL Server 2005 | 2005 | SIGMOD |
| 5 | 5,134 | Auto-FuzzyJoin: Auto-Program Fuzzy Similarity Joins Without Labeled Examples | 2021 | SIGMOD |
| 6 | 6,827 | Sampling Dirty Data for Matching Attributes | 2010 | SIGMOD |
| 7 | 1,248 | Entity Matching: How Similar Is Similar | 2011 | VLDB |
| 8 | 3,610 | Merging the Results of Approximate Match Operations | 2004 | VLDB |
| 9 | 4,009 | Flexible String Matching Against Large Databases in Practice | 2004 | VLDB |
| 10 | 306 | Eliminating Fuzzy Duplicates in Data Warehouses | 2002 | VLDB |