Incorporating String Transformations in Record Matching
Summary: Extends record matching by allowing user-defined string transformations (e.g., Robert/Bob) to define string similarity. Proposes an index-aware fuzzy-lookup framework that combines transformations with a base similarity, achieving better match quality and faster retrieval. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Arvind Arasu (Microsoft)
- 2. Surajit Chaudhuri (Microsoft)
- 3. Kris Ganjam (Microsoft)
- 4. Raghav Kaushik (Microsoft)
BibTeX Citation
@inproceedings{arasu_sigmod08,
title = {{Incorporating String Transformations in Record Matching}},
author = {Arasu, Arvind and Chaudhuri, Surajit and Ganjam, Kris and Kaushik, Raghav},
series = {{SIGMOD} '08},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/1376616.1376742},
url = {https://dl.acm.org/doi/10.1145/1376616.1376742},
year = {2008}
}
Incoming Citations (Sorted by Pagerank)
Showing 2 of 2 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 3,446 | Efficient Approximate Entity Extraction with Edit Distance Constraints | 2009 | SIGMOD | 7.4087786e-05 |
| 5,667 | Efficient Approximate Search on String Collections (Tutorial) | 2009 | VLDB | 6.1279762e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 4 of 4 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 21 | Similarity Search in High Dimensions via Hashing | 1999 | VLDB | 0.00056760516 |
| 169 | Efficient Exact Set-Similarity Joins | 2006 | VLDB | 0.0002743469 |
| 254 | Record Linkage: Similarity Measures and Algorithms | 2006 | SIGMOD | 0.00023199211 |
| 7,719 | Data Cleaning in Microsoft SQL Server 2005 | 2005 | SIGMOD | 5.55926e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 158 | Robust and Efficient Fuzzy Match for Online Data Cleaning | 2003 | SIGMOD |
| 2 | 9,702 | Towards a Unified Framework for String Similarity Joins | 2019 | VLDB |
| 3 | 4,743 | Probabilistic String Similarity Joins | 2010 | SIGMOD |
| 4 | 4,396 | Approximate String Joins with Abbreviations | 2018 | VLDB |
| 5 | 12,177 | Similarity Joins for Uncertain Strings | 2014 | SIGMOD |
| 6 | 1,560 | Example-driven Design of Efficient Record Matching Queries | 2007 | VLDB |
| 7 | 5,306 | On Indexing Error-Tolerant Set Containment | 2010 | SIGMOD |
| 8 | 5,088 | String Similarity Measures and Joins with Synonyms | 2013 | SIGMOD |
| 9 | 4,009 | Flexible String Matching Against Large Databases in Practice | 2004 | VLDB |
| 10 | 3,509 | Learning String Transformations From Examples | 2009 | VLDB |