Flexible String Matching Against Large Databases in Practice
Summary: Practical deployment of tf-idf/cosine fuzzy matching on large AT&T databases, extending it across multiple string attributes and semantic equivalences. Engineering optimizations trade slight accuracy loss for major speedups in real data-cleaning workloads. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Nick Koudas (AT&T)
- 2. Amit Marathe (AT&T)
- 3. Divesh Srivastava (AT&T)
BibTeX Citation
@article{koudas_vldb04,
title = {{Flexible String Matching Against Large Databases in Practice}},
author = {Koudas, Nick and Marathe, Amit and Srivastava, Divesh},
journal = {PVLDB},
series = {{VLDB} '04},
pages = {1078--1087},
doi = {10.1016/B978-012088469-8.50094-2},
url = {https://doi.org/10.1016/B978-012088469-8.50094-2},
year = {2004}
}
Incoming Citations (Sorted by Pagerank)
Showing 9 of 9 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,497 | Generic Schema Matching, Ten Years Later | 2011 | VLDB | 0.00010568859 |
| 1,560 | Example-driven Design of Efficient Record Matching Queries | 2007 | VLDB | 0.00010361222 |
| 3,294 | Multi-column Substring Matching for Database Schema Translation | 2006 | VLDB | 7.5491351e-05 |
| 3,356 | Benchmarking Declarative Approximate Selection Predicates | 2007 | SIGMOD | 7.490881e-05 |
| 4,359 | Incremental Maintenance of Length Normalized Indexes for Approximate String Matching | 2009 | SIGMOD | 6.7451753e-05 |
| 6,716 | SigMatch: Fast and Scalable Multi-Pattern Matching | 2010 | VLDB | 5.7958663e-05 |
| 7,692 | Efficient Top-k Algorithms for Approximate Substring Matching | 2013 | SIGMOD | 5.5664611e-05 |
| 12,737 | SPIDER: Flexible Matching in Databases | 2005 | SIGMOD | 5.093636e-05 |
| 13,812 | Using SPIDER: An Experience Report | 2006 | SIGMOD | - |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 1 of 1 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 158 | Robust and Efficient Fuzzy Match for Online Data Cleaning | 2003 | SIGMOD | 0.00028199923 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 13,812 | Using SPIDER: An Experience Report | 2006 | SIGMOD |
| 2 | 9,702 | Towards a Unified Framework for String Similarity Joins | 2019 | VLDB |
| 3 | 3,610 | Merging the Results of Approximate Match Operations | 2004 | VLDB |
| 4 | 107 | Approximate String Joins in a Database (Almost) for Free | 2001 | VLDB |
| 5 | 6,827 | Sampling Dirty Data for Matching Attributes | 2010 | SIGMOD |
| 6 | 12,177 | Similarity Joins for Uncertain Strings | 2014 | SIGMOD |
| 7 | 4,743 | Probabilistic String Similarity Joins | 2010 | SIGMOD |
| 8 | 7,714 | Incorporating String Transformations in Record Matching | 2008 | SIGMOD |
| 9 | 2,186 | String Similarity Joins: An Experimental Evaluation | 2014 | VLDB |
| 10 | 158 | Robust and Efficient Fuzzy Match for Online Data Cleaning | 2003 | SIGMOD |