Eliminating Fuzzy Duplicates in Data Warehouses
Summary: Eliminating fuzzy duplicates in data-warehouse dimensional tables by leveraging hierarchies to resolve domain-specific abbreviations. A scalable, high-quality duplicate-elimination algorithm that outperforms generic text-similarity, validated on real operational DW datasets. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Rohit Ananthakrishna (Cornell University)
- 2. Surajit Chaudhuri (Microsoft)
- 3. Venkatesh Ganti (Microsoft)
BibTeX Citation
@article{ananthakrishna_vldb02,
title = {{Eliminating Fuzzy Duplicates in Data Warehouses}},
author = {Ananthakrishna, Rohit and Chaudhuri, Surajit and Ganti, Venkatesh},
journal = {PVLDB},
series = {{VLDB} '02},
doi = {10.1016/B978-155860869-6/50058-5},
url = {https://doi.org/10.1016/B978-155860869-6/50058-5},
year = {2002}
}
Incoming Citations (Sorted by Pagerank)
Showing 37 of 37 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 8 of 8 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 60 | The Merge/Purge Problem for Large Databases | 1995 | SIGMOD | 0.000394583 |
| 95 | Potter's Wheel: An Interactive Data Cleaning System | 2001 | VLDB | 0.00034382643 |
| 108 | Approximate String Joins in a Database (Almost) for Free | 2001 | VLDB | 0.0003305531 |
| 162 | Integration of Heterogeneous Databases Without Common Domains Using Queries Based on Textual Similarity | 1998 | SIGMOD | 0.00027536748 |
| 204 | Declarative Data Cleaning: Language, Model, and Algorithms | 2001 | VLDB | 0.00025190386 |
| 296 | Generic Schema Matching with Cupid | 2001 | VLDB | 0.00021867512 |
| 735 | Automatic segmentation of text into structured records | 2001 | SIGMOD | 0.00014373044 |
| 1,749 | Clustering Categorical Data: An Approach Based on Dynamical Systems | 1998 | VLDB | 9.7329666e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 4,087 | Flexible String Matching Against Large Databases in Practice | 2004 | VLDB |
| 2 | 1,247 | Entity Matching: How Similar Is Similar | 2011 | VLDB |
| 3 | 4,051 | Crowd-Based Deduplication: An Adaptive Approach | 2015 | SIGMOD |
| 4 | 2,948 | Distributed Data Deduplication | 2016 | VLDB |
| 5 | 5,849 | Industry-Scale Duplicate Detection | 2008 | VLDB |
| 6 | 5,522 | MDedup: Duplicate Detection with Matching Dependencies | 2020 | VLDB |
| 7 | 7,878 | Data Cleaning in Microsoft SQL Server 2005 | 2005 | SIGMOD |
| 8 | 885 | Framework for Evaluating Clustering Algorithms in Duplicate Detection | 2009 | VLDB |
| 9 | 3,426 | Modeling and Querying Possible Repairs in Duplicate Detection | 2009 | VLDB |
| 10 | 161 | Robust and Efficient Fuzzy Match for Online Data Cleaning | 2003 | SIGMOD |