Eliminating Fuzzy Duplicates in Data Warehouses
Summary: Eliminating fuzzy duplicates in data-warehouse dimensional tables by leveraging hierarchies to resolve domain-specific abbreviations. A scalable, high-quality duplicate-elimination algorithm that outperforms generic text-similarity, validated on real operational DW datasets. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Rohit Ananthakrishna (Cornell University)
- 2. Surajit Chaudhuri (Microsoft)
- 3. Venkatesh Ganti (Microsoft)
BibTeX Citation
@article{ananthakrishna_vldb02,
title = {{Eliminating Fuzzy Duplicates in Data Warehouses}},
author = {Ananthakrishna, Rohit and Chaudhuri, Surajit and Ganti, Venkatesh},
journal = {PVLDB},
series = {{VLDB} '02},
doi = {10.1016/B978-155860869-6/50058-5},
url = {https://doi.org/10.1016/B978-155860869-6/50058-5},
year = {2002}
}
Incoming Citations (Sorted by Pagerank)
Showing 37 of 37 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 8 of 8 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 58 | The Merge/Purge Problem for Large Databases | 1995 | SIGMOD | 0.00040116748 |
| 94 | Potter's Wheel: An Interactive Data Cleaning System | 2001 | VLDB | 0.00034616103 |
| 107 | Approximate String Joins in a Database (Almost) for Free | 2001 | VLDB | 0.00033511706 |
| 160 | Integration of Heterogeneous Databases Without Common Domains Using Queries Based on Textual Similarity | 1998 | SIGMOD | 0.0002802209 |
| 201 | Declarative Data Cleaning: Language, Model, and Algorithms | 2001 | VLDB | 0.00025558602 |
| 297 | Generic Schema Matching with Cupid | 2001 | VLDB | 0.00022157284 |
| 717 | Automatic segmentation of text into structured records | 2001 | SIGMOD | 0.00014649121 |
| 1,713 | Clustering Categorical Data: An Approach Based on Dynamical Systems | 1998 | VLDB | 9.94811e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 4,009 | Flexible String Matching Against Large Databases in Practice | 2004 | VLDB |
| 2 | 1,248 | Entity Matching: How Similar Is Similar | 2011 | VLDB |
| 3 | 3,983 | Crowd-Based Deduplication: An Adaptive Approach | 2015 | SIGMOD |
| 4 | 2,893 | Distributed Data Deduplication | 2016 | VLDB |
| 5 | 5,754 | Industry-Scale Duplicate Detection | 2008 | VLDB |
| 6 | 5,391 | MDedup: Duplicate Detection with Matching Dependencies | 2020 | VLDB |
| 7 | 7,719 | Data Cleaning in Microsoft SQL Server 2005 | 2005 | SIGMOD |
| 8 | 871 | Framework for Evaluating Clustering Algorithms in Duplicate Detection | 2009 | VLDB |
| 9 | 3,371 | Modeling and Querying Possible Repairs in Duplicate Detection | 2009 | VLDB |
| 10 | 158 | Robust and Efficient Fuzzy Match for Online Data Cleaning | 2003 | SIGMOD |