Framework for Evaluating Clustering Algorithms in Duplicate Detection
Summary: Stringer is an evaluation framework for scalable duplicate-detection clustering, combining approximate joins with unconstrained algorithms. Its experiments show previously overlooked clusterers can deliver strong accuracy and scalability. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Oktie Hassanzadeh (University of Toronto)
- 2. Fei Chiang (University of Toronto)
- 3. Hyun Chul Lee (Thoora Inc.)
- 4. Renée J. Miller (University of Toronto)
BibTeX Citation
@article{hassanzadeh_vldb09,
title = {{Framework for Evaluating Clustering Algorithms in Duplicate Detection}},
author = {Hassanzadeh, Oktie and Chiang, Fei and Lee, Hyun Chul and Miller, Renée J.},
journal = {PVLDB},
series = {{VLDB} '09},
doi = {10.14778/1687627.1687771},
url = {https://doi.org/10.14778/1687627.1687771},
year = {2009}
}
Incoming Citations (Sorted by Pagerank)
Showing 22 of 22 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 131 | Discovering Large Dense Subgraphs in Massive Graphs | 2005 | VLDB | 0.00030242586 |
| 168 | Efficient Exact Set-Similarity Joins | 2006 | VLDB | 0.00027163517 |
| 201 | Efficient set joins on similarity predicates | 2004 | SIGMOD | 0.00025331535 |
| 1,061 | VGRAM: Improving Performance of Approximate Queries on String Collections Using Variable-Length Grams | 2007 | VLDB | 0.00012213729 |
| 2,884 | Seeking Stable Clusters in the Blogosphere | 2007 | VLDB | 7.9108054e-05 |
| 3,403 | Benchmarking Declarative Approximate Selection Predicates | 2007 | SIGMOD | 7.3300953e-05 |
| 4,728 | Finding Near Neighbors Through Cluster Pruning | 2007 | PODS | 6.4473669e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 3,812 | Incremental Record Linkage | 2014 | VLDB |
| 2 | 2,461 | Leveraging Aggregate Constraints For Deduplication | 2007 | SIGMOD |
| 3 | 6,849 | Record Linkage with Uniqueness Constraints and Erroneous Values | 2010 | VLDB |
| 4 | 11,280 | Evaluating Methods for Efficient Entity Count Estimation | 2025 | VLDB |
| 5 | 10,180 | Progressive Entity Matching: A Design Space Exploration | 2025 | SIGMOD |
| 6 | 4,051 | Crowd-Based Deduplication: An Adaptive Approach | 2015 | SIGMOD |
| 7 | 2,220 | String Similarity Joins: An Experimental Evaluation | 2014 | VLDB |
| 8 | 262 | Record Linkage: Similarity Measures and Algorithms | 2006 | SIGMOD |
| 9 | 307 | Eliminating Fuzzy Duplicates in Data Warehouses | 2002 | VLDB |
| 10 | 3,426 | Modeling and Querying Possible Repairs in Duplicate Detection | 2009 | VLDB |