Framework for Evaluating Clustering Algorithms in Duplicate Detection
Summary: Stringer is an evaluation framework for scalable duplicate-detection clustering, combining approximate joins with unconstrained algorithms. Its experiments show previously overlooked clusterers can deliver strong accuracy and scalability. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Oktie Hassanzadeh (University of Toronto)
- 2. Fei Chiang (University of Toronto)
- 3. Hyun Chul Lee (Thoora Inc.)
- 4. Renée J. Miller (University of Toronto)
BibTeX Citation
@article{hassanzadeh_vldb09,
title = {{Framework for Evaluating Clustering Algorithms in Duplicate Detection}},
author = {Hassanzadeh, Oktie and Chiang, Fei and Lee, Hyun Chul and Miller, Renée J.},
journal = {PVLDB},
series = {{VLDB} '09},
doi = {10.14778/1687627.1687771},
url = {https://doi.org/10.14778/1687627.1687771},
year = {2009}
}
Incoming Citations (Sorted by Pagerank)
Showing 22 of 22 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 131 | Discovering Large Dense Subgraphs in Massive Graphs | 2005 | VLDB | 0.00030236369 |
| 168 | Efficient Exact Set-Similarity Joins | 2006 | VLDB | 0.00027151132 |
| 201 | Efficient set joins on similarity predicates | 2004 | SIGMOD | 0.00025319937 |
| 1,062 | VGRAM: Improving Performance of Approximate Queries on String Collections Using Variable-Length Grams | 2007 | VLDB | 0.00012208031 |
| 2,885 | Seeking Stable Clusters in the Blogosphere | 2007 | VLDB | 7.9072958e-05 |
| 3,403 | Benchmarking Declarative Approximate Selection Predicates | 2007 | SIGMOD | 7.3266624e-05 |
| 4,730 | Finding Near Neighbors Through Cluster Pruning | 2007 | PODS | 6.4444432e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 3,814 | Incremental Record Linkage | 2014 | VLDB |
| 2 | 2,461 | Leveraging Aggregate Constraints For Deduplication | 2007 | SIGMOD |
| 3 | 6,853 | Record Linkage with Uniqueness Constraints and Erroneous Values | 2010 | VLDB |
| 4 | 11,288 | Evaluating Methods for Efficient Entity Count Estimation | 2025 | VLDB |
| 5 | 10,184 | Progressive Entity Matching: A Design Space Exploration | 2025 | SIGMOD |
| 6 | 4,053 | Crowd-Based Deduplication: An Adaptive Approach | 2015 | SIGMOD |
| 7 | 2,223 | String Similarity Joins: An Experimental Evaluation | 2014 | VLDB |
| 8 | 262 | Record Linkage: Similarity Measures and Algorithms | 2006 | SIGMOD |
| 9 | 307 | Eliminating Fuzzy Duplicates in Data Warehouses | 2002 | VLDB |
| 10 | 3,426 | Modeling and Querying Possible Repairs in Duplicate Detection | 2009 | VLDB |