Framework for Evaluating Clustering Algorithms in Duplicate Detection
Summary: Stringer is an evaluation framework for scalable duplicate-detection clustering, combining approximate joins with unconstrained algorithms. Its experiments show previously overlooked clusterers can deliver strong accuracy and scalability. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Oktie Hassanzadeh (University of Toronto)
- 2. Fei Chiang (University of Toronto)
- 3. Hyun Chul Lee (Thoora Inc.)
- 4. Renée J. Miller (University of Toronto)
BibTeX Citation
@article{hassanzadeh_vldb09,
title = {{Framework for Evaluating Clustering Algorithms in Duplicate Detection}},
author = {Hassanzadeh, Oktie and Chiang, Fei and Lee, Hyun Chul and Miller, Renée J.},
journal = {PVLDB},
series = {{VLDB} '09},
doi = {10.14778/1687627.1687771},
url = {https://doi.org/10.14778/1687627.1687771},
year = {2009}
}
Incoming Citations (Sorted by Pagerank)
Showing 22 of 22 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 138 | Discovering Large Dense Subgraphs in Massive Graphs | 2005 | VLDB | 0.00029823423 |
| 169 | Efficient Exact Set-Similarity Joins | 2006 | VLDB | 0.0002743469 |
| 200 | Efficient set joins on similarity predicates | 2004 | SIGMOD | 0.00025597287 |
| 1,040 | VGRAM: Improving Performance of Approximate Queries on String Collections Using Variable-Length Grams | 2007 | VLDB | 0.00012466499 |
| 2,833 | Seeking Stable Clusters in the Blogosphere | 2007 | VLDB | 8.0782625e-05 |
| 3,356 | Benchmarking Declarative Approximate Selection Predicates | 2007 | SIGMOD | 7.490881e-05 |
| 4,657 | Finding Near Neighbors Through Cluster Pruning | 2007 | PODS | 6.5830079e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 3,916 | Incremental Record Linkage | 2014 | VLDB |
| 2 | 2,410 | Leveraging Aggregate Constraints For Deduplication | 2007 | SIGMOD |
| 3 | 6,723 | Record Linkage with Uniqueness Constraints and Erroneous Values | 2010 | VLDB |
| 4 | 10,878 | Evaluating Methods for Efficient Entity Count Estimation | 2025 | VLDB |
| 5 | 9,992 | Progressive Entity Matching: A Design Space Exploration | 2025 | SIGMOD |
| 6 | 3,983 | Crowd-Based Deduplication: An Adaptive Approach | 2015 | SIGMOD |
| 7 | 2,186 | String Similarity Joins: An Experimental Evaluation | 2014 | VLDB |
| 8 | 254 | Record Linkage: Similarity Measures and Algorithms | 2006 | SIGMOD |
| 9 | 306 | Eliminating Fuzzy Duplicates in Data Warehouses | 2002 | VLDB |
| 10 | 3,371 | Modeling and Querying Possible Repairs in Duplicate Detection | 2009 | VLDB |