DBScholar

Back to papers

Industry-Scale Duplicate Detection

Summary: Extends the DogmatiX duplicate-detection prototype from hierarchical XML to a 60M-person industrial credit database. Addresses both matching accuracy and scalability, with evaluation on real-world data where errors affect credit access and identity fraud. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
9942
Venue
VLDB
Year
2008
Pagerank
6.0977946e-05
Overall Rank
5,754 | 60.53%
DOI
10.14778/1454159.1454165

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{weis_vldb08,
        title = {{Industry-Scale Duplicate Detection}},
        author = {Weis, Melanie and Naumann, Felix and Jehle, Ulrich and Lufter, Jens and Schuster, Holger},
        journal = {PVLDB},
        series = {{VLDB} '08},
        pages = {1253},
        doi = {10.14778/1454159.1454165},
        url = {https://doi.org/10.14778/1454159.1454165},
        year = {2008}
}

Incoming Citations (Sorted by Pagerank)

Showing 2 of 2 citing papers.

Rank Citing Paper Year Venue Pagerank
619 Reasoning about Record Matching Rules 2009 VLDB 0.00015707247
7,880 Learning Over Dirty Data Without Cleaning 2020 SIGMOD 5.5244204e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 4 of 4 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
58 The Merge/Purge Problem for Large Databases 1995 SIGMOD 0.00040116748
306 Eliminating Fuzzy Duplicates in Data Warehouses 2002 VLDB 0.00021839661
1,560 Example-driven Design of Efficient Record Matching Queries 2007 VLDB 0.00010361222
2,531 DogmatiX Tracks down Duplicates in XML 2005 SIGMOD 8.4574631e-05
Previous Page 1 / 1 Next

Semantically Similar Papers