DBScholar

Back to papers

Distributed Data Deduplication

Summary: Distributed data deduplication in a shared-nothing setting; leverages parallelism to prune pairwise comparisons after blocking. Dis-Dedup—a distribution strategy that minimizes the maximum per-node workload with theoretical guarantees; experiments on synthetic and real data show scalable speedups. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
11562
Venue
VLDB
Year
2016
Pagerank
7.983961e-05
Overall Rank
2,893 | 80.16%
DOI
10.14778/2983200.2983206

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{chu_vldb16,
        title = {{Distributed Data Deduplication}},
        author = {Chu, Xu and Ilyas, Ihab F. and Koutris, Paraschos},
        journal = {PVLDB},
        series = {{VLDB} '16},
        volume = {9},
        number = {11},
        pages = {864--875},
        doi = {10.14778/2983200.2983206},
        url = {https://doi.org/10.14778/2983200.2983206},
        year = {2016}
}

Incoming Citations (Sorted by Pagerank)

Showing 16 of 16 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 12 of 12 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers