R2D2: Reducing Redundancy and Duplication in Data Lakes
Summary: R2D2 tackles table-level containment in data lakes with a three-stage pipeline: schema containment graph, min-max pruning, and content-level pruning—for scalable detection. It trims storage and access costs by deleting redundant datasets and reconstructing on demand under latency bounds; built on Spark (Azure Databricks/ADLS Gen2, AWS) for TB-scale lakes. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Raunak Shah (University of Illinois Urbana-Champaign)
- 2. Koyel Mukherjee (Adobe)
- 3. Atharv Tyagi (Adobe)
- 4. Sai Keerthana Karnam (Adobe)
- 5. Dhruv Joshi (Indian Institute of Technology Kharagpur)
- 6. Shivam Bhosale (Adobe)
- 7. Subrata Mitra (Adobe)
BibTeX Citation
@inproceedings{shah_sigmod23,
title = {{R2D2: Reducing Redundancy and Duplication in Data Lakes}},
author = {Shah, Raunak and Mukherjee, Koyel and Tyagi, Atharv and Karnam, Sai Keerthana and Joshi, Dhruv and Bhosale, Shivam and Mitra, Subrata},
series = {{SIGMOD} '23},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3626762},
url = {https://dl.acm.org/doi/10.1145/3626762},
year = {2023}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 11,090 | T-Assess: An Efficient Data Quality Assessment System Tailored for Trajectory Data | 2025 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 15 of 15 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 5,021 | Online Deduplication for Databases | 2017 | SIGMOD |
| 2 | 306 | Eliminating Fuzzy Duplicates in Data Warehouses | 2002 | VLDB |
| 3 | 5,983 | Adaptive and Robust Query Execution for Lakehouses at Scale | 2024 | VLDB |
| 4 | 2,410 | Leveraging Aggregate Constraints For Deduplication | 2007 | SIGMOD |
| 5 | 9,378 | AutoComp: Automated Data Compaction for Log-Structured Tables in Data Lakes | 2025 | SIGMOD |
| 6 | 1,350 | Automating Large-Scale Data Quality Verification | 2018 | VLDB |
| 7 | 9,550 | Data Imputation with Limited Data Redundancy Using Data Lakes | 2025 | VLDB |
| 8 | 2,893 | Distributed Data Deduplication | 2016 | VLDB |
| 9 | 7,780 | Petabyte-Scale Row-Level Operations in Data Lakehouses | 2024 | VLDB |
| 10 | 6,708 | Serving Deep Learning Models with Deduplication from Relational Databases | 2022 | VLDB |