Discovering Related Data At Scale
Summary: ML-driven model to detect related columns across data streams from a month of lake queries. Scales to tens of millions of column-pairs and builds a data-relationship graph over 4.5 PB in ~80 minutes, with ~23% gains over state-of-the-art on labeled data. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Sagar Bharadwaj (Microsoft)
- 2. Praveen Gupta (Microsoft)
- 3. Ranjita Bhagwan (Microsoft)
- 4. Saikat Guha (Microsoft)
BibTeX Citation
@article{bharadwaj_vldb21,
title = {{Discovering Related Data At Scale}},
author = {Bharadwaj, Sagar and Gupta, Praveen and Bhagwan, Ranjita and Guha, Saikat},
journal = {PVLDB},
series = {{VLDB} '21},
volume = {14},
number = {8},
pages = {1392--1400},
doi = {10.14778/3457390.3457403},
url = {https://doi.org/10.14778/3457390.3457403},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 7 of 7 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 3,073 | DeepJoin: Joinable Table Discovery with Pre-trained Language Models | 2023 | VLDB | 7.785842e-05 |
| 8,144 | WarpGate: A Semantic Join Discovery System for Cloud Data Warehouses | 2023 | CIDR | 5.4800625e-05 |
| 9,065 | R2D2: Reducing Redundancy and Duplication in Data Lakes | 2023 | SIGMOD | 5.3251649e-05 |
| 10,080 | Fainder: A Fast and Accurate Index for Distribution-Aware Dataset Search | 2024 | VLDB | 5.158939e-05 |
| 10,638 | A Theoretical Framework for Distribution-Aware Dataset Search | 2025 | PODS | 5.093636e-05 |
| 10,987 | OmniMatch: Joinability Discovery in Data Products | 2025 | VLDB | 5.093636e-05 |
| 11,062 | Data Discovery in Data Lakes: Operations, Indexes, Systems | 2025 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 30 | SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets | 2008 | VLDB | 0.00051174276 |
| 107 | Approximate String Joins in a Database (Almost) for Free | 2001 | VLDB | 0.00033511706 |
| 158 | Robust and Efficient Fuzzy Match for Online Data Cleaning | 2003 | SIGMOD | 0.00028199923 |
| 749 | Data Lake Management: Challenges and Opportunities | 2019 | VLDB | 0.00014379989 |
| 779 | JOSIE: Overlap Set Similarity Search for Finding Joinable Tables in Data Lakes | 2019 | SIGMOD | 0.00014092047 |
| 963 | The Data Civilizer System | 2017 | CIDR | 0.00012935145 |
| 1,495 | LSH Ensemble: Internet-Scale Domain Search | 2016 | VLDB | 0.00010571481 |
| 4,260 | Set Similarity Joins on MapReduce: An Experimental Survey | 2018 | VLDB | 6.7984322e-05 |
| 4,511 | Juneau: Data Lake Management for Jupyter | 2019 | VLDB | 6.6517623e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 1,532 | Summarizing Relational Databases | 2009 | VLDB |
| 2 | 1,402 | Creating Embeddings of Heterogeneous Relational Datasets for Data Integration Tasks | 2020 | SIGMOD |
| 3 | 9,962 | Structure-Aware Machine Learning over Multi-Relational Databases | 2021 | SIGMOD |
| 4 | 7,330 | Discovering Association Rules from Big Graphs | 2022 | VLDB |
| 5 | 9,071 | Data Lakes Empowered by Knowledge Graph Technologies | 2021 | SIGMOD |
| 6 | 3,456 | Automatic Discovery of Attributes in Relational Databases | 2011 | SIGMOD |
| 7 | 7,572 | Cross Modal Data Discovery over Structured and Unstructured Data Lakes | 2023 | VLDB |
| 8 | 11,270 | Searching Data Lakes for Nested and Joined Data | 2024 | VLDB |
| 9 | 1,303 | Finding Related Tables in Data Lakes for Interactive Data Science | 2020 | SIGMOD |
| 10 | 739 | Finding Related Tables | 2012 | SIGMOD |