Discovering Related Data At Scale
Summary: ML-driven model to detect related columns across data streams from a month of lake queries. Scales to tens of millions of column-pairs and builds a data-relationship graph over 4.5 PB in ~80 minutes, with ~23% gains over state-of-the-art on labeled data. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Sagar Bharadwaj (Microsoft)
- 2. Praveen Gupta (Microsoft)
- 3. Ranjita Bhagwan (Microsoft)
- 4. Saikat Guha (Microsoft)
BibTeX Citation
@article{bharadwaj_vldb21,
title = {{Discovering Related Data At Scale}},
author = {Bharadwaj, Sagar and Gupta, Praveen and Bhagwan, Ranjita and Guha, Saikat},
journal = {PVLDB},
series = {{VLDB} '21},
volume = {14},
number = {8},
pages = {1392--1400},
doi = {10.14778/3457390.3457403},
url = {https://doi.org/10.14778/3457390.3457403},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 8 of 8 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,382 | DeepJoin: Joinable Table Discovery with Pre-trained Language Models | 2023 | VLDB | 8.5458532e-05 |
| 5,701 | Data Discovery in Data Lakes: Operations, Indexes, Systems | 2025 | VLDB | 6.0315322e-05 |
| 7,161 | WarpGate: A Semantic Join Discovery System for Cloud Data Warehouses | 2023 | CIDR | 5.596048e-05 |
| 8,574 | OmniMatch: Joinability Discovery in Data Products | 2025 | VLDB | 5.312217e-05 |
| 9,025 | Fainder: A Fast and Accurate Index for Distribution-Aware Dataset Search | 2024 | VLDB | 5.2346997e-05 |
| 9,242 | R2D2: Reducing Redundancy and Duplication in Data Lakes | 2023 | SIGMOD | 5.2056825e-05 |
| 10,798 | NeurIDA: Dynamic Modeling for Effective In-Database Analytics | 2026 | VLDB | 4.9793485e-05 |
| 11,081 | A Theoretical Framework for Distribution-Aware Dataset Search | 2025 | PODS | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 30 | SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets | 2008 | VLDB | 0.00050495102 |
| 108 | Approximate String Joins in a Database (Almost) for Free | 2001 | VLDB | 0.0003305531 |
| 161 | Robust and Efficient Fuzzy Match for Online Data Cleaning | 2003 | SIGMOD | 0.00027718195 |
| 624 | Data Lake Management: Challenges and Opportunities | 2019 | VLDB | 0.00015479826 |
| 694 | JOSIE: Overlap Set Similarity Search for Finding Joinable Tables in Data Lakes | 2019 | SIGMOD | 0.00014727089 |
| 972 | The Data Civilizer System | 2017 | CIDR | 0.00012763234 |
| 1,318 | LSH Ensemble: Internet-Scale Domain Search | 2016 | VLDB | 0.00011047393 |
| 4,274 | Set Similarity Joins on MapReduce: An Experimental Survey | 2018 | VLDB | 6.6918599e-05 |
| 4,491 | Juneau: Data Lake Management for Jupyter | 2019 | VLDB | 6.5778095e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 1,558 | Summarizing Relational Databases | 2009 | VLDB |
| 2 | 1,391 | Creating Embeddings of Heterogeneous Relational Datasets for Data Integration Tasks | 2020 | SIGMOD |
| 3 | 10,154 | Structure-Aware Machine Learning over Multi-Relational Databases | 2021 | SIGMOD |
| 4 | 7,342 | Discovering Association Rules from Big Graphs | 2022 | VLDB |
| 5 | 9,247 | Data Lakes Empowered by Knowledge Graph Technologies | 2021 | SIGMOD |
| 6 | 3,470 | Automatic Discovery of Attributes in Relational Databases | 2011 | SIGMOD |
| 7 | 7,197 | Cross Modal Data Discovery over Structured and Unstructured Data Lakes | 2023 | VLDB |
| 8 | 11,598 | Searching Data Lakes for Nested and Joined Data | 2024 | VLDB |
| 9 | 1,293 | Finding Related Tables in Data Lakes for Interactive Data Science | 2020 | SIGMOD |
| 10 | 717 | Finding Related Tables | 2012 | SIGMOD |