DBScholar

Back to papers

Goods: Organizing Google's Datasets

Summary: Goods crawls diverse enterprise datasets to build a scalable metadata catalog and infer relationships (similarity, provenance) across billions in a distributed landscape. Provides discovery, monitoring, annotation, and relationship-analysis services for enterprise data at scale. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h7fdc70a31a3c5bc6
Venue
SIGMOD
Year
2016
Pagerank
0.00017063491
Overall Rank
509 | 96.59%
DOI
10.1145/2882903.2903730

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{halevy_sigmod16,
        title = {{Goods: Organizing Google's Datasets}},
        author = {Halevy, Alon and Korn, Flip and Noy, Natalya F. and Olston, Christopher and Polyzotis, Neoklis and Roy, Sudip and Whang, Steven Euijong},
        series = {{SIGMOD} '16},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/2882903.2903730},
        url = {https://dl.acm.org/doi/10.1145/2882903.2903730},
        year = {2016}
}

Incoming Citations (Sorted by Pagerank)

Showing 45 of 45 citing papers.

Rank Citing Paper Year Venue Pagerank
624 Data Lake Management: Challenges and Opportunities 2019 VLDB 0.00015476813
972 The Data Civilizer System 2017 CIDR 0.00012757732
1,038 ARDA: Automatic Relational Data Augmentation for Machine Learning 2020 VLDB 0.000123653
1,153 Data Management Challenges in Production Machine Learning 2017 SIGMOD 0.00011793347
1,293 Finding Related Tables in Data Lakes for Interactive Data Science 2020 SIGMOD 0.00011144991
1,308 Automating Large-Scale Data Quality Verification 2018 VLDB 0.00011073863
1,847 Data Market Platforms: Trading Data Assets to Solve Data Problems 2020 VLDB 9.5091361e-05
1,881 Ground: A Data Context Service 2017 CIDR 9.4372419e-05
2,160 Open Data Integration 2018 VLDB 8.9373643e-05
2,189 Production Machine Learning Pipelines: Empirical Analysis and Optimization Opportunities 2021 SIGMOD 8.8854572e-05
2,908 AI Meets Database: AI4DB and DB4AI 2021 SIGMOD 7.8716173e-05
2,927 Organizing Data Lakes for Navigation 2020 SIGMOD 7.8407133e-05
3,419 Data Profiling – A Tutorial 2017 SIGMOD 7.3175997e-05
3,521 Ember: No-Code Context Enrichment via Similarity-Based Keyless Joins 2022 VLDB 7.2320843e-05
3,544 Computation Reuse in Analytics Job Service at Microsoft 2018 SIGMOD 7.2108612e-05
3,913 Navigating the Data Lake with DATAMARAN: Automatically Extracting Structure from Log Datasets 2018 SIGMOD 6.9274273e-05
4,171 Data Integration and Machine Learning: A Natural Synergy 2018 SIGMOD 6.7576159e-05
4,334 LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems 2021 SIGMOD 6.6537801e-05
4,411 LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data Lakes 2024 VLDB 6.6089506e-05
4,675 Data Platform for Machine Learning 2019 SIGMOD 6.4728435e-05
4,804 Improving Reproducibility of Data Science Pipelines through Transparent Provenance Capture 2020 VLDB 6.4074452e-05
5,078 Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond 2021 SIGMOD 6.2832055e-05
5,109 Data-Driven Domain Discovery for Structured Datasets 2020 VLDB 6.2679672e-05
6,146 Efficient Construction of Approximate Ad-Hoc ML models Through Materialization and Reuse 2018 VLDB 5.869373e-05
6,410 Schemas and Types for JSON Data: from Theory to Practice 2019 SIGMOD 5.7915684e-05
6,979 Solo: Data Discovery Using Natural Language Questions Via A Self-Supervised Approach 2023 SIGMOD 5.6272426e-05
7,199 Cross Modal Data Discovery over Structured and Unstructured Data Lakes 2023 VLDB 5.5864204e-05
7,277 Computational Fact Checking: A Content Management Perspective 2018 VLDB 5.5655929e-05
7,418 Data Integration and Machine Learning: A Natural Synergy 2018 VLDB 5.5306468e-05
7,588 Unity Catalog: Open and Universal Governance for the Lakehouse and Beyond 2025 SIGMOD 5.4884845e-05
8,033 Dependency-Driven Analytics: a Compass for Uncharted Data Oceans 2017 CIDR 5.4017733e-05
8,407 Crossing the finish line faster when paddling the Data Lake with KAYAK 2017 VLDB 5.3365937e-05
8,590 OneProvenance: Efficient Extraction of Dynamic Coarse-Grained Provenance From Database Query Event Logs 2023 VLDB 5.3067446e-05
9,033 Fainder: A Fast and Accurate Index for Distribution-Aware Dataset Search 2024 VLDB 5.2322216e-05
10,344 QueryArtisan: Generating Data Manipulation Codes for Ad-hoc Analysis in Data Lakes 2025 VLDB 5.0176429e-05
10,903 MosaicJoin: Compact Semantic Sketches for Value-Level Join Discovery 2026 VLDB 4.9769913e-05
11,292 CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines 2025 VLDB 4.9769913e-05
11,303 OpenForge: Probabilistic Metadata Integration 2025 VLDB 4.9769913e-05
11,472 Towards an Objective Metric for Data Value Through Relevance 2024 CIDR 4.9769913e-05
11,604 Searching Data Lakes for Nested and Joined Data 2024 VLDB 4.9769913e-05
11,830 Kyrix-J: Visual Discovery of Connected Datasets in a Data Lake 2022 CIDR 4.9769913e-05
11,891 Fast Dataset Search with Earth Mover's Distance 2022 VLDB 4.9769913e-05
12,025 A Demonstration of RELIC: A System for REtrospective Lineage InferenCe of Data Workflows 2021 VLDB 4.9769913e-05
12,166 Ursprung: Provenance for Large-Scale Analytics Environments 2019 SIGMOD 4.9769913e-05
12,168 Peering through the Dark: An Owl's View of Inter-job Dependencies and Jobs' Impact in Shared Clusters 2019 SIGMOD 4.9769913e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 9 of 9 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers