DBScholar

Back to papers

Goods: Organizing Google's Datasets

Summary: Goods crawls diverse enterprise datasets to build a scalable metadata catalog and infer relationships (similarity, provenance) across billions in a distributed landscape. Provides discovery, monitoring, annotation, and relationship-analysis services for enterprise data at scale. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h7fdc70a31a3c5bc6
Venue
SIGMOD
Year
2016
Pagerank
0.00017071087
Overall Rank
509 | 96.58%
DOI
10.1145/2882903.2903730

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{halevy_sigmod16,
        title = {{Goods: Organizing Google's Datasets}},
        author = {Halevy, Alon and Korn, Flip and Noy, Natalya F. and Olston, Christopher and Polyzotis, Neoklis and Roy, Sudip and Whang, Steven Euijong},
        series = {{SIGMOD} '16},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/2882903.2903730},
        url = {https://dl.acm.org/doi/10.1145/2882903.2903730},
        year = {2016}
}

Incoming Citations (Sorted by Pagerank)

Showing 45 of 45 citing papers.

Rank Citing Paper Year Venue Pagerank
624 Data Lake Management: Challenges and Opportunities 2019 VLDB 0.00015479826
972 The Data Civilizer System 2017 CIDR 0.00012763234
1,038 ARDA: Automatic Relational Data Augmentation for Machine Learning 2020 VLDB 0.00012370691
1,153 Data Management Challenges in Production Machine Learning 2017 SIGMOD 0.00011798912
1,293 Finding Related Tables in Data Lakes for Interactive Data Science 2020 SIGMOD 0.00011149857
1,308 Automating Large-Scale Data Quality Verification 2018 VLDB 0.0001107886
1,845 Data Market Platforms: Trading Data Assets to Solve Data Problems 2020 VLDB 9.5136397e-05
1,880 Ground: A Data Context Service 2017 CIDR 9.4415148e-05
2,158 Open Data Integration 2018 VLDB 8.941016e-05
2,187 Production Machine Learning Pipelines: Empirical Analysis and Optimization Opportunities 2021 SIGMOD 8.8896655e-05
2,908 AI Meets Database: AI4DB and DB4AI 2021 SIGMOD 7.8742664e-05
2,926 Organizing Data Lakes for Navigation 2020 SIGMOD 7.8442582e-05
3,418 Data Profiling – A Tutorial 2017 SIGMOD 7.3210362e-05
3,521 Ember: No-Code Context Enrichment via Similarity-Based Keyless Joins 2022 VLDB 7.2351481e-05
3,545 Computation Reuse in Analytics Job Service at Microsoft 2018 SIGMOD 7.2134803e-05
3,912 Navigating the Data Lake with DATAMARAN: Automatically Extracting Structure from Log Datasets 2018 SIGMOD 6.9306061e-05
4,171 Data Integration and Machine Learning: A Natural Synergy 2018 SIGMOD 6.7608137e-05
4,334 LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems 2021 SIGMOD 6.6569314e-05
4,409 LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data Lakes 2024 VLDB 6.6120807e-05
4,672 Data Platform for Machine Learning 2019 SIGMOD 6.4759004e-05
4,801 Improving Reproducibility of Data Science Pipelines through Transparent Provenance Capture 2020 VLDB 6.4104797e-05
5,076 Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond 2021 SIGMOD 6.2860582e-05
5,106 Data-Driven Domain Discovery for Structured Datasets 2020 VLDB 6.2706946e-05
6,143 Efficient Construction of Approximate Ad-Hoc ML models Through Materialization and Reuse 2018 VLDB 5.8721471e-05
6,407 Schemas and Types for JSON Data: from Theory to Practice 2019 SIGMOD 5.7942821e-05
6,993 Solo: Data Discovery Using Natural Language Questions Via A Self-Supervised Approach 2023 SIGMOD 5.626765e-05
7,197 Cross Modal Data Discovery over Structured and Unstructured Data Lakes 2023 VLDB 5.5890662e-05
7,274 Computational Fact Checking: A Content Management Perspective 2018 VLDB 5.5682289e-05
7,415 Data Integration and Machine Learning: A Natural Synergy 2018 VLDB 5.5332662e-05
7,582 Unity Catalog: Open and Universal Governance for the Lakehouse and Beyond 2025 SIGMOD 5.4910839e-05
8,027 Dependency-Driven Analytics: a Compass for Uncharted Data Oceans 2017 CIDR 5.4043169e-05
8,402 Crossing the finish line faster when paddling the Data Lake with KAYAK 2017 VLDB 5.3391207e-05
8,583 OneProvenance: Efficient Extraction of Dynamic Coarse-Grained Provenance From Database Query Event Logs 2023 VLDB 5.3092552e-05
9,025 Fainder: A Fast and Accurate Index for Distribution-Aware Dataset Search 2024 VLDB 5.2346997e-05
10,337 QueryArtisan: Generating Data Manipulation Codes for Ad-hoc Analysis in Data Lakes 2025 VLDB 5.0200193e-05
10,894 MosaicJoin: Compact Semantic Sketches for Value-Level Join Discovery 2026 VLDB 4.9793485e-05
11,284 CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines 2025 VLDB 4.9793485e-05
11,295 OpenForge: Probabilistic Metadata Integration 2025 VLDB 4.9793485e-05
11,466 Towards an Objective Metric for Data Value Through Relevance 2024 CIDR 4.9793485e-05
11,598 Searching Data Lakes for Nested and Joined Data 2024 VLDB 4.9793485e-05
11,824 Kyrix-J: Visual Discovery of Connected Datasets in a Data Lake 2022 CIDR 4.9793485e-05
11,885 Fast Dataset Search with Earth Mover's Distance 2022 VLDB 4.9793485e-05
12,019 A Demonstration of RELIC: A System for REtrospective Lineage InferenCe of Data Workflows 2021 VLDB 4.9793485e-05
12,160 Ursprung: Provenance for Large-Scale Analytics Environments 2019 SIGMOD 4.9793485e-05
12,162 Peering through the Dark: An Owl's View of Inter-job Dependencies and Jobs' Impact in Shared Clusters 2019 SIGMOD 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 9 of 9 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers