Gobblin: Unifying Data Ingestion for Hadoop
Summary: Gobblin unifies Hadoop data ingestion into a single, extensible framework. Out-of-the-box support for relational, NoSQL, streaming, REST, and file sources; emphasizes generality, extensibility, operability, and end-to-end production metrics. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Lin Qiao (LinkedIn Inc.)
- 2. Yinan Li (LinkedIn Inc.)
- 3. Sahil Takiar (LinkedIn Inc.)
- 4. Ziyang Liu (LinkedIn Inc.)
- 5. Narasimha Veeramreddy (LinkedIn Inc.)
- 6. Min Tu (LinkedIn Inc.)
- 7. Ying Dai (LinkedIn Inc.)
- 8. Issac Buenrostro (LinkedIn Inc.)
- 9. Kapil Surlaker (LinkedIn Inc.)
- 10. Shirshanka Das (LinkedIn Inc.)
- 11. Chavdar Botev (LinkedIn Inc.)
BibTeX Citation
@article{qiao_vldb15,
title = {{Gobblin: Unifying Data Ingestion for Hadoop}},
author = {Qiao, Lin and Li, Yinan and Takiar, Sahil and Liu, Ziyang and Veeramreddy, Narasimha and Tu, Min and Dai, Ying and Buenrostro, Issac and Surlaker, Kapil and Das, Shirshanka and Botev, Chavdar},
journal = {PVLDB},
series = {{VLDB} '15},
volume = {8},
number = {12},
pages = {1764--1775},
doi = {10.14778/2824032.2824062},
url = {https://doi.org/10.14778/2824032.2824062},
year = {2015}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 12,006 | Query-able Kafka: An agile data analytics pipeline for mobile wireless networks | 2017 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 3 of 3 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 330 | Impala: A Modern, Open-Source SQL Engine for Hadoop | 2015 | CIDR | 0.0002104801 |
| 1,791 | Cheetah: A High Performance, Custom Data Warehouse on Top of MapReduce | 2010 | VLDB | 9.7470504e-05 |
| 1,999 | On Brewing Fresh Espresso: LinkedIn’s Distributed Data Serving Platform | 2013 | SIGMOD | 9.3349118e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 3,555 | Apache Hive: From MapReduce to Enterprise-grade Big Data Warehousing | 2019 | SIGMOD |
| 2 | 9,508 | An IDEA: An Ingestion Framework for Data Enrichment in AsterixDB | 2019 | VLDB |
| 3 | 2,642 | CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop | 2011 | VLDB |
| 4 | 6,004 | Data Ingestion for the Connected World | 2017 | CIDR |
| 5 | 6,734 | Liquid: Unifying Nearline and Offline Big Data Integration | 2015 | CIDR |
| 6 | 4,169 | The Unified Logging Infrastructure for Data Analytics at Twitter | 2012 | VLDB |
| 7 | 8,069 | Execution Primitives for Scalable Joins and Aggregations in Map Reduce | 2014 | VLDB |
| 8 | 32 | Hive - A Warehousing Solution Over a Map-Reduce Framework | 2009 | VLDB |
| 9 | 2,191 | Data Warehousing and Analytics Infrastructure at Facebook | 2010 | SIGMOD |
| 10 | 4,595 | The "Big Data" Ecosystem at LinkedIn | 2013 | SIGMOD |