The "Big Data" Ecosystem at LinkedIn
Summary: LinkedIn's Hadoop-based analytics stack abstracts distributed systems for data scientists, enabling end-to-end analytics on massive data. Novelty: seamless last-mile integration—ingress/egress to online systems and production workflows—via a 1-line Pig command, with use cases in recommendations and feeds. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Roshan Sumbaly (LinkedIn)
- 2. Jay Kreps (LinkedIn)
- 3. Sam Shah (LinkedIn)
BibTeX Citation
@inproceedings{sumbaly_sigmod13,
title = {{The "Big Data" Ecosystem at LinkedIn}},
author = {Sumbaly, Roshan and Kreps, Jay and Shah, Sam},
series = {{SIGMOD} '13},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2463676.2463707},
url = {https://dl.acm.org/doi/10.1145/2463676.2463707},
year = {2013}
}
Incoming Citations (Sorted by Pagerank)
Showing 6 of 6 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,568 | HELIX: Holistic Optimization for Accelerating Iterative Machine Learning | 2019 | VLDB | 0.00010208225 |
| 6,248 | The Cosmos Big Data Platform at Microsoft: Over a Decade of Progress and a Decade to Look Forward | 2021 | VLDB | 5.8345657e-05 |
| 8,184 | Data Management for Social Networking | 2016 | PODS | 5.3804156e-05 |
| 8,381 | Meta-Dataflows: Efficient Exploratory Dataflow Jobs | 2018 | SIGMOD | 5.3411788e-05 |
| 9,557 | Redoop Infrastructure for Recurring Big Data Queries | 2014 | VLDB | 5.1577635e-05 |
| 12,453 | Shared Execution of Recurring Workloads in MapReduce | 2015 | VLDB | 4.9769913e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 10 of 10 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 6 | Pig Latin: A Not-So-Foreign Language for Data Processing | 2008 | SIGMOD | 0.0010515896 |
| 13 | Mining Association Rules between Sets of Items in Large Databases | 1993 | SIGMOD | 0.00064391979 |
| 1,302 | Apache Hadoop Goes Realtime at Facebook | 2011 | SIGMOD | 0.00011106173 |
| 2,231 | Data Warehousing and Analytics Infrastructure at Facebook | 2010 | SIGMOD | 8.7958881e-05 |
| 3,641 | Efficient Bulk Insertion into a Distributed Ordered Table | 2008 | SIGMOD | 7.141045e-05 |
| 3,790 | Large-Scale Machine Learning at Twitter | 2012 | SIGMOD | 7.0162259e-05 |
| 4,141 | Efficient Type-Ahead Search on Relational Data: a TASTIER Approach | 2009 | SIGMOD | 6.7822757e-05 |
| 4,250 | The Unified Logging Infrastructure for Data Analytics at Twitter | 2012 | VLDB | 6.7005978e-05 |
| 7,602 | Avatara: OLAP for Web-scale Analytics Products | 2012 | VLDB | 5.484773e-05 |
| 8,811 | A Batch of PNUTS: Experiences Connecting Cloud Batch and Serving Systems | 2011 | SIGMOD | 5.2708766e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 9,446 | MapReduce Algorithms for Big Data Analysis | 2012 | VLDB |
| 2 | 6,279 | HadoopDB in Action: Building Real World Applications | 2010 | SIGMOD |
| 3 | 6,792 | Hadoop's Adolescence: An analysis of Hadoop usage in scientific workloads | 2013 | VLDB |
| 4 | 651 | Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience | 2009 | VLDB |
| 5 | 2,285 | A Platform for Scalable One-Pass Analytics using MapReduce | 2011 | SIGMOD |
| 6 | 120 | HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads | 2009 | VLDB |
| 7 | 8,074 | Emerging Trends in the Enterprise Data Analytics: Connecting Hadoop and DB2 Warehouse | 2011 | SIGMOD |
| 8 | 9,266 | Gobblin: Unifying Data Ingestion for Hadoop | 2015 | VLDB |
| 9 | 2,231 | Data Warehousing and Analytics Infrastructure at Facebook | 2010 | SIGMOD |
| 10 | 3,790 | Large-Scale Machine Learning at Twitter | 2012 | SIGMOD |