The "Big Data" Ecosystem at LinkedIn
Summary: LinkedIn's Hadoop-based analytics stack abstracts distributed systems for data scientists, enabling end-to-end analytics on massive data. Novelty: seamless last-mile integration—ingress/egress to online systems and production workflows—via a 1-line Pig command, with use cases in recommendations and feeds. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Roshan Sumbaly (LinkedIn)
- 2. Jay Kreps (LinkedIn)
- 3. Sam Shah (LinkedIn)
BibTeX Citation
@inproceedings{sumbaly_sigmod13,
title = {{The "Big Data" Ecosystem at LinkedIn}},
author = {Sumbaly, Roshan and Kreps, Jay and Shah, Sam},
series = {{SIGMOD} '13},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2463676.2463707},
url = {https://dl.acm.org/doi/10.1145/2463676.2463707},
year = {2013}
}
Incoming Citations (Sorted by Pagerank)
Showing 6 of 6 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,569 | HELIX: Holistic Optimization for Accelerating Iterative Machine Learning | 2019 | VLDB | 0.00010335423 |
| 6,121 | The Cosmos Big Data Platform at Microsoft: Over a Decade of Progress and a Decade to Look Forward | 2021 | VLDB | 5.9688569e-05 |
| 8,017 | Data Management for Social Networking | 2016 | PODS | 5.5064531e-05 |
| 8,237 | Meta-Dataflows: Efficient Exploratory Dataflow Jobs | 2018 | SIGMOD | 5.4601966e-05 |
| 9,369 | Redoop Infrastructure for Recurring Big Data Queries | 2014 | VLDB | 5.2774963e-05 |
| 12,156 | Shared Execution of Recurring Workloads in MapReduce | 2015 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 10 of 10 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 6 | Pig Latin: A Not-So-Foreign Language for Data Processing | 2008 | SIGMOD | 0.0010686205 |
| 13 | Mining Association Rules between Sets of Items in Large Databases | 1993 | SIGMOD | 0.0006567919 |
| 1,324 | Apache Hadoop Goes Realtime at Facebook | 2011 | SIGMOD | 0.00011149314 |
| 2,191 | Data Warehousing and Analytics Infrastructure at Facebook | 2010 | SIGMOD | 8.9780154e-05 |
| 3,566 | Efficient Bulk Insertion into a Distributed Ordered Table | 2008 | SIGMOD | 7.303217e-05 |
| 3,714 | Large-Scale Machine Learning at Twitter | 2012 | SIGMOD | 7.1764857e-05 |
| 4,048 | Efficient Type-Ahead Search on Relational Data: a TASTIER Approach | 2009 | SIGMOD | 6.9372074e-05 |
| 4,169 | The Unified Logging Infrastructure for Data Analytics at Twitter | 2012 | VLDB | 6.8561463e-05 |
| 7,470 | Avatara: OLAP for Web-scale Analytics Products | 2012 | VLDB | 5.6095423e-05 |
| 8,642 | A Batch of PNUTS: Experiences Connecting Cloud Batch and Serving Systems | 2011 | SIGMOD | 5.3942591e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 9,263 | MapReduce Algorithms for Big Data Analysis | 2012 | VLDB |
| 2 | 6,147 | HadoopDB in Action: Building Real World Applications | 2010 | SIGMOD |
| 3 | 6,736 | Hadoop's Adolescence: An analysis of Hadoop usage in scientific workloads | 2013 | VLDB |
| 4 | 642 | Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience | 2009 | VLDB |
| 5 | 2,265 | A Platform for Scalable One-Pass Analytics using MapReduce | 2011 | SIGMOD |
| 6 | 120 | HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads | 2009 | VLDB |
| 7 | 7,965 | Emerging Trends in the Enterprise Data Analytics: Connecting Hadoop and DB2 Warehouse | 2011 | SIGMOD |
| 8 | 9,081 | Gobblin: Unifying Data Ingestion for Hadoop | 2015 | VLDB |
| 9 | 2,191 | Data Warehousing and Analytics Infrastructure at Facebook | 2010 | SIGMOD |
| 10 | 3,714 | Large-Scale Machine Learning at Twitter | 2012 | SIGMOD |