The "Big Data" Ecosystem at LinkedIn
Summary: LinkedIn's Hadoop-based analytics stack abstracts distributed systems for data scientists, enabling end-to-end analytics on massive data. Novelty: seamless last-mile integration—ingress/egress to online systems and production workflows—via a 1-line Pig command, with use cases in recommendations and feeds. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Roshan Sumbaly (LinkedIn)
- 2. Jay Kreps (LinkedIn)
- 3. Sam Shah (LinkedIn)
BibTeX Citation
@inproceedings{sumbaly_sigmod13,
title = {{The "Big Data" Ecosystem at LinkedIn}},
author = {Sumbaly, Roshan and Kreps, Jay and Shah, Sam},
series = {{SIGMOD} '13},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2463676.2463707},
url = {https://dl.acm.org/doi/10.1145/2463676.2463707},
year = {2013}
}
Incoming Citations (Sorted by Pagerank)
Showing 6 of 6 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,568 | HELIX: Holistic Optimization for Accelerating Iterative Machine Learning | 2019 | VLDB | 0.0001021302 |
| 6,245 | The Cosmos Big Data Platform at Microsoft: Over a Decade of Progress and a Decade to Look Forward | 2021 | VLDB | 5.837329e-05 |
| 8,177 | Data Management for Social Networking | 2016 | PODS | 5.3829638e-05 |
| 8,376 | Meta-Dataflows: Efficient Exploratory Dataflow Jobs | 2018 | SIGMOD | 5.3437059e-05 |
| 9,548 | Redoop Infrastructure for Recurring Big Data Queries | 2014 | VLDB | 5.1601681e-05 |
| 12,447 | Shared Execution of Recurring Workloads in MapReduce | 2015 | VLDB | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 10 of 10 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 6 | Pig Latin: A Not-So-Foreign Language for Data Processing | 2008 | SIGMOD | 0.001052036 |
| 13 | Mining Association Rules between Sets of Items in Large Databases | 1993 | SIGMOD | 0.00064420972 |
| 1,302 | Apache Hadoop Goes Realtime at Facebook | 2011 | SIGMOD | 0.00011111126 |
| 2,230 | Data Warehousing and Analytics Infrastructure at Facebook | 2010 | SIGMOD | 8.799877e-05 |
| 3,639 | Efficient Bulk Insertion into a Distributed Ordered Table | 2008 | SIGMOD | 7.1444069e-05 |
| 3,788 | Large-Scale Machine Learning at Twitter | 2012 | SIGMOD | 7.0195464e-05 |
| 4,140 | Efficient Type-Ahead Search on Relational Data: a TASTIER Approach | 2009 | SIGMOD | 6.7854783e-05 |
| 4,249 | The Unified Logging Infrastructure for Data Analytics at Twitter | 2012 | VLDB | 6.7037533e-05 |
| 7,596 | Avatara: OLAP for Web-scale Analytics Products | 2012 | VLDB | 5.4873503e-05 |
| 8,803 | A Batch of PNUTS: Experiences Connecting Cloud Batch and Serving Systems | 2011 | SIGMOD | 5.2733728e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 9,437 | MapReduce Algorithms for Big Data Analysis | 2012 | VLDB |
| 2 | 6,276 | HadoopDB in Action: Building Real World Applications | 2010 | SIGMOD |
| 3 | 6,786 | Hadoop's Adolescence: An analysis of Hadoop usage in scientific workloads | 2013 | VLDB |
| 4 | 651 | Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience | 2009 | VLDB |
| 5 | 2,282 | A Platform for Scalable One-Pass Analytics using MapReduce | 2011 | SIGMOD |
| 6 | 120 | HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads | 2009 | VLDB |
| 7 | 8,068 | Emerging Trends in the Enterprise Data Analytics: Connecting Hadoop and DB2 Warehouse | 2011 | SIGMOD |
| 8 | 9,256 | Gobblin: Unifying Data Ingestion for Hadoop | 2015 | VLDB |
| 9 | 2,230 | Data Warehousing and Analytics Infrastructure at Facebook | 2010 | SIGMOD |
| 10 | 3,788 | Large-Scale Machine Learning at Twitter | 2012 | SIGMOD |