Llama: Leveraging Columnar Storage for Scalable Join Processing in the MapReduce Framework
Summary: Llama uses columnar storage with vertical partitioning via correlation groups to enable scalable join processing on a DFS-based MapReduce engine. A new join algorithm and TPC-H evaluation show faster loading and superior query performance vs Hive on EC2. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Yuting Lin (National University of Singapore)
- 2. Divyakant Agrawal (University of California Santa Barbara)
- 3. Chun Chen (Zhejiang University)
- 4. Beng Chin Ooi (National University of Singapore)
- 5. Sai Wu (National University of Singapore)
BibTeX Citation
@inproceedings{lin_sigmod11,
title = {{Llama: Leveraging Columnar Storage for Scalable Join Processing in the MapReduce Framework}},
author = {Lin, Yuting and Agrawal, Divyakant and Chen, Chun and Ooi, Beng Chin and Wu, Sai},
series = {{SIGMOD} '11},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/1989323.1989424},
url = {https://dl.acm.org/doi/10.1145/1989323.1989424},
year = {2011}
}
Incoming Citations (Sorted by Pagerank)
Showing 11 of 11 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,573 | Minimal MapReduce Algorithms | 2013 | SIGMOD | 8.2821647e-05 |
| 2,689 | Major Technical Advancements in Apache Hive | 2014 | SIGMOD | 8.1264718e-05 |
| 3,066 | Scalable Big Graph Processing in MapReduce | 2014 | SIGMOD | 7.6877117e-05 |
| 3,788 | Large-Scale Machine Learning at Twitter | 2012 | SIGMOD | 7.0195464e-05 |
| 4,249 | The Unified Logging Infrastructure for Data Analytics at Twitter | 2012 | VLDB | 6.7037533e-05 |
| 6,322 | Clydesdale: Structured Data Processing on Hadoop | 2012 | SIGMOD | 5.8148318e-05 |
| 6,763 | Understanding Insights into the Basic Structure and Essential Issues of Table Placement Methods in Clusters | 2013 | VLDB | 5.6874593e-05 |
| 8,442 | CARTILAGE: Adding Flexibility to the Hadoop Skeleton | 2013 | SIGMOD | 5.3350162e-05 |
| 9,681 | Rank Join Queries in NoSQL Databases | 2014 | VLDB | 5.1427987e-05 |
| 9,693 | Efficient Big Data Processing in Hadoop MapReduce | 2012 | VLDB | 5.1399537e-05 |
| 12,495 | YZStack: Provisioning Customizable Solution for Big Data | 2014 | VLDB | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 11 of 11 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 12 | C-Store: A Column-oriented DBMS | 2005 | VLDB | 0.00068998927 |
| 27 | Database Architecture Optimized for the New Bottleneck: Memory Access | 1999 | VLDB | 0.0005158963 |
| 48 | Weaving Relations for Cache Performance | 2001 | VLDB | 0.00043805923 |
| 49 | Dremel: Interactive Analysis of Web-Scale Datasets | 2010 | VLDB | 0.00043160717 |
| 61 | Integrating Compression and Execution in Column-Oriented Database Systems | 2006 | SIGMOD | 0.000392237 |
| 114 | A Decomposition Storage Model | 1985 | SIGMOD | 0.00031928929 |
| 120 | HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads | 2009 | VLDB | 0.000311132 |
| 673 | Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) | 2010 | VLDB | 0.0001488755 |
| 799 | A Comparison of Join Algorithms for Log Processing in MapReduce | 2010 | SIGMOD | 0.00013889081 |
| 980 | Sybase IQ Multiplex – Designed For Analytics | 2004 | VLDB | 0.00012720677 |
| 1,466 | The Performance of MapReduce: An In-depth Study | 2010 | VLDB | 0.00010569837 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 1,466 | The Performance of MapReduce: An In-depth Study | 2010 | VLDB |
| 2 | 799 | A Comparison of Join Algorithms for Log Processing in MapReduce | 2010 | SIGMOD |
| 3 | 2,577 | ClusterJoin: A Similarity Joins Framework using Map-Reduce | 2014 | VLDB |
| 4 | 9,681 | Rank Join Queries in NoSQL Databases | 2014 | VLDB |
| 5 | 1,837 | Cheetah: A High Performance, Custom Data Warehouse on Top of MapReduce | 2010 | VLDB |
| 6 | 12,383 | Let's Rethink Join Optimization in Distributed Systems | 2015 | CIDR |
| 7 | 889 | Rack-Scale In-Memory Join Processing using RDMA | 2015 | SIGMOD |
| 8 | 1,884 | SQL-on-Hadoop: Full Circle Back to Shared-Nothing Database Architectures | 2014 | VLDB |
| 9 | 2,173 | Efficient Processing of Data Warehousing Queries in a Split Execution Environment | 2011 | SIGMOD |
| 10 | 2,903 | Column-Oriented Storage Techniques for MapReduce | 2011 | VLDB |