Llama: Leveraging Columnar Storage for Scalable Join Processing in the MapReduce Framework
Summary: Llama uses columnar storage with vertical partitioning via correlation groups to enable scalable join processing on a DFS-based MapReduce engine. A new join algorithm and TPC-H evaluation show faster loading and superior query performance vs Hive on EC2. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Yuting Lin (National University of Singapore)
- 2. Divyakant Agrawal (University of California Santa Barbara)
- 3. Chun Chen (Zhejiang University)
- 4. Beng Chin Ooi (National University of Singapore)
- 5. Sai Wu (National University of Singapore)
BibTeX Citation
@inproceedings{lin_sigmod11,
title = {{Llama: Leveraging Columnar Storage for Scalable Join Processing in the MapReduce Framework}},
author = {Lin, Yuting and Agrawal, Divyakant and Chen, Chun and Ooi, Beng Chin and Wu, Sai},
series = {{SIGMOD} '11},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/1989323.1989424},
url = {https://dl.acm.org/doi/10.1145/1989323.1989424},
year = {2011}
}
Incoming Citations (Sorted by Pagerank)
Showing 11 of 11 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,539 | Minimal MapReduce Algorithms | 2013 | SIGMOD | 8.4526595e-05 |
| 2,706 | Major Technical Advancements in Apache Hive | 2014 | SIGMOD | 8.2287564e-05 |
| 3,008 | Scalable Big Graph Processing in MapReduce | 2014 | SIGMOD | 7.8578871e-05 |
| 3,714 | Large-Scale Machine Learning at Twitter | 2012 | SIGMOD | 7.1764857e-05 |
| 4,169 | The Unified Logging Infrastructure for Data Analytics at Twitter | 2012 | VLDB | 6.8561463e-05 |
| 6,199 | Clydesdale: Structured Data Processing on Hadoop | 2012 | SIGMOD | 5.9458654e-05 |
| 6,640 | Understanding Insights into the Basic Structure and Essential Issues of Table Placement Methods in Clusters | 2013 | VLDB | 5.8162526e-05 |
| 8,276 | CARTILAGE: Adding Flexibility to the Hadoop Skeleton | 2013 | SIGMOD | 5.4574671e-05 |
| 9,497 | Rank Join Queries in NoSQL Databases | 2014 | VLDB | 5.2608378e-05 |
| 9,510 | Efficient Big Data Processing in Hadoop MapReduce | 2012 | VLDB | 5.2576928e-05 |
| 12,204 | YZStack: Provisioning Customizable Solution for Big Data | 2014 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 11 of 11 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 12 | C-Store: A Column-oriented DBMS | 2005 | VLDB | 0.00069513174 |
| 29 | Database Architecture Optimized for the New Bottleneck: Memory Access | 1999 | VLDB | 0.00052093615 |
| 49 | Weaving Relations for Cache Performance | 2001 | VLDB | 0.00043781096 |
| 51 | Dremel: Interactive Analysis of Web-Scale Datasets | 2010 | VLDB | 0.0004291425 |
| 60 | Integrating Compression and Execution in Column-Oriented Database Systems | 2006 | SIGMOD | 0.0003955489 |
| 115 | A Decomposition Storage Model | 1985 | SIGMOD | 0.00032338948 |
| 120 | HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads | 2009 | VLDB | 0.00031680027 |
| 660 | Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) | 2010 | VLDB | 0.00015198804 |
| 769 | A Comparison of Join Algorithms for Log Processing in MapReduce | 2010 | SIGMOD | 0.00014166872 |
| 970 | Sybase IQ Multiplex – Designed For Analytics | 2004 | VLDB | 0.00012882125 |
| 1,436 | The Performance of MapReduce: An In-depth Study | 2010 | VLDB | 0.00010797443 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 1,436 | The Performance of MapReduce: An In-depth Study | 2010 | VLDB |
| 2 | 769 | A Comparison of Join Algorithms for Log Processing in MapReduce | 2010 | SIGMOD |
| 3 | 2,567 | ClusterJoin: A Similarity Joins Framework using Map-Reduce | 2014 | VLDB |
| 4 | 9,497 | Rank Join Queries in NoSQL Databases | 2014 | VLDB |
| 5 | 1,791 | Cheetah: A High Performance, Custom Data Warehouse on Top of MapReduce | 2010 | VLDB |
| 6 | 12,090 | Let's Rethink Join Optimization in Distributed Systems | 2015 | CIDR |
| 7 | 892 | Rack-Scale In-Memory Join Processing using RDMA | 2015 | SIGMOD |
| 8 | 1,889 | SQL-on-Hadoop: Full Circle Back to Shared-Nothing Database Architectures | 2014 | VLDB |
| 9 | 2,159 | Efficient Processing of Data Warehousing Queries in a Split Execution Environment | 2011 | SIGMOD |
| 10 | 2,849 | Column-Oriented Storage Techniques for MapReduce | 2011 | VLDB |