Impala: A Modern, Open-Source SQL Engine for Hadoop
Summary: Impala: an open-source MPP SQL engine for Hadoop providing low-latency, high-concurrency execution for BI/read-mostly analytic queries where batch frameworks (e.g., Hive) fall short. Paper presents architecture/components and empirical superiority vs other SQL-on-Hadoop systems. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Marcel Kornacker (Cloudera)
- 2. Alexander Behm (Cloudera)
- 3. Victor Bittorf (Cloudera)
- 4. Taras Bobrovytsky (Cloudera)
- 5. Casey Ching (Cloudera)
- 6. Alan Choi (Cloudera)
- 7. Justin Erickson (Cloudera)
- 8. Martin Grund (Cloudera)
- 9. Daniel Hecht (Cloudera)
- 10. Matthew Jacobs (Cloudera)
- 11. Ishaan Joshi (Cloudera)
- 12. Lenni Kuff (Cloudera)
- 13. Dileep Kumar (Cloudera)
- 14. Alex Leblang (Cloudera)
- 15. Nong Li (Cloudera)
- 16. Ippokratis Pandis (Cloudera)
- 17. Henry Robinson (Cloudera)
- 18. David Rorke (Cloudera)
- 19. Silvius Rus (Cloudera)
- 20. John Russell (Cloudera)
- 21. Dimitris Tsirogiannis (Cloudera)
- 22. Skye Wanderman-Milne (Cloudera)
- 23. Michael Yoder (Cloudera)
BibTeX Citation
@inproceedings{kornacker_cidr15,
address = {Amsterdam, Netherlands},
series = {{CIDR} '15},
title = {{Impala: A Modern, Open-Source SQL Engine for Hadoop}},
booktitle = {Proceedings of the {Conference} on {Innovative} {Data} {Systems} {Research}},
author = {Kornacker, Marcel and Behm, Alexander and Bittorf, Victor and Bobrovytsky, Taras and Ching, Casey and Choi, Alan and Erickson, Justin and Grund, Martin and Hecht, Daniel and Jacobs, Matthew and Joshi, Ishaan and Kuff, Lenni and Kumar, Dileep and Leblang, Alex and Li, Nong and Pandis, Ippokratis and Robinson, Henry and Rorke, David and Rus, Silvius and Russell, John and Tsirogiannis, Dimitris and Wanderman-Milne, Skye and Yoder, Michael},
year = {2015}
}
Incoming Citations (Sorted by Pagerank)
Showing 12 of 62 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 49 | Weaving Relations for Cache Performance | 2001 | VLDB | 0.00043781096 |
| 51 | Dremel: Interactive Analysis of Web-Scale Datasets | 2010 | VLDB | 0.0004291425 |
| 98 | Encapsulation of Parallelism in the Volcano Query Processing System | 1990 | SIGMOD | 0.00034510605 |
| 165 | DB2 with BLU Acceleration: So Much More than Just a Column Store | 2013 | VLDB | 0.00027693424 |
| 216 | SIMD-Scan: Ultra Fast in-Memory Table Scan using on-Chip Vector Processing Units | 2009 | VLDB | 0.00024498128 |
| 1,889 | SQL-on-Hadoop: Full Circle Back to Shared-Nothing Database Architectures | 2014 | VLDB | 9.5335988e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 6,122 | iOLAP: Managing Uncertainty for Efficient Incremental OLAP | 2016 | SIGMOD |
| 2 | 11,885 | Integration of Large-Scale Data Processing Systems and Traditional Parallel Database Technology | 2019 | VLDB |
| 3 | 9,640 | Supporting Scalable Analytics with Latency Constraints | 2015 | VLDB |
| 4 | 7,689 | Oracle In-Database Hadoop: When MapReduce Meets RDBMS | 2012 | SIGMOD |
| 5 | 8,072 | Operational Analytics Data Management Systems | 2016 | VLDB |
| 6 | 120 | HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads | 2009 | VLDB |
| 7 | 2,850 | HAWQ: A Massively Parallel Processing SQL Engine in Hadoop | 2014 | SIGMOD |
| 8 | 12,146 | Tutorial: SQL-on-Hadoop Systems | 2015 | VLDB |
| 9 | 3,555 | Apache Hive: From MapReduce to Enterprise-grade Big Data Warehousing | 2019 | SIGMOD |
| 10 | 1,889 | SQL-on-Hadoop: Full Circle Back to Shared-Nothing Database Architectures | 2014 | VLDB |