Apache Hive: From MapReduce to Enterprise-grade Big Data Warehousing
Summary: Apache Hive evolves from MapReduce to an enterprise-grade data warehouse via a hybrid MPP/big-data architecture that integrates SQL, storage formats, and cloud concepts. Innovations cover Transactions, optimizer, runtime, and federation, with experiments on typical workloads and a forward-looking community roadmap. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Jesús Camacho-Rodríguez (Hortonworks)
- 2. Ashutosh Chauhan (Hortonworks)
- 3. Alan Gates (Hortonworks)
- 4. Eugene Koifman (Hortonworks)
- 5. Owen O’Malley (Hortonworks)
- 6. Vineet Garg (Hortonworks)
- 7. Zoltan Haindrich (Hortonworks)
- 8. Sergey Shelukhin (Hortonworks)
- 9. Prasanth Jayachandran (Hortonworks)
- 10. Siddharth Seth (Hortonworks)
- 11. Deepak Jaiswal (Hortonworks)
- 12. Slim Bouguerra (Hortonworks)
- 13. Nishant Bangarwa (Hortonworks)
- 14. Sankar Hariappan (Hortonworks)
- 15. Anishek Agarwal (Hortonworks)
- 16. Jason Dere (Hortonworks)
- 17. Daniel Dai (Hortonworks)
- 18. Thejas Nair (Hortonworks)
- 19. Nita Dembla (Hortonworks)
- 20. Gopal Vijayaraghavan (Hortonworks)
- 21. Günther Hagleitner (Hortonworks)
BibTeX Citation
@inproceedings{camachorodriguez_sigmod19,
title = {{Apache Hive: From MapReduce to Enterprise-grade Big Data Warehousing}},
author = {Camacho-Rodríguez, Jesús and Chauhan, Ashutosh and Gates, Alan and Koifman, Eugene and O’Malley, Owen and Garg, Vineet and Haindrich, Zoltan and Shelukhin, Sergey and Jayachandran, Prasanth and Seth, Siddharth and Jaiswal, Deepak and Bouguerra, Slim and Bangarwa, Nishant and Hariappan, Sankar and Agarwal, Anishek and Dere, Jason and Dai, Daniel and Nair, Thejas and Dembla, Nita and Vijayaraghavan, Gopal and Hagleitner, Günther},
series = {{SIGMOD} '19},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3299869.3314045},
url = {https://dl.acm.org/doi/10.1145/3299869.3314045},
year = {2019}
}
Incoming Citations (Sorted by Pagerank)
Showing 12 of 12 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 20 of 20 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 1,791 | Cheetah: A High Performance, Custom Data Warehouse on Top of MapReduce | 2010 | VLDB |
| 2 | 445 | Apache Calcite: A Foundational Framework for Optimized Query Processing Over Heterogeneous Data Sources | 2018 | SIGMOD |
| 3 | 1,889 | SQL-on-Hadoop: Full Circle Back to Shared-Nothing Database Architectures | 2014 | VLDB |
| 4 | 11,485 | DHive: Query Execution Performance Analysis via Dataflow in Apache Hive | 2023 | VLDB |
| 5 | 2,191 | Data Warehousing and Analytics Infrastructure at Facebook | 2010 | SIGMOD |
| 6 | 2,159 | Efficient Processing of Data Warehousing Queries in a Split Execution Environment | 2011 | SIGMOD |
| 7 | 9,079 | QMapper for Smart Grid: Migrating SQL-based Application to Hive | 2015 | SIGMOD |
| 8 | 120 | HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads | 2009 | VLDB |
| 9 | 2,706 | Major Technical Advancements in Apache Hive | 2014 | SIGMOD |
| 10 | 32 | Hive - A Warehousing Solution Over a Map-Reduce Framework | 2009 | VLDB |