Database Paper Browser

Back to papers

Hive - A Warehousing Solution Over a Map-Reduce Framework

Summary: Hive provides a data warehousing layer on Hadoop with SQL-like HiveQL compiled to MapReduce for scalable analytics on commodity hardware. It adds an extensible IO layer, a nested type system, and a centralized Hive-Metastore catalog for statistics and optimization, enabling large-scale deployments (thousands of tables, TB-scale data). (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
9963
Venue
VLDB
Year
2009
Pagerank
0.00059744625
Overall Rank
70 | 99.52%
DOI
-

Incoming Non-self Citations Over Time

Authors

Incoming Citations (Sorted by Pagerank)

Showing 33 of 133 citing papers.

Rank Citing Paper Year Venue Pagerank
7,533 Enabling Efficient and General Subpopulation Analytics in Multidimensional Data Streams 2022 VLDB 4.7134753e-05
7,600 Quill: Efficient, Transferable, and Rich Analytics at Scale 2016 VLDB 4.697054e-05
7,778 Runtime Variation in Big Data Analytics 2023 SIGMOD 4.6491879e-05
8,460 Piranha: Optimizing Short Jobs in Hadoop 2013 VLDB 4.5008938e-05
8,502 New Query Optimization Techniques in the Spark Engine of Azure Synapse 2022 VLDB 4.491819e-05
8,585 A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning 2024 VLDB 4.4856045e-05
8,754 Hyperspace: The Indexing Subsystem of Azure Synapse 2021 VLDB 4.4520434e-05
8,927 QMapper for Smart Grid: Migrating SQL-based Application to Hive 2015 SIGMOD 4.4229886e-05
9,116 Towards Observability for Production Machine Learning Pipelines 2022 VLDB 4.3886184e-05
9,138 Phoebe: A Learning-based Checkpoint Optimizer 2021 VLDB 4.3842765e-05
9,203 F3: The Open-Source Data File Format for the Future 2026 SIGMOD 4.3701616e-05
9,353 Rank Join Queries in NoSQL Databases 2014 VLDB 4.3485003e-05
9,384 Versatile Optimization of UDF-heavy Data Flows with Sofa 2014 SIGMOD 4.3432098e-05
9,391 Sapprox: Enabling Efficient and Accurate Approximations on Sub-datasets with Distribution-aware Online Sampling 2017 VLDB 4.3402854e-05
9,691 GHive: A Demonstration of GPU-Accelerated Query Processing in Apache Hive 2022 SIGMOD 4.2987288e-05
9,893 OceanRT: Real-Time Analytics over Large Temporal Data 2014 SIGMOD 4.2561802e-05
10,196 PTO: A Workload-driven Predictive Table Optimizer for Lakehouse Systems 2026 SIGMOD 4.1905499e-05
10,422 OpenMLDB: A Real-Time Relational Data Feature Computation System for Online ML 2025 SIGMOD 4.1905499e-05
10,577 QOVIS: Understanding and Diagnosing Query Optimizer via a Visualization-assisted Approach 2025 VLDB 4.1905499e-05
10,644 Concurrency Control as a Service 2025 VLDB 4.1905499e-05
10,670 ArrayMorph: Optimizing Hyperslab Queries on the Cloud for Machine Learning Pipelines 2025 VLDB 4.1905499e-05
11,087 Presto’s History-based Query Optimizer 2024 VLDB 4.1905499e-05
11,391 CDI-E: An Elastic Cloud Service for Data Engineering 2022 VLDB 4.1905499e-05
11,695 Integration of Large-Scale Data Processing Systems and Traditional Parallel Database Technology 2019 VLDB 4.1905499e-05
11,839 Logical Aspects of Massively Parallel and Distributed Systems 2016 PODS 4.1905499e-05
11,867 dmapply: A functional primitive to express distributed machine learning algorithms in R 2016 VLDB 4.1905499e-05
11,898 Let's Rethink Join Optimization in Distributed Systems 2015 CIDR 4.1905499e-05
11,924 A Demonstration of Rubato DB: A Highly Scalable NewSQL Database System for OLTP and Big Data Applications 2015 SIGMOD 4.1905499e-05
11,984 Anti-Combining for MapReduce 2014 SIGMOD 4.1905499e-05
12,007 Getting Your Big Data Priorities Straight: A Demonstration of Priority-based QoS using Social-network-driven Stock Recommendation 2014 VLDB 4.1905499e-05
12,013 Design and Implementation of a Real-Time Interactive Analytics System for Large Spatio-Temporal Data 2014 VLDB 4.1905499e-05
12,117 Declarative Error Management for Robust Data-Intensive Applications 2012 SIGMOD 4.1905499e-05
12,211 Resiliency-Aware Data Management 2011 VLDB 4.1905499e-05
Previous Page 3 / 3 Next

Outgoing Citations (Sorted by Pagerank)

Showing 2 of 2 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
22 SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets 2008 VLDB 0.00084679526
42 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.00073570328
Previous Page 1 / 1 Next

Semantically Similar Papers