Database Paper Browser

Back to papers

Hive - A Warehousing Solution Over a Map-Reduce Framework

Summary: Hive provides a data warehousing layer on Hadoop with SQL-like HiveQL compiled to MapReduce for scalable analytics on commodity hardware. It adds an extensible IO layer, a nested type system, and a centralized Hive-Metastore catalog for statistics and optimization, enabling large-scale deployments (thousands of tables, TB-scale data). (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
9963
Venue
VLDB
Year
2009
Pagerank
0.00059744625
Overall Rank
70 | 99.52%
DOI
-

Incoming Non-self Citations Over Time

Authors

Incoming Citations (Sorted by Pagerank)

Showing 50 of 133 citing papers.

Rank Citing Paper Year Venue Pagerank
3,247 Early Accurate Results for Advanced Analytics on MapReduce 2012 VLDB 7.3237519e-05
3,344 F1 Query: Declarative Querying at Scale 2018 VLDB 7.1944106e-05
3,510 Integrating Hadoop and Parallel DBMS 2010 SIGMOD 7.0280199e-05
3,511 M3R: Increased Performance for In-Memory Hadoop Jobs 2012 VLDB 7.0242945e-05
3,548 Adaptive Query Processing on RAW Data 2014 VLDB 6.9798836e-05
3,709 Multi-Query Optimization in MapReduce Framework 2014 VLDB 6.8211506e-05
3,753 Choosing A Cloud DBMS: Architectures and Tradeoffs 2019 VLDB 6.7850001e-05
3,839 GTS: A Fast and Scalable Graph Processing Method based on Streaming Topology to GPUs 2016 SIGMOD 6.7108721e-05
3,893 Slalom: Coasting Through Raw Data via Adaptive Partitioning and Indexing 2017 VLDB 6.653922e-05
3,923 Pushing Data-Induced Predicates Through Joins in Big-Data Clusters 2020 VLDB 6.6232068e-05
3,948 Unicorn: A System for Searching the Social Graph 2013 VLDB 6.5968941e-05
4,171 Computation Reuse in Analytics Job Service at Microsoft 2018 SIGMOD 6.3800823e-05
4,186 Apache Tez: A Unifying Framework for Modeling and Building Data Processing Applications 2015 SIGMOD 6.3701959e-05
4,199 Meet Charles, big data query advisor 2013 CIDR 6.3618577e-05
4,319 Fast Queries Over Heterogeneous Data Through Engine Customization 2016 VLDB 6.2823814e-05
4,516 An Empirical Evaluation of Columnar Storage Formats 2024 VLDB 6.1146215e-05
4,641 Algorithmic Aspects of Parallel Query Processing 2018 SIGMOD 6.0215749e-05
4,702 JSON Tiles: Fast Analytics on Semi-Structured Data 2021 SIGMOD 5.9796907e-05
4,766 Pinot: Realtime OLAP for 530 Million Users 2018 SIGMOD 5.933206e-05
4,862 OctopusFS: A Distributed File System with Tiered Storage Management 2017 SIGMOD 5.8659921e-05
5,012 Dynamically Optimizing Queries over Large Scale Data Platforms 2014 SIGMOD 5.7543101e-05
5,187 Yugong: Geo-Distributed Data and Job Placement at Scale 2019 VLDB 5.6351231e-05
5,292 StreamOps: Cloud-Native Runtime Management for Streaming Services in ByteDance 2023 VLDB 5.5784769e-05
5,304 ReCache: Reactive Caching for Fast Analytics over Heterogeneous Data 2018 VLDB 5.5738777e-05
5,318 LOCAT: Low-Overhead Online Configuration Auto-Tuning of Spark SQL Applications 2022 SIGMOD 5.5685434e-05
5,373 Fine-Grained Modeling and Optimization for Intelligent Resource Management in Big Data Processing 2022 VLDB 5.5410059e-05
5,543 Lightweight Cardinality Estimation in LSM-based Systems 2018 SIGMOD 5.4486922e-05
5,800 AQWA: Adaptive Query-Workload-Aware Partitioning of Big Spatial Data 2015 VLDB 5.3218628e-05
5,855 MIFO: A Query-Semantic Aware Resource Allocation Policy 2019 SIGMOD 5.297903e-05
5,908 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 5.2731311e-05
5,985 The Era of Big Spatial Data 2017 VLDB 5.2399365e-05
6,120 REEF: Retainable Evaluator Execution Framework 2015 SIGMOD 5.199013e-05
6,301 Elastic Pipelining in an In-Memory Database Cluster 2016 SIGMOD 5.1172165e-05
6,302 Hillview: A trillion-cell spreadsheet for big data 2019 VLDB 5.1166201e-05
6,397 BigLake: BigQuery’s Evolution toward a Multi-Cloud Lakehouse 2024 SIGMOD 5.0749432e-05
6,403 Just-In-Time Data Virtualization: Lightweight Data Management with ViDa 2015 CIDR 5.0717043e-05
6,477 Towards Unified Ad-hoc Data Processing 2014 SIGMOD 5.0408007e-05
6,493 Memory-Aware Framework for Efficient Second-Order Random Walk on Large Graphs 2020 SIGMOD 5.0344095e-05
6,590 Interactive Demonstration of Probabilistic Predicates 2018 SIGMOD 4.996343e-05
6,671 Incorporating Super-Operators in Big-Data Query Optimizers 2020 VLDB 4.9625353e-05
6,675 Exploiting Common Patterns for Tree-Structured Data 2017 SIGMOD 4.9615691e-05
6,856 Liquid: Unifying Nearline and Offline Big Data Integration 2015 CIDR 4.901516e-05
7,079 JetScope: Reliable and Interactive Analytics at Cloud Scale 2015 VLDB 4.8353804e-05
7,099 KEA: Tuning an Exabyte-Scale Data Infrastructure 2021 SIGMOD 4.8263529e-05
7,195 BSMA: A Benchmark for Analytical Queries over Social Media Data 2014 VLDB 4.799032e-05
7,205 Kodiak: Leveraging Materialized Views For Very Low-Latency Analytics Over High-Dimensional Web-Scale Data 2016 VLDB 4.7965293e-05
7,269 Oracle In-Database Hadoop: When MapReduce Meets RDBMS 2012 SIGMOD 4.7768528e-05
7,398 SmartBench: A Benchmark For Data Management In Smart Spaces 2020 VLDB 4.7364678e-05
7,459 Lachesis: Automatic Partitioning for UDF-Centric Analytics 2021 VLDB 4.7199075e-05
7,470 Bullion: A Column Store for Machine Learning 2025 CIDR 4.7159125e-05
Previous Page 2 / 3 Next

Outgoing Citations (Sorted by Pagerank)

Showing 2 of 2 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
22 SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets 2008 VLDB 0.00084679526
42 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.00073570328
Previous Page 1 / 1 Next

Semantically Similar Papers