DBScholar

Back to papers

Hive - A Warehousing Solution Over a Map-Reduce Framework

Summary: Hive brings SQL-like declarative warehousing to Hadoop by compiling HiveQL into MapReduce, while retaining extensibility via custom scripts and data formats. Its rich nested type system and catalog of schemas/statistics enable large-scale exploration and optimization. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
hcc7c6e4b04247911
Venue
VLDB
Year
2009
Pagerank
0.00049839909
Overall Rank
31 | 99.80%
DOI
10.14778/1687553.1687609

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{thusoo_vldb09,
        title = {{Hive - A Warehousing Solution Over a Map-Reduce Framework}},
        author = {Thusoo, Ashish and Sarma, Joydeep Sen and Jain, Namit and Shao, Zheng and Chakka, Prasad and Anthony, Suresh and Liu, Hao and Wyckoff, Pete and Murthy, Raghotham},
        journal = {PVLDB},
        series = {{VLDB} '09},
        volume = {2},
        number = {2},
        pages = {1626--1629},
        doi = {10.14778/1687553.1687609},
        url = {https://doi.org/10.14778/1687553.1687609},
        year = {2009}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 136 citing papers.

Rank Citing Paper Year Venue Pagerank
148 Gorilla: A Fast, Scalable, In-Memory Time Series Database 2015 VLDB 0.0002900671
179 The Vertica Analytic Database: C-Store 7 Years Later 2012 VLDB 0.00026611886
281 Accelerating Machine Learning Inference with Probabilistic Predicates 2018 SIGMOD 0.00022295232
325 The Dataflow Model: A Practical Approach to Balancing Correctness, Latency, and Cost in Massive-Scale, Unbounded, Out-of-Order Data Processing 2015 VLDB 0.00020964941
379 Apache Calcite: A Foundational Framework for Optimized Query Processing Over Heterogeneous Data Sources 2018 SIGMOD 0.00019514689
394 One Trillion Edges: Graph Processing at Facebook-Scale 2015 VLDB 0.00019191286
673 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.0001488755
823 MRShare: Sharing Across Multiple Queries in MapReduce 2010 VLDB 0.00013648332
840 Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters 2016 SIGMOD 0.0001354605
876 Tenzing: A SQL Implementation On The MapReduce Framework 2011 VLDB 0.00013309176
977 Parallel Evaluation of Conjunctive Queries 2011 PODS 0.00012731794
1,009 Simulation of Database-Valued Markov Chains Using SimSQL 2013 SIGMOD 0.00012539827
1,036 Fine-grained Partitioning for Aggressive Data Skipping 2014 SIGMOD 0.00012377471
1,037 Jaql: A Scripting Language for Large Scale Semistructured Data Analysis 2011 VLDB 0.00012376819
1,072 Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads 2012 VLDB 0.00012168947
1,080 Hadoop-GIS: A High Performance Spatial Data Warehousing System over MapReduce 2013 VLDB 0.00012131978
1,081 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012123917
1,082 Approximate Query Processing: No Silver Bullet 2017 SIGMOD 0.00012122749
1,178 Simba: Efficient In-Memory Spatial Analytics 2016 SIGMOD 0.00011632691
1,231 Communication Steps for Parallel Query Processing 2013 PODS 0.00011410373
1,371 Ricardo: Integrating R and Hadoop 2010 SIGMOD 0.00010899041
1,381 Scuba: Diving into Data at Facebook 2013 VLDB 0.00010860462
1,388 Provenance for Generalized Map and Reduce Workflows 2011 CIDR 0.00010821908
1,421 V-SMART-Join: A Scalable MapReduce Framework for All-Pair Similarity Joins of Multisets and Vectors 2012 VLDB 0.00010726757
1,428 Knowing When You’re Wrong: Building Fast and Reliable Approximate Query Processing Systems 2014 SIGMOD 0.00010693831
1,433 Towards a Learning Optimizer for Shared Clouds 2019 VLDB 0.00010677711
1,462 Manu: A Cloud Native Vector Database Management System 2022 VLDB 0.0001058099
1,466 The Performance of MapReduce: An In-depth Study 2010 VLDB 0.00010569837
1,513 Cloud-Native Database Systems at Alibaba: Opportunities and Challenges 2019 VLDB 0.00010429438
1,564 Titian: Data Provenance Support in Spark 2016 VLDB 0.00010222394
1,622 Mesa: Geo-Replicated, Near Real-Time, Scalable Data Warehousing 2014 VLDB 0.00010051462
1,745 Selecting Subexpressions to Materialize at Datacenter Scale 2018 VLDB 9.7343818e-05
1,771 Samza: Stateful Scalable Stream Processing at LinkedIn 2017 VLDB 9.6803752e-05
1,829 DBEst: Revisiting Approximate Query Processing Engines with Machine Learning Models 2019 SIGMOD 9.5510333e-05
1,837 Cheetah: A High Performance, Custom Data Warehouse on Top of MapReduce 2010 VLDB 9.5350058e-05
1,924 ReStore: Reusing Results of MapReduce Jobs 2012 VLDB 9.3687009e-05
1,931 Instant Loading for Main Memory Databases 2013 VLDB 9.3474242e-05
2,195 Opening the Black Boxes in Data Flow Optimization 2012 VLDB 8.8781177e-05
2,247 Cumulon: Optimizing Statistical Data Analysis in the Cloud 2013 SIGMOD 8.7585767e-05
2,282 A Platform for Scalable One-Pass Analytics using MapReduce 2011 SIGMOD 8.6964781e-05
2,361 Online Aggregation and Continuous Query support in MapReduce 2010 SIGMOD 8.5761274e-05
2,410 epiC: an Extensible and Scalable System for Processing Big Data 2014 VLDB 8.5114032e-05
2,423 BigDansing: A System for Big Data Cleansing 2015 SIGMOD 8.4877894e-05
2,484 AnalyticDB: Real-time OLAP Database System at Alibaba Cloud 2019 VLDB 8.3993882e-05
2,499 Towards Scalable Real-time Analytics: An Architecture for Scale-out of OLxP Workloads 2015 VLDB 8.3815142e-05
2,550 SQLShare: Results from a Multi-Year SQL-as-a-Service Experiment 2016 SIGMOD 8.3121033e-05
2,635 Big Data Analytics with Datalog Queries on Spark 2016 SIGMOD 8.1965216e-05
2,687 CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop 2011 VLDB 8.1301312e-05
2,754 Implicit Parallelism through Deep Language Embedding 2015 SIGMOD 8.0534972e-05
2,898 HAWQ: A Massively Parallel Processing SQL Engine in Hadoop 2014 SIGMOD 7.8853204e-05
Previous Page 1 / 3 Next

Outgoing Citations (Sorted by Pagerank)

Showing 2 of 2 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
30 SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets 2008 VLDB 0.00050495102
43 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.00045546775
Previous Page 1 / 1 Next

Semantically Similar Papers