DBScholar

Back to papers

Hive - A Warehousing Solution Over a Map-Reduce Framework

Summary: Hive brings SQL-like declarative warehousing to Hadoop by compiling HiveQL into MapReduce, while retaining extensibility via custom scripts and data formats. Its rich nested type system and catalog of schemas/statistics enable large-scale exploration and optimization. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
hcc7c6e4b04247911
Venue
VLDB
Year
2009
Pagerank
0.00049821554
Overall Rank
31 | 99.80%
DOI
10.14778/1687553.1687609

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{thusoo_vldb09,
        title = {{Hive - A Warehousing Solution Over a Map-Reduce Framework}},
        author = {Thusoo, Ashish and Sarma, Joydeep Sen and Jain, Namit and Shao, Zheng and Chakka, Prasad and Anthony, Suresh and Liu, Hao and Wyckoff, Pete and Murthy, Raghotham},
        journal = {PVLDB},
        series = {{VLDB} '09},
        volume = {2},
        number = {2},
        pages = {1626--1629},
        doi = {10.14778/1687553.1687609},
        url = {https://doi.org/10.14778/1687553.1687609},
        year = {2009}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 136 citing papers.

Rank Citing Paper Year Venue Pagerank
148 Gorilla: A Fast, Scalable, In-Memory Time Series Database 2015 VLDB 0.00029001141
178 The Vertica Analytic Database: C-Store 7 Years Later 2012 VLDB 0.00026620521
282 Accelerating Machine Learning Inference with Probabilistic Predicates 2018 SIGMOD 0.00022302793
325 The Dataflow Model: A Practical Approach to Balancing Correctness, Latency, and Cost in Massive-Scale, Unbounded, Out-of-Order Data Processing 2015 VLDB 0.0002095522
379 Apache Calcite: A Foundational Framework for Optimized Query Processing Over Heterogeneous Data Sources 2018 SIGMOD 0.00019507406
394 One Trillion Edges: Graph Processing at Facebook-Scale 2015 VLDB 0.00019182322
675 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.00014880686
823 MRShare: Sharing Across Multiple Queries in MapReduce 2010 VLDB 0.00013642066
841 Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters 2016 SIGMOD 0.00013543
876 Tenzing: A SQL Implementation On The MapReduce Framework 2011 VLDB 0.00013304112
978 Parallel Evaluation of Conjunctive Queries 2011 PODS 0.00012725823
1,009 Simulation of Database-Valued Markov Chains Using SimSQL 2013 SIGMOD 0.0001253551
1,036 Fine-grained Partitioning for Aggressive Data Skipping 2014 SIGMOD 0.00012372946
1,037 Jaql: A Scripting Language for Large Scale Semistructured Data Analysis 2011 VLDB 0.00012371105
1,061 Approximate Query Processing: No Silver Bullet 2017 SIGMOD 0.00012208639
1,073 Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads 2012 VLDB 0.00012163258
1,081 Hadoop-GIS: A High Performance Spatial Data Warehousing System over MapReduce 2013 VLDB 0.00012126286
1,082 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012118261
1,178 Simba: Efficient In-Memory Spatial Analytics 2016 SIGMOD 0.00011627256
1,233 Communication Steps for Parallel Query Processing 2013 PODS 0.00011404978
1,371 Ricardo: Integrating R and Hadoop 2010 SIGMOD 0.00010894415
1,381 Scuba: Diving into Data at Facebook 2013 VLDB 0.00010856731
1,389 Provenance for Generalized Map and Reduce Workflows 2011 CIDR 0.00010818511
1,421 V-SMART-Join: A Scalable MapReduce Framework for All-Pair Similarity Joins of Multisets and Vectors 2012 VLDB 0.00010722146
1,428 Knowing When You’re Wrong: Building Fast and Reliable Approximate Query Processing Systems 2014 SIGMOD 0.0001069161
1,432 Towards a Learning Optimizer for Shared Clouds 2019 VLDB 0.00010676754
1,456 Manu: A Cloud Native Vector Database Management System 2022 VLDB 0.00010588868
1,466 The Performance of MapReduce: An In-depth Study 2010 VLDB 0.00010565007
1,512 Cloud-Native Database Systems at Alibaba: Opportunities and Challenges 2019 VLDB 0.00010425349
1,564 Titian: Data Provenance Support in Spark 2016 VLDB 0.00010219222
1,623 Mesa: Geo-Replicated, Near Real-Time, Scalable Data Warehousing 2014 VLDB 0.00010047769
1,747 Selecting Subexpressions to Materialize at Datacenter Scale 2018 VLDB 9.7303647e-05
1,771 Samza: Stateful Scalable Stream Processing at LinkedIn 2017 VLDB 9.6757926e-05
1,828 DBEst: Revisiting Approximate Query Processing Engines with Machine Learning Models 2019 SIGMOD 9.547768e-05
1,838 Cheetah: A High Performance, Custom Data Warehouse on Top of MapReduce 2010 VLDB 9.5305061e-05
1,925 ReStore: Reusing Results of MapReduce Jobs 2012 VLDB 9.3643089e-05
1,933 Instant Loading for Main Memory Databases 2013 VLDB 9.3430609e-05
2,197 Opening the Black Boxes in Data Flow Optimization 2012 VLDB 8.8740089e-05
2,249 Cumulon: Optimizing Statistical Data Analysis in the Cloud 2013 SIGMOD 8.7544468e-05
2,285 A Platform for Scalable One-Pass Analytics using MapReduce 2011 SIGMOD 8.6924793e-05
2,362 Online Aggregation and Continuous Query support in MapReduce 2010 SIGMOD 8.5729053e-05
2,411 epiC: an Extensible and Scalable System for Processing Big Data 2014 VLDB 8.5073756e-05
2,424 BigDansing: A System for Big Data Cleansing 2015 SIGMOD 8.483813e-05
2,483 AnalyticDB: Real-time OLAP Database System at Alibaba Cloud 2019 VLDB 8.3973995e-05
2,499 Towards Scalable Real-time Analytics: An Architecture for Scale-out of OLxP Workloads 2015 VLDB 8.3795899e-05
2,549 SQLShare: Results from a Multi-Year SQL-as-a-Service Experiment 2016 SIGMOD 8.3096391e-05
2,636 Big Data Analytics with Datalog Queries on Spark 2016 SIGMOD 8.1926426e-05
2,688 CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop 2011 VLDB 8.1262948e-05
2,753 Implicit Parallelism through Deep Language Embedding 2015 SIGMOD 8.0513412e-05
2,898 HAWQ: A Massively Parallel Processing SQL Engine in Hadoop 2014 SIGMOD 7.8816694e-05
Previous Page 1 / 3 Next

Outgoing Citations (Sorted by Pagerank)

Showing 2 of 2 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
30 SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets 2008 VLDB 0.00050475202
43 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.0004552807
Previous Page 1 / 1 Next

Semantically Similar Papers