DBScholar

Back to papers

Hive - A Warehousing Solution Over a Map-Reduce Framework

Summary: Hive brings SQL-like declarative warehousing to Hadoop by compiling HiveQL into MapReduce, while retaining extensibility via custom scripts and data formats. Its rich nested type system and catalog of schemas/statistics enable large-scale exploration and optimization. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
10153
Venue
VLDB
Year
2009
Pagerank
0.00050111008
Overall Rank
32 | 99.79%
DOI
10.14778/1687553.1687609

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{thusoo_vldb09,
        title = {{Hive - A Warehousing Solution Over a Map-Reduce Framework}},
        author = {Thusoo, Ashish and Sarma, Joydeep Sen and Jain, Namit and Shao, Zheng and Chakka, Prasad and Anthony, Suresh and Liu, Hao and Wyckoff, Pete and Murthy, Raghotham},
        journal = {PVLDB},
        series = {{VLDB} '09},
        volume = {2},
        number = {2},
        pages = {1626--1629},
        doi = {10.14778/1687553.1687609},
        url = {https://doi.org/10.14778/1687553.1687609},
        year = {2009}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 133 citing papers.

Rank Citing Paper Year Venue Pagerank
148 Gorilla: A Fast, Scalable, In-Memory Time Series Database 2015 VLDB 0.00029250767
186 The Vertica Analytic Database: C-Store 7 Years Later 2012 VLDB 0.00026182534
295 Accelerating Machine Learning Inference with Probabilistic Predicates 2018 SIGMOD 0.00022238183
361 The Dataflow Model: A Practical Approach to Balancing Correctness, Latency, and Cost in Massive-Scale, Unbounded, Out-of-Order Data Processing 2015 VLDB 0.00020138717
389 One Trillion Edges: Graph Processing at Facebook-Scale 2015 VLDB 0.00019386526
445 Apache Calcite: A Foundational Framework for Optimized Query Processing Over Heterogeneous Data Sources 2018 SIGMOD 0.00018336751
660 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.00015198804
803 MRShare: Sharing Across Multiple Queries in MapReduce 2010 VLDB 0.00013899943
819 Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters 2016 SIGMOD 0.00013815639
872 Tenzing: A SQL Implementation On The MapReduce Framework 2011 VLDB 0.00013486409
954 Parallel Evaluation of Conjunctive Queries 2011 PODS 0.00012997301
993 Simulation of Database-Valued Markov Chains Using SimSQL 2013 SIGMOD 0.00012789598
1,021 Jaql: A Scripting Language for Large Scale Semistructured Data Analysis 2011 VLDB 0.00012606673
1,044 Fine-grained Partitioning for Aggressive Data Skipping 2014 SIGMOD 0.0001244236
1,054 Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads 2012 VLDB 0.00012390673
1,067 Hadoop-GIS: A High Performance Spatial Data Warehousing System over MapReduce 2013 VLDB 0.00012327784
1,079 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012258469
1,108 Approximate Query Processing: No Silver Bullet 2017 SIGMOD 0.00012145154
1,175 Simba: Efficient In-Memory Spatial Analytics 2016 SIGMOD 0.00011812263
1,207 Communication Steps for Parallel Query Processing 2013 PODS 0.00011663155
1,333 Ricardo: Integrating R and Hadoop 2010 SIGMOD 0.0001112858
1,352 Scuba: Diving into Data at Facebook 2013 VLDB 0.00011064595
1,366 Provenance for Generalized Map and Reduce Workflows 2011 CIDR 0.00011016972
1,401 Knowing When You’re Wrong: Building Fast and Reliable Approximate Query Processing Systems 2014 SIGMOD 0.00010889902
1,415 V-SMART-Join: A Scalable MapReduce Framework for All-Pair Similarity Joins of Multisets and Vectors 2012 VLDB 0.00010840141
1,436 The Performance of MapReduce: An In-depth Study 2010 VLDB 0.00010797443
1,468 Towards a Learning Optimizer for Shared Clouds 2019 VLDB 0.00010686496
1,559 Cloud-Native Database Systems at Alibaba: Opportunities and Challenges 2019 VLDB 0.00010362456
1,589 Manu: A Cloud Native Vector Database Management System 2022 VLDB 0.00010264469
1,606 Mesa: Geo-Replicated, Near Real-Time, Scalable Data Warehousing 2014 VLDB 0.00010228576
1,617 Titian: Data Provenance Support in Spark 2016 VLDB 0.00010209397
1,765 Selecting Subexpressions to Materialize at Datacenter Scale 2018 VLDB 9.8079546e-05
1,791 Cheetah: A High Performance, Custom Data Warehouse on Top of MapReduce 2010 VLDB 9.7470504e-05
1,799 DBEst: Revisiting Approximate Query Processing Engines with Machine Learning Models 2019 SIGMOD 9.7326398e-05
1,883 ReStore: Reusing Results of MapReduce Jobs 2012 VLDB 9.5421713e-05
1,895 Samza: Stateful Scalable Stream Processing at LinkedIn 2017 VLDB 9.5260291e-05
1,903 Instant Loading for Main Memory Databases 2013 VLDB 9.5049156e-05
2,164 Opening the Black Boxes in Data Flow Optimization 2012 VLDB 9.0521951e-05
2,217 Cumulon: Optimizing Statistical Data Analysis in the Cloud 2013 SIGMOD 8.9332438e-05
2,265 A Platform for Scalable One-Pass Analytics using MapReduce 2011 SIGMOD 8.8398946e-05
2,312 Online Aggregation and Continuous Query support in MapReduce 2010 SIGMOD 8.7642158e-05
2,354 epiC: an Extensible and Scalable System for Processing Big Data 2014 VLDB 8.7060612e-05
2,398 BigDansing: A System for Big Data Cleansing 2015 SIGMOD 8.631172e-05
2,563 AnalyticDB: Real-time OLAP Database System at Alibaba Cloud 2019 VLDB 8.412445e-05
2,583 Towards Scalable Real-time Analytics: An Architecture for Scale-out of OLxP Workloads 2015 VLDB 8.3849987e-05
2,592 SQLShare: Results from a Multi-Year SQL-as-a-Service Experiment 2016 SIGMOD 8.3657526e-05
2,594 Big Data Analytics with Datalog Queries on Spark 2016 SIGMOD 8.3646367e-05
2,642 CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop 2011 VLDB 8.3059948e-05
2,717 Implicit Parallelism through Deep Language Embedding 2015 SIGMOD 8.2102313e-05
2,850 HAWQ: A Massively Parallel Processing SQL Engine in Hadoop 2014 SIGMOD 8.0499212e-05
Previous Page 1 / 3 Next

Outgoing Citations (Sorted by Pagerank)

Showing 2 of 2 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
30 SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets 2008 VLDB 0.00051174276
44 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.00046055057
Previous Page 1 / 1 Next

Semantically Similar Papers