DBScholar

Back to papers

Hive - A Warehousing Solution Over a Map-Reduce Framework

Summary: Hive brings SQL-like declarative warehousing to Hadoop by compiling HiveQL into MapReduce, while retaining extensibility via custom scripts and data formats. Its rich nested type system and catalog of schemas/statistics enable large-scale exploration and optimization. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
hcc7c6e4b04247911
Venue
VLDB
Year
2009
Pagerank
0.00049839909
Overall Rank
31 | 99.80%
DOI
10.14778/1687553.1687609

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{thusoo_vldb09,
        title = {{Hive - A Warehousing Solution Over a Map-Reduce Framework}},
        author = {Thusoo, Ashish and Sarma, Joydeep Sen and Jain, Namit and Shao, Zheng and Chakka, Prasad and Anthony, Suresh and Liu, Hao and Wyckoff, Pete and Murthy, Raghotham},
        journal = {PVLDB},
        series = {{VLDB} '09},
        volume = {2},
        number = {2},
        pages = {1626--1629},
        doi = {10.14778/1687553.1687609},
        url = {https://doi.org/10.14778/1687553.1687609},
        year = {2009}
}

Incoming Citations (Sorted by Pagerank)

Showing 36 of 136 citing papers.

Rank Citing Paper Year Venue Pagerank
7,836 Quill: Efficient, Transferable, and Rich Analytics at Scale 2016 VLDB 5.4435399e-05
7,843 Oracle In-Database Hadoop: When MapReduce Meets RDBMS 2012 SIGMOD 5.4424628e-05
7,925 Runtime Variation in Big Data Analytics 2023 SIGMOD 5.4256585e-05
8,004 A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning 2024 VLDB 5.4089097e-05
8,535 New Query Optimization Techniques in the Spark Engine of Azure Synapse 2022 VLDB 5.320973e-05
8,552 F3: The Open-Source Data File Format for the Future 2026 SIGMOD 5.316708e-05
8,624 Piranha: Optimizing Short Jobs in Hadoop 2013 VLDB 5.3009314e-05
9,088 Hyperspace: The Indexing Subsystem of Azure Synapse 2021 VLDB 5.2283159e-05
9,113 Presto’s History-based Query Optimizer 2024 VLDB 5.2276066e-05
9,254 QMapper for Smart Grid: Migrating SQL-based Application to Hive 2015 SIGMOD 5.2056825e-05
9,358 Phoebe: A Learning-based Checkpoint Optimizer 2021 VLDB 5.1869771e-05
9,417 Towards Observability for Production Machine Learning Pipelines 2022 VLDB 5.1803615e-05
9,681 Rank Join Queries in NoSQL Databases 2014 VLDB 5.1427987e-05
9,692 Versatile Optimization of UDF-heavy Data Flows with Sofa 2014 SIGMOD 5.140168e-05
9,764 Sapprox: Enabling Efficient and Accurate Approximations on Sub-datasets with Distribution-aware Online Sampling 2017 VLDB 5.1344728e-05
10,015 GHive: A Demonstration of GPU-Accelerated Query Processing in Apache Hive 2022 SIGMOD 5.0961867e-05
10,230 OceanRT: Real-Time Analytics over Large Temporal Data 2014 SIGMOD 5.0570085e-05
10,672 PTO: A Workload-driven Predictive Table Optimizer for Lakehouse Systems 2026 SIGMOD 4.9793485e-05
10,887 MFTune: An Efficient Multi-fidelity Framework for Spark SQL Configuration Tuning 2026 VLDB 4.9793485e-05
10,915 Reaching the Pinnacle of TPC-DS: Co-design of Architecture, Executor, and Storage in TDSQL 2026 VLDB 4.9793485e-05
10,940 Bridging NL2SQL for CAN Signal Analytics at NIO: From Externalized Schemas to CTE Pipelines 2026 VLDB 4.9793485e-05
11,130 OpenMLDB: A Real-Time Relational Data Feature Computation System for Online ML 2025 SIGMOD 4.9793485e-05
11,242 QOVIS: Understanding and Diagnosing Query Optimizer via a Visualization-assisted Approach 2025 VLDB 4.9793485e-05
11,288 Concurrency Control as a Service 2025 VLDB 4.9793485e-05
11,306 ArrayMorph: Optimizing Hyperslab Queries on the Cloud for Machine Learning Pipelines 2025 VLDB 4.9793485e-05
11,895 CDI-E: An Elastic Cloud Service for Data Engineering 2022 VLDB 4.9793485e-05
12,185 Integration of Large-Scale Data Processing Systems and Traditional Parallel Database Technology 2019 VLDB 4.9793485e-05
12,327 Logical Aspects of Massively Parallel and Distributed Systems 2016 PODS 4.9793485e-05
12,353 dmapply: A functional primitive to express distributed machine learning algorithms in R 2016 VLDB 4.9793485e-05
12,383 Let's Rethink Join Optimization in Distributed Systems 2015 CIDR 4.9793485e-05
12,408 A Demonstration of Rubato DB: A Highly Scalable NewSQL Database System for OLTP and Big Data Applications 2015 SIGMOD 4.9793485e-05
12,465 Anti-Combining for MapReduce 2014 SIGMOD 4.9793485e-05
12,488 Getting Your Big Data Priorities Straight: A Demonstration of Priority-based QoS using Social-network-driven Stock Recommendation 2014 VLDB 4.9793485e-05
12,494 Design and Implementation of a Real-Time Interactive Analytics System for Large Spatio-Temporal Data 2014 VLDB 4.9793485e-05
12,597 Declarative Error Management for Robust Data-Intensive Applications 2012 SIGMOD 4.9793485e-05
12,689 Resiliency-Aware Data Management 2011 VLDB 4.9793485e-05
Previous Page 3 / 3 Next

Outgoing Citations (Sorted by Pagerank)

Showing 2 of 2 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
30 SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets 2008 VLDB 0.00050495102
43 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.00045546775
Previous Page 1 / 1 Next

Semantically Similar Papers