DBScholar

Back to papers

HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads

Summary: HadoopDB hybridizes MapReduce’s scalability, fault tolerance, and data flexibility with parallel-DBMS query processing and storage. Its prototype targets commodity shared-nothing/cloud clusters, approaching DBMS performance without sacrificing MapReduce robustness. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
hf795d89d038f12ee
Venue
VLDB
Year
2009
Pagerank
0.000311132
Overall Rank
120 | 99.20%
DOI
10.14778/1687627.1687731

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{abouzeid_vldb09,
        title = {{HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads}},
        author = {Abouzeid, Azza and Bajda-Pawlikowski, Kamil and Abadi, Daniel and Silberschatz, Avi and Rasin, Alexander},
        journal = {PVLDB},
        series = {{VLDB} '09},
        doi = {10.14778/1687627.1687731},
        url = {https://doi.org/10.14778/1687627.1687731},
        year = {2009}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 72 citing papers.

Rank Citing Paper Year Venue Pagerank
49 Dremel: Interactive Analysis of Web-Scale Datasets 2010 VLDB 0.00043160717
384 HaLoop: Efficient Iterative Data Processing on Large Clusters 2010 VLDB 0.0001948031
432 Shark: SQL and Rich Analytics at Scale 2013 SIGMOD 0.00018339357
511 Scalable SPARQL Querying of Large RDF Graphs 2011 VLDB 0.00017061883
673 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.0001488755
753 Profiling, What-if Analysis, and Cost-based Optimization of MapReduce Programs 2011 VLDB 0.00014237583
823 MRShare: Sharing Across Multiple Queries in MapReduce 2010 VLDB 0.00013648332
876 Tenzing: A SQL Implementation On The MapReduce Framework 2011 VLDB 0.00013309176
940 Starfish: A Self-tuning System for Big Data Analytics 2011 CIDR 0.00012964445
1,009 Simulation of Database-Valued Markov Chains Using SimSQL 2013 SIGMOD 0.00012539827
1,080 Hadoop-GIS: A High Performance Spatial Data Warehousing System over MapReduce 2013 VLDB 0.00012131978
1,281 Automatic Optimization for MapReduce Programs 2011 VLDB 0.00011213384
1,292 From Theory to Practice: Efficient Join Query Evaluation in a Parallel Database System 2015 SIGMOD 0.00011152286
1,466 The Performance of MapReduce: An In-depth Study 2010 VLDB 0.00010569837
1,622 Mesa: Geo-Replicated, Near Real-Time, Scalable Data Warehousing 2014 VLDB 0.00010051462
1,837 Cheetah: A High Performance, Custom Data Warehouse on Top of MapReduce 2010 VLDB 9.5350058e-05
1,883 Automated Partitioning Design in Parallel Database Systems 2011 SIGMOD 9.4391795e-05
1,884 SQL-on-Hadoop: Full Circle Back to Shared-Nothing Database Architectures 2014 VLDB 9.4388097e-05
1,931 Instant Loading for Main Memory Databases 2013 VLDB 9.3474242e-05
2,173 Efficient Processing of Data Warehousing Queries in a Split Execution Environment 2011 SIGMOD 8.91924e-05
2,410 epiC: an Extensible and Scalable System for Processing Big Data 2014 VLDB 8.5114032e-05
2,480 REX: Recursive, Delta-Based Data-Centric Computation 2012 VLDB 8.4035081e-05
2,573 Minimal MapReduce Algorithms 2013 SIGMOD 8.2821647e-05
2,605 dbTouch: Analytics at your Fingertips 2013 CIDR 8.229938e-05
2,687 CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop 2011 VLDB 8.1301312e-05
2,898 HAWQ: A Massively Parallel Processing SQL Engine in Hadoop 2014 SIGMOD 7.8853204e-05
2,903 Column-Oriented Storage Techniques for MapReduce 2011 VLDB 7.8810111e-05
3,007 Llama: Leveraging Columnar Storage for Scalable Join Processing in the MapReduce Framework 2011 SIGMOD 7.7607173e-05
3,143 Query Optimization Techniques for Partitioned Tables 2011 SIGMOD 7.5981026e-05
3,258 How to Fit when No One Size Fits 2013 CIDR 7.4873069e-05
3,339 Optimizing Analytic Data Flows for Multiple Execution Engines 2012 SIGMOD 7.4080114e-05
3,505 Fast Queries Over Heterogeneous Data Through Engine Customization 2016 VLDB 7.2508161e-05
3,685 The Myria Big Data Management and Analytics System and Cloud Service 2017 CIDR 7.0988339e-05
3,788 Large-Scale Machine Learning at Twitter 2012 SIGMOD 7.0195464e-05
3,823 Integrating Hadoop and Parallel DBMS 2010 SIGMOD 7.0029055e-05
3,840 MISO: Souping Up Big Data Query Processing with a Multistore System 2014 SIGMOD 6.9920793e-05
3,895 Hyper Dimension Shuffle: Efficient Data Repartition at Petabyte Scale in SCOPE 2019 VLDB 6.9393157e-05
3,930 Flexible Rule-Based Decomposition and Metadata Independence in Modin: A Parallel Dataframe System 2022 VLDB 6.9185978e-05
4,154 Asynchronous and Fault-Tolerant Recursive Datalog Evaluation in Shared-Nothing Engines 2015 VLDB 6.776227e-05
4,249 The Unified Logging Infrastructure for Data Analytics at Twitter 2012 VLDB 6.7037533e-05
5,005 Using Cloud Functions as Accelerator for Elastic Data Analytics 2023 SIGMOD 6.3177105e-05
5,513 Big Graphs: Challenges and Opportunities 2022 VLDB 6.0986544e-05
5,652 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 6.0515502e-05
6,057 Towards Energy-Efficient Database Cluster Design 2012 VLDB 5.900701e-05
6,276 HadoopDB in Action: Building Real World Applications 2010 SIGMOD 5.8268854e-05
6,322 Clydesdale: Structured Data Processing on Hadoop 2012 SIGMOD 5.8148318e-05
6,558 Replication at the Speed of Change - a Fast, Scalable Replication Solution for Near Real-Time HTAP Processing 2020 VLDB 5.7497828e-05
7,031 The Next Generation Operational Data Historian for IoT Based on Informix 2014 SIGMOD 5.6172818e-05
7,308 Optimization for iterative queries on MapReduce 2014 VLDB 5.5572894e-05
7,611 Are We Experiencing a Big Data Bubble? 2014 SIGMOD 5.4843914e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 7 of 7 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers