DBScholar

Back to papers

HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads

Summary: HadoopDB hybridizes MapReduce’s scalability, fault tolerance, and data flexibility with parallel-DBMS query processing and storage. Its prototype targets commodity shared-nothing/cloud clusters, approaching DBMS performance without sacrificing MapReduce robustness. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
10148
Venue
VLDB
Year
2009
Pagerank
0.00031680027
Overall Rank
120 | 99.18%
DOI
10.14778/1687627.1687731

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{abouzeid_vldb09,
        title = {{HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads}},
        author = {Abouzeid, Azza and Bajda-Pawlikowski, Kamil and Abadi, Daniel and Silberschatz, Avi and Rasin, Alexander},
        journal = {PVLDB},
        series = {{VLDB} '09},
        doi = {10.14778/1687627.1687731},
        url = {https://doi.org/10.14778/1687627.1687731},
        year = {2009}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 72 citing papers.

Rank Citing Paper Year Venue Pagerank
51 Dremel: Interactive Analysis of Web-Scale Datasets 2010 VLDB 0.0004291425
372 HaLoop: Efficient Iterative Data Processing on Large Clusters 2010 VLDB 0.0001981521
425 Shark: SQL and Rich Analytics at Scale 2013 SIGMOD 0.00018704491
500 Scalable SPARQL Querying of Large RDF Graphs 2011 VLDB 0.00017413839
660 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.00015198804
735 Profiling, What-if Analysis, and Cost-based Optimization of MapReduce Programs 2011 VLDB 0.00014522606
803 MRShare: Sharing Across Multiple Queries in MapReduce 2010 VLDB 0.00013899943
872 Tenzing: A SQL Implementation On The MapReduce Framework 2011 VLDB 0.00013486409
923 Starfish: A Self-tuning System for Big Data Analytics 2011 CIDR 0.00013189886
993 Simulation of Database-Valued Markov Chains Using SimSQL 2013 SIGMOD 0.00012789598
1,067 Hadoop-GIS: A High Performance Spatial Data Warehousing System over MapReduce 2013 VLDB 0.00012327784
1,257 Automatic Optimization for MapReduce Programs 2011 VLDB 0.000114432
1,320 From Theory to Practice: Efficient Join Query Evaluation in a Parallel Database System 2015 SIGMOD 0.00011166426
1,436 The Performance of MapReduce: An In-depth Study 2010 VLDB 0.00010797443
1,606 Mesa: Geo-Replicated, Near Real-Time, Scalable Data Warehousing 2014 VLDB 0.00010228576
1,791 Cheetah: A High Performance, Custom Data Warehouse on Top of MapReduce 2010 VLDB 9.7470504e-05
1,852 Automated Partitioning Design in Parallel Database Systems 2011 SIGMOD 9.6134443e-05
1,889 SQL-on-Hadoop: Full Circle Back to Shared-Nothing Database Architectures 2014 VLDB 9.5335988e-05
1,903 Instant Loading for Main Memory Databases 2013 VLDB 9.5049156e-05
2,159 Efficient Processing of Data Warehousing Queries in a Split Execution Environment 2011 SIGMOD 9.061086e-05
2,354 epiC: an Extensible and Scalable System for Processing Big Data 2014 VLDB 8.7060612e-05
2,454 REX: Recursive, Delta-Based Data-Centric Computation 2012 VLDB 8.5560058e-05
2,539 Minimal MapReduce Algorithms 2013 SIGMOD 8.4526595e-05
2,559 dbTouch: Analytics at your Fingertips 2013 CIDR 8.4156672e-05
2,642 CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop 2011 VLDB 8.3059948e-05
2,849 Column-Oriented Storage Techniques for MapReduce 2011 VLDB 8.053191e-05
2,850 HAWQ: A Massively Parallel Processing SQL Engine in Hadoop 2014 SIGMOD 8.0499212e-05
2,942 Llama: Leveraging Columnar Storage for Scalable Join Processing in the MapReduce Framework 2011 SIGMOD 7.9358593e-05
3,199 Query Optimization Techniques for Partitioned Tables 2011 SIGMOD 7.6423984e-05
3,202 How to Fit when No One Size Fits 2013 CIDR 7.6409287e-05
3,275 Optimizing Analytic Data Flows for Multiple Execution Engines 2012 SIGMOD 7.5747814e-05
3,618 The Myria Big Data Management and Analytics System and Cloud Service 2017 CIDR 7.2523695e-05
3,638 Fast Queries Over Heterogeneous Data Through Engine Customization 2016 VLDB 7.2338361e-05
3,714 Large-Scale Machine Learning at Twitter 2012 SIGMOD 7.1764857e-05
3,749 Integrating Hadoop and Parallel DBMS 2010 SIGMOD 7.1539955e-05
3,783 MISO: Souping Up Big Data Query Processing with a Multistore System 2014 SIGMOD 7.1298683e-05
3,964 Hyper Dimension Shuffle: Efficient Data Repartition at Petabyte Scale in SCOPE 2019 VLDB 6.9855158e-05
4,078 Asynchronous and Fault-Tolerant Recursive Datalog Evaluation in Shared-Nothing Engines 2015 VLDB 6.9209348e-05
4,079 Flexible Rule-Based Decomposition and Metadata Independence in Modin: A Parallel Dataframe System 2022 VLDB 6.9188052e-05
4,169 The Unified Logging Infrastructure for Data Analytics at Twitter 2012 VLDB 6.8561463e-05
4,908 Using Cloud Functions as Accelerator for Elastic Data Analytics 2023 SIGMOD 6.4488784e-05
5,522 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 6.1871697e-05
5,608 Big Graphs: Challenges and Opportunities 2022 VLDB 6.1514145e-05
5,935 Towards Energy-Efficient Database Cluster Design 2012 VLDB 6.0361286e-05
6,147 HadoopDB in Action: Building Real World Applications 2010 SIGMOD 5.9601328e-05
6,199 Clydesdale: Structured Data Processing on Hadoop 2012 SIGMOD 5.9458654e-05
6,447 Replication at the Speed of Change - a Fast, Scalable Replication Solution for Near Real-Time HTAP Processing 2020 VLDB 5.8761879e-05
6,888 The Next Generation Operational Data Historian for IoT Based on Informix 2014 SIGMOD 5.7455844e-05
7,168 Optimization for iterative queries on MapReduce 2014 VLDB 5.6841364e-05
7,471 Are We Experiencing a Big Data Bubble? 2014 SIGMOD 5.6095411e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 7 of 7 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers