DBScholar

Back to papers

HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads

Summary: HadoopDB hybridizes MapReduce’s scalability, fault tolerance, and data flexibility with parallel-DBMS query processing and storage. Its prototype targets commodity shared-nothing/cloud clusters, approaching DBMS performance without sacrificing MapReduce robustness. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
hf795d89d038f12ee
Venue
VLDB
Year
2009
Pagerank
0.00031099083
Overall Rank
120 | 99.20%
DOI
10.14778/1687627.1687731

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{abouzeid_vldb09,
        title = {{HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads}},
        author = {Abouzeid, Azza and Bajda-Pawlikowski, Kamil and Abadi, Daniel and Silberschatz, Avi and Rasin, Alexander},
        journal = {PVLDB},
        series = {{VLDB} '09},
        doi = {10.14778/1687627.1687731},
        url = {https://doi.org/10.14778/1687627.1687731},
        year = {2009}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 72 citing papers.

Rank Citing Paper Year Venue Pagerank
49 Dremel: Interactive Analysis of Web-Scale Datasets 2010 VLDB 0.0004314366
384 HaLoop: Efficient Iterative Data Processing on Large Clusters 2010 VLDB 0.00019471648
432 Shark: SQL and Rich Analytics at Scale 2013 SIGMOD 0.00018331051
512 Scalable SPARQL Querying of Large RDF Graphs 2011 VLDB 0.00017053842
675 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.00014880686
754 Profiling, What-if Analysis, and Cost-based Optimization of MapReduce Programs 2011 VLDB 0.00014231311
823 MRShare: Sharing Across Multiple Queries in MapReduce 2010 VLDB 0.00013642066
876 Tenzing: A SQL Implementation On The MapReduce Framework 2011 VLDB 0.00013304112
940 Starfish: A Self-tuning System for Big Data Analytics 2011 CIDR 0.00012959992
1,009 Simulation of Database-Valued Markov Chains Using SimSQL 2013 SIGMOD 0.0001253551
1,081 Hadoop-GIS: A High Performance Spatial Data Warehousing System over MapReduce 2013 VLDB 0.00012126286
1,282 Automatic Optimization for MapReduce Programs 2011 VLDB 0.00011208192
1,292 From Theory to Practice: Efficient Join Query Evaluation in a Parallel Database System 2015 SIGMOD 0.00011147959
1,466 The Performance of MapReduce: An In-depth Study 2010 VLDB 0.00010565007
1,623 Mesa: Geo-Replicated, Near Real-Time, Scalable Data Warehousing 2014 VLDB 0.00010047769
1,838 Cheetah: A High Performance, Custom Data Warehouse on Top of MapReduce 2010 VLDB 9.5305061e-05
1,884 Automated Partitioning Design in Parallel Database Systems 2011 SIGMOD 9.4349703e-05
1,885 SQL-on-Hadoop: Full Circle Back to Shared-Nothing Database Architectures 2014 VLDB 9.4345374e-05
1,933 Instant Loading for Main Memory Databases 2013 VLDB 9.3430609e-05
2,176 Efficient Processing of Data Warehousing Queries in a Split Execution Environment 2011 SIGMOD 8.9150466e-05
2,411 epiC: an Extensible and Scalable System for Processing Big Data 2014 VLDB 8.5073756e-05
2,481 REX: Recursive, Delta-Based Data-Centric Computation 2012 VLDB 8.3995962e-05
2,573 Minimal MapReduce Algorithms 2013 SIGMOD 8.2782871e-05
2,607 dbTouch: Analytics at your Fingertips 2013 CIDR 8.2260445e-05
2,688 CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop 2011 VLDB 8.1262948e-05
2,898 HAWQ: A Massively Parallel Processing SQL Engine in Hadoop 2014 SIGMOD 7.8816694e-05
2,903 Column-Oriented Storage Techniques for MapReduce 2011 VLDB 7.877293e-05
3,009 Llama: Leveraging Columnar Storage for Scalable Join Processing in the MapReduce Framework 2011 SIGMOD 7.7570485e-05
3,145 Query Optimization Techniques for Partitioned Tables 2011 SIGMOD 7.594509e-05
3,258 How to Fit when No One Size Fits 2013 CIDR 7.4839611e-05
3,339 Optimizing Analytic Data Flows for Multiple Execution Engines 2012 SIGMOD 7.4045097e-05
3,505 Fast Queries Over Heterogeneous Data Through Engine Customization 2016 VLDB 7.2474175e-05
3,687 The Myria Big Data Management and Analytics System and Cloud Service 2017 CIDR 7.0954736e-05
3,790 Large-Scale Machine Learning at Twitter 2012 SIGMOD 7.0162259e-05
3,824 Integrating Hadoop and Parallel DBMS 2010 SIGMOD 6.9996014e-05
3,841 MISO: Souping Up Big Data Query Processing with a Multistore System 2014 SIGMOD 6.98877e-05
3,895 Hyper Dimension Shuffle: Efficient Data Repartition at Petabyte Scale in SCOPE 2019 VLDB 6.9360458e-05
3,931 Flexible Rule-Based Decomposition and Metadata Independence in Modin: A Parallel Dataframe System 2022 VLDB 6.9153238e-05
4,153 Asynchronous and Fault-Tolerant Recursive Datalog Evaluation in Shared-Nothing Engines 2015 VLDB 6.7730626e-05
4,250 The Unified Logging Infrastructure for Data Analytics at Twitter 2012 VLDB 6.7005978e-05
5,008 Using Cloud Functions as Accelerator for Elastic Data Analytics 2023 SIGMOD 6.3147226e-05
5,516 Big Graphs: Challenges and Opportunities 2022 VLDB 6.0957673e-05
5,653 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 6.048773e-05
6,058 Towards Energy-Efficient Database Cluster Design 2012 VLDB 5.8979077e-05
6,279 HadoopDB in Action: Building Real World Applications 2010 SIGMOD 5.8241273e-05
6,326 Clydesdale: Structured Data Processing on Hadoop 2012 SIGMOD 5.8120855e-05
6,560 Replication at the Speed of Change - a Fast, Scalable Replication Solution for Near Real-Time HTAP Processing 2020 VLDB 5.747061e-05
7,033 The Next Generation Operational Data Historian for IoT Based on Informix 2014 SIGMOD 5.6147004e-05
7,311 Optimization for iterative queries on MapReduce 2014 VLDB 5.5546605e-05
7,617 Are We Experiencing a Big Data Bubble? 2014 SIGMOD 5.4817965e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 7 of 7 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers