DBScholar

Back to papers

Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing)

Summary: Hadoop++ transparently accelerates Hadoop by injecting indexing and join optimizations through UDFs, without modifying its framework or interface. It substantially outperforms Hadoop and HadoopDB while remaining compatible with future Hadoop changes. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
h36bb8849cf7e7883
Venue
VLDB
Year
2010
Pagerank
0.00014880686
Overall Rank
675 | 95.47%
DOI
10.14778/1920841.1920908

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{dittrich_vldb10,
        title = {{Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing)}},
        author = {Dittrich, Jens and Kargin, Yagiz and Quiané-Ruiz, Jorge-Arnulfo and Setty, Vinay and Jindal, Alekh and Schad, Jörg},
        journal = {PVLDB},
        series = {{VLDB} '10},
        volume = {3},
        number = {1},
        pages = {518--529},
        doi = {10.14778/1920841.1920908},
        url = {https://doi.org/10.14778/1920841.1920908},
        year = {2010}
}

Incoming Citations (Sorted by Pagerank)

Showing 39 of 39 citing papers.

Rank Citing Paper Year Venue Pagerank
754 Profiling, What-if Analysis, and Cost-based Optimization of MapReduce Programs 2011 VLDB 0.00014231311
940 Starfish: A Self-tuning System for Big Data Analytics 2011 CIDR 0.00012959992
1,073 Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads 2012 VLDB 0.00012163258
1,282 Automatic Optimization for MapReduce Programs 2011 VLDB 0.00011208192
1,351 SkewTune: Mitigating Skew in MapReduce Applications 2012 SIGMOD 0.00010929229
1,708 PerfXplain: Debugging MapReduce Job Performance 2012 VLDB 9.8184303e-05
2,176 Efficient Processing of Data Warehousing Queries in a Split Execution Environment 2011 SIGMOD 8.9150466e-05
2,343 TriAD: A Distributed Shared-Nothing RDF Engine based on Asynchronous Message Passing 2014 SIGMOD 8.6030486e-05
2,573 Minimal MapReduce Algorithms 2013 SIGMOD 8.2782871e-05
2,688 CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop 2011 VLDB 8.1262948e-05
2,903 Column-Oriented Storage Techniques for MapReduce 2011 VLDB 7.877293e-05
2,932 Efficient Multi-way Theta-Join Processing Using MapReduce 2012 VLDB 7.8368679e-05
3,009 Llama: Leveraging Columnar Storage for Scalable Join Processing in the MapReduce Framework 2011 SIGMOD 7.7570485e-05
3,068 Scalable Big Graph Processing in MapReduce 2014 SIGMOD 7.6841028e-05
3,296 Lightning Fast and Space Efficient Inequality Joins 2015 VLDB 7.444648e-05
3,339 Optimizing Analytic Data Flows for Multiple Execution Engines 2012 SIGMOD 7.4045097e-05
3,895 Hyper Dimension Shuffle: Efficient Data Repartition at Petabyte Scale in SCOPE 2019 VLDB 6.9360458e-05
4,250 The Unified Logging Infrastructure for Data Analytics at Twitter 2012 VLDB 6.7005978e-05
4,534 AdaptDB: Adaptive Partitioning for Distributed Joins 2017 VLDB 6.5535468e-05
5,134 Holistic Indexing in Main-memory Column-stores 2015 SIGMOD 6.2573427e-05
5,278 Only Aggressive Elephants are Fast Elephants 2012 VLDB 6.1976381e-05
5,653 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 6.048773e-05
6,076 Exploiting Soft and Hard Correlations in Big Data Query Optimization 2016 VLDB 5.8918218e-05
6,326 Clydesdale: Structured Data Processing on Hadoop 2012 SIGMOD 5.8120855e-05
6,424 A Hadoop Based Distributed Loading Approach to Parallel Data Warehouses 2011 SIGMOD 5.7873769e-05
6,922 Petabyte-Scale Row-Level Operations in Data Lakehouses 2024 VLDB 5.6415103e-05
7,388 Lachesis: Automatic Partitioning for UDF-Centric Analytics 2021 VLDB 5.5355684e-05
8,164 ScalaGiST: Scalable Generalized Search Trees for MapReduce Systems [Innovative Systems Paper] 2014 VLDB 5.3850352e-05
8,171 Pangea: Monolithic Distributed Storage for Data Analytics 2019 VLDB 5.3826392e-05
8,177 Indexing HDFS Data in PDW: Splitting the data from the index 2014 VLDB 5.3818353e-05
8,451 CARTILAGE: Adding Flexibility to the Hadoop Skeleton 2013 SIGMOD 5.3324907e-05
8,631 Piranha: Optimizing Short Jobs in Hadoop 2013 VLDB 5.298422e-05
9,688 Rank Join Queries in NoSQL Databases 2014 VLDB 5.1403642e-05
9,699 Efficient Big Data Processing in Hadoop MapReduce 2012 VLDB 5.1375205e-05
12,482 DGFIndex for Smart Grid: Enhancing Hive with a Cost-Effective Multidimensional Range Index 2014 VLDB 4.9769913e-05
12,522 D-Hive: Data Bees Pollinating RDF, Text, and Time 2013 CIDR 4.9769913e-05
12,524 How Achaeans Would Construct Columns in Troy 2013 CIDR 4.9769913e-05
12,565 Mosquito: Another One Bites the Data Upload STream 2013 VLDB 4.9769913e-05
14,003 RAFT at Work: Speeding-Up MapReduce Applications under Task and Node Failures 2011 SIGMOD -
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 10 of 10 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers