DBScholar

Back to papers

Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing)

Summary: Hadoop++ transparently accelerates Hadoop by injecting indexing and join optimizations through UDFs, without modifying its framework or interface. It substantially outperforms Hadoop and HadoopDB while remaining compatible with future Hadoop changes. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
10292
Venue
VLDB
Year
2010
Pagerank
0.00015198804
Overall Rank
660 | 95.48%
DOI
10.14778/1920841.1920908

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{dittrich_vldb10,
        title = {{Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing)}},
        author = {Dittrich, Jens and Kargin, Yagiz and Quiané-Ruiz, Jorge-Arnulfo and Setty, Vinay and Jindal, Alekh and Schad, Jörg},
        journal = {PVLDB},
        series = {{VLDB} '10},
        volume = {3},
        number = {1},
        pages = {518--529},
        doi = {10.14778/1920841.1920908},
        url = {https://doi.org/10.14778/1920841.1920908},
        year = {2010}
}

Incoming Citations (Sorted by Pagerank)

Showing 39 of 39 citing papers.

Rank Citing Paper Year Venue Pagerank
735 Profiling, What-if Analysis, and Cost-based Optimization of MapReduce Programs 2011 VLDB 0.00014522606
923 Starfish: A Self-tuning System for Big Data Analytics 2011 CIDR 0.00013189886
1,054 Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads 2012 VLDB 0.00012390673
1,257 Automatic Optimization for MapReduce Programs 2011 VLDB 0.000114432
1,319 SkewTune: Mitigating Skew in MapReduce Applications 2012 SIGMOD 0.00011175005
1,694 PerfXplain: Debugging MapReduce Job Performance 2012 VLDB 9.9953156e-05
2,159 Efficient Processing of Data Warehousing Queries in a Split Execution Environment 2011 SIGMOD 9.061086e-05
2,287 TriAD: A Distributed Shared-Nothing RDF Engine based on Asynchronous Message Passing 2014 SIGMOD 8.8034872e-05
2,539 Minimal MapReduce Algorithms 2013 SIGMOD 8.4526595e-05
2,642 CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop 2011 VLDB 8.3059948e-05
2,849 Column-Oriented Storage Techniques for MapReduce 2011 VLDB 8.053191e-05
2,887 Efficient Multi-way Theta-Join Processing Using MapReduce 2012 VLDB 7.9952432e-05
2,942 Llama: Leveraging Columnar Storage for Scalable Join Processing in the MapReduce Framework 2011 SIGMOD 7.9358593e-05
3,008 Scalable Big Graph Processing in MapReduce 2014 SIGMOD 7.8578871e-05
3,275 Optimizing Analytic Data Flows for Multiple Execution Engines 2012 SIGMOD 7.5747814e-05
3,295 Lightning Fast and Space Efficient Inequality Joins 2015 VLDB 7.5477715e-05
3,964 Hyper Dimension Shuffle: Efficient Data Repartition at Petabyte Scale in SCOPE 2019 VLDB 6.9855158e-05
4,169 The Unified Logging Infrastructure for Data Analytics at Twitter 2012 VLDB 6.8561463e-05
4,456 AdaptDB: Adaptive Partitioning for Distributed Joins 2017 VLDB 6.692321e-05
5,039 Holistic Indexing in Main-memory Column-stores 2015 SIGMOD 6.3909067e-05
5,156 Only Aggressive Elephants are Fast Elephants 2012 VLDB 6.3410921e-05
5,522 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 6.1871697e-05
6,044 Exploiting Soft and Hard Correlations in Big Data Query Optimization 2016 VLDB 5.995949e-05
6,199 Clydesdale: Structured Data Processing on Hadoop 2012 SIGMOD 5.9458654e-05
6,312 A Hadoop Based Distributed Loading Approach to Parallel Data Warehouses 2011 SIGMOD 5.9170938e-05
7,311 Lachesis: Automatic Partitioning for UDF-Centric Analytics 2021 VLDB 5.6491618e-05
7,780 Petabyte-Scale Row-Level Operations in Data Lakehouses 2024 VLDB 5.5457298e-05
7,990 ScalaGiST: Scalable Generalized Search Trees for MapReduce Systems [Innovative Systems Paper] 2014 VLDB 5.5107935e-05
8,001 Pangea: Monolithic Distributed Storage for Data Analytics 2019 VLDB 5.508791e-05
8,014 Indexing HDFS Data in PDW: Splitting the data from the index 2014 VLDB 5.5073329e-05
8,276 CARTILAGE: Adding Flexibility to the Hadoop Skeleton 2013 SIGMOD 5.4574671e-05
8,455 Piranha: Optimizing Short Jobs in Hadoop 2013 VLDB 5.4226e-05
9,497 Rank Join Queries in NoSQL Databases 2014 VLDB 5.2608378e-05
9,510 Efficient Big Data Processing in Hadoop MapReduce 2012 VLDB 5.2576928e-05
12,185 DGFIndex for Smart Grid: Enhancing Hive with a Cost-Effective Multidimensional Range Index 2014 VLDB 5.093636e-05
12,225 D-Hive: Data Bees Pollinating RDF, Text, and Time 2013 CIDR 5.093636e-05
12,227 How Achaeans Would Construct Columns in Troy 2013 CIDR 5.093636e-05
12,268 Mosquito: Another One Bites the Data Upload STream 2013 VLDB 5.093636e-05
13,686 RAFT at Work: Speeding-Up MapReduce Applications under Task and Node Failures 2011 SIGMOD -
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 10 of 10 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers