DBScholar

Back to papers

Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing)

Summary: Hadoop++ transparently accelerates Hadoop by injecting indexing and join optimizations through UDFs, without modifying its framework or interface. It substantially outperforms Hadoop and HadoopDB while remaining compatible with future Hadoop changes. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
h36bb8849cf7e7883
Venue
VLDB
Year
2010
Pagerank
0.0001488755
Overall Rank
673 | 95.48%
DOI
10.14778/1920841.1920908

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{dittrich_vldb10,
        title = {{Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing)}},
        author = {Dittrich, Jens and Kargin, Yagiz and Quiané-Ruiz, Jorge-Arnulfo and Setty, Vinay and Jindal, Alekh and Schad, Jörg},
        journal = {PVLDB},
        series = {{VLDB} '10},
        volume = {3},
        number = {1},
        pages = {518--529},
        doi = {10.14778/1920841.1920908},
        url = {https://doi.org/10.14778/1920841.1920908},
        year = {2010}
}

Incoming Citations (Sorted by Pagerank)

Showing 39 of 39 citing papers.

Rank Citing Paper Year Venue Pagerank
753 Profiling, What-if Analysis, and Cost-based Optimization of MapReduce Programs 2011 VLDB 0.00014237583
940 Starfish: A Self-tuning System for Big Data Analytics 2011 CIDR 0.00012964445
1,072 Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads 2012 VLDB 0.00012168947
1,281 Automatic Optimization for MapReduce Programs 2011 VLDB 0.00011213384
1,351 SkewTune: Mitigating Skew in MapReduce Applications 2012 SIGMOD 0.00010934347
1,708 PerfXplain: Debugging MapReduce Job Performance 2012 VLDB 9.8224587e-05
2,173 Efficient Processing of Data Warehousing Queries in a Split Execution Environment 2011 SIGMOD 8.91924e-05
2,340 TriAD: A Distributed Shared-Nothing RDF Engine based on Asynchronous Message Passing 2014 SIGMOD 8.6071228e-05
2,573 Minimal MapReduce Algorithms 2013 SIGMOD 8.2821647e-05
2,687 CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop 2011 VLDB 8.1301312e-05
2,903 Column-Oriented Storage Techniques for MapReduce 2011 VLDB 7.8810111e-05
2,931 Efficient Multi-way Theta-Join Processing Using MapReduce 2012 VLDB 7.8405483e-05
3,007 Llama: Leveraging Columnar Storage for Scalable Join Processing in the MapReduce Framework 2011 SIGMOD 7.7607173e-05
3,066 Scalable Big Graph Processing in MapReduce 2014 SIGMOD 7.6877117e-05
3,295 Lightning Fast and Space Efficient Inequality Joins 2015 VLDB 7.448168e-05
3,339 Optimizing Analytic Data Flows for Multiple Execution Engines 2012 SIGMOD 7.4080114e-05
3,895 Hyper Dimension Shuffle: Efficient Data Repartition at Petabyte Scale in SCOPE 2019 VLDB 6.9393157e-05
4,249 The Unified Logging Infrastructure for Data Analytics at Twitter 2012 VLDB 6.7037533e-05
4,533 AdaptDB: Adaptive Partitioning for Distributed Joins 2017 VLDB 6.5565658e-05
5,131 Holistic Indexing in Main-memory Column-stores 2015 SIGMOD 6.2602133e-05
5,274 Only Aggressive Elephants are Fast Elephants 2012 VLDB 6.2005706e-05
5,652 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 6.0515502e-05
6,074 Exploiting Soft and Hard Correlations in Big Data Query Optimization 2016 VLDB 5.8946121e-05
6,322 Clydesdale: Structured Data Processing on Hadoop 2012 SIGMOD 5.8148318e-05
6,422 A Hadoop Based Distributed Loading Approach to Parallel Data Warehouses 2011 SIGMOD 5.7901142e-05
6,920 Petabyte-Scale Row-Level Operations in Data Lakehouses 2024 VLDB 5.6441822e-05
7,386 Lachesis: Automatic Partitioning for UDF-Centric Analytics 2021 VLDB 5.538189e-05
8,158 ScalaGiST: Scalable Generalized Search Trees for MapReduce Systems [Innovative Systems Paper] 2014 VLDB 5.3875848e-05
8,165 Pangea: Monolithic Distributed Storage for Data Analytics 2019 VLDB 5.3851885e-05
8,171 Indexing HDFS Data in PDW: Splitting the data from the index 2014 VLDB 5.3843795e-05
8,442 CARTILAGE: Adding Flexibility to the Hadoop Skeleton 2013 SIGMOD 5.3350162e-05
8,624 Piranha: Optimizing Short Jobs in Hadoop 2013 VLDB 5.3009314e-05
9,681 Rank Join Queries in NoSQL Databases 2014 VLDB 5.1427987e-05
9,693 Efficient Big Data Processing in Hadoop MapReduce 2012 VLDB 5.1399537e-05
12,476 DGFIndex for Smart Grid: Enhancing Hive with a Cost-Effective Multidimensional Range Index 2014 VLDB 4.9793485e-05
12,516 D-Hive: Data Bees Pollinating RDF, Text, and Time 2013 CIDR 4.9793485e-05
12,518 How Achaeans Would Construct Columns in Troy 2013 CIDR 4.9793485e-05
12,559 Mosquito: Another One Bites the Data Upload STream 2013 VLDB 4.9793485e-05
13,998 RAFT at Work: Speeding-Up MapReduce Applications under Task and Node Failures 2011 SIGMOD -
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 10 of 10 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers