Peta-Scale Data Warehousing at Yahoo!
Summary: EVEREST is a SQL-compliant, parallel data warehouse for Yahoo!, built on a columnar architecture with commodity hardware. Unlike commercial engines, it delivers scale, analytics, and lower admin costs, in production since 2007, managing six PB. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Mona Ahuja (Yahoo)
- 2. Cheng Che Chen (Yahoo)
- 3. Ravi Gottapu (Yahoo)
- 4. Jörg Hallmann (Yahoo)
- 5. Waqar Hasan (Yahoo)
- 6. Richard Johnson (Yahoo)
- 7. Maciek Kozyrczak (Yahoo)
- 8. Ramesh Pabbati (Yahoo)
- 9. Neeta Pandit (Yahoo)
- 10. Sreenivasulu Pokuri (Yahoo)
- 11. Krishna Uppala (Yahoo)
BibTeX Citation
@inproceedings{ahuja_sigmod09,
title = {{Peta-Scale Data Warehousing at Yahoo!}},
author = {Ahuja, Mona and Chen, Cheng Che and Gottapu, Ravi and Hallmann, Jörg and Hasan, Waqar and Johnson, Richard and Kozyrczak, Maciek and Pabbati, Ramesh and Pandit, Neeta and Pokuri, Sreenivasulu and Uppala, Krishna},
series = {{SIGMOD} '09},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/1559845.1559935},
url = {https://dl.acm.org/doi/10.1145/1559845.1559935},
year = {2009}
}
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,883 | ReStore: Reusing Results of MapReduce Jobs | 2012 | VLDB | 9.5421713e-05 |
| 4,463 | Consistency in a Stream Warehouse | 2011 | CIDR | 6.6860367e-05 |
| 9,125 | Data Stream Warehousing | 2013 | SIGMOD | 5.3191726e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 1 of 1 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,107 | Simultaneous Optimization and Evaluation of Multiple Dimensional Queries | 1998 | SIGMOD | 0.00012145695 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 3,555 | Apache Hive: From MapReduce to Enterprise-grade Big Data Warehousing | 2019 | SIGMOD |
| 2 | 32 | Hive - A Warehousing Solution Over a Map-Reduce Framework | 2009 | VLDB |
| 3 | 3,749 | Integrating Hadoop and Parallel DBMS | 2010 | SIGMOD |
| 4 | 8,957 | What is the data warehousing problem? (Are materialized views the answer?) | 1996 | VLDB |
| 5 | 9,711 | PNUTS to Sherpa: Lessons from Yahoo!’s Cloud Database | 2019 | VLDB |
| 6 | 237 | Amazon Redshift and the Case for Simpler Data Warehouses | 2015 | SIGMOD |
| 7 | 9,274 | DataGarage: Warehousing Massive Performance Data on Commodity Servers | 2010 | VLDB |
| 8 | 2,159 | Efficient Processing of Data Warehousing Queries in a Split Execution Environment | 2011 | SIGMOD |
| 9 | 2,191 | Data Warehousing and Analytics Infrastructure at Facebook | 2010 | SIGMOD |
| 10 | 1,791 | Cheetah: A High Performance, Custom Data Warehouse on Top of MapReduce | 2010 | VLDB |