Apache Tez: A Unifying Framework for Modeling and Building Data Processing Applications
Summary: Tez is an open framework to build data-flow engines on YARN, enabling component reuse with a flexible data plane. It unifies blocks to curb fragmentation and enables dynamic partition pruning; Tez-backed Hive, Pig, Spark, Cascading beat native YARN on TPC-DS/H. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Bikas Saha (Hortonworks)
- 2. Hitesh Shah (Hortonworks)
- 3. Siddharth Seth (Hortonworks)
- 4. Gopal Vijayaraghavan (Hortonworks)
- 5. Arun Murthy (Hortonworks)
- 6. Carlo Curino (Microsoft)
BibTeX Citation
@inproceedings{saha_sigmod15,
title = {{Apache Tez: A Unifying Framework for Modeling and Building Data Processing Applications}},
author = {Saha, Bikas and Shah, Hitesh and Seth, Siddharth and Vijayaraghavan, Gopal and Murthy, Arun and Curino, Carlo},
series = {{SIGMOD} '15},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2723372.2742790},
url = {https://dl.acm.org/doi/10.1145/2723372.2742790},
year = {2015}
}
Incoming Citations (Sorted by Pagerank)
Showing 9 of 9 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 445 | Apache Calcite: A Foundational Framework for Optimized Query Processing Over Heterogeneous Data Sources | 2018 | SIGMOD | 0.00018336751 |
| 1,799 | DBEst: Revisiting Approximate Query Processing Engines with Machine Learning Models | 2019 | SIGMOD | 9.7326398e-05 |
| 3,555 | Apache Hive: From MapReduce to Enterprise-grade Big Data Warehousing | 2019 | SIGMOD | 7.3115321e-05 |
| 4,908 | Using Cloud Functions as Accelerator for Elastic Data Analytics | 2023 | SIGMOD | 6.4488784e-05 |
| 6,430 | REEF: Retainable Evaluator Execution Framework | 2015 | SIGMOD | 5.880521e-05 |
| 11,106 | IcedTea: Efficient and Responsive Time-Travel Debugging in Dataflow Systems | 2025 | VLDB | 5.093636e-05 |
| 11,728 | Fangorn: Adaptive Execution Framework for Heterogeneous Workloads on Shared Clusters | 2021 | VLDB | 5.093636e-05 |
| 11,885 | Integration of Large-Scale Data Processing Systems and Traditional Parallel Database Technology | 2019 | VLDB | 5.093636e-05 |
| 12,146 | Tutorial: SQL-on-Hadoop Systems | 2015 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 6 | Pig Latin: A Not-So-Foreign Language for Data Processing | 2008 | SIGMOD | 0.0010686205 |
| 32 | Hive - A Warehousing Solution Over a Map-Reduce Framework | 2009 | VLDB | 0.00050111008 |
| 51 | Dremel: Interactive Analysis of Web-Scale Datasets | 2010 | VLDB | 0.0004291425 |
| 330 | Impala: A Modern, Open-Source SQL Engine for Hadoop | 2015 | CIDR | 0.0002104801 |
| 3,494 | WANalytics: Analytics for a Geo-Distributed Data-Intensive World | 2015 | CIDR | 7.3642231e-05 |
| 7,082 | REEF: Retainable Evaluator Execution Framework | 2013 | VLDB | 5.708011e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 2,265 | A Platform for Scalable One-Pass Analytics using MapReduce | 2011 | SIGMOD |
| 2 | 2,706 | Major Technical Advancements in Apache Hive | 2014 | SIGMOD |
| 3 | 6,430 | REEF: Retainable Evaluator Execution Framework | 2015 | SIGMOD |
| 4 | 642 | Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience | 2009 | VLDB |
| 5 | 2,018 | tf.data: A Machine Learning Data Processing Framework | 2021 | VLDB |
| 6 | 872 | Tenzing: A SQL Implementation On The MapReduce Framework | 2011 | VLDB |
| 7 | 660 | Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) | 2010 | VLDB |
| 8 | 1,889 | SQL-on-Hadoop: Full Circle Back to Shared-Nothing Database Architectures | 2014 | VLDB |
| 9 | 445 | Apache Calcite: A Foundational Framework for Optimized Query Processing Over Heterogeneous Data Sources | 2018 | SIGMOD |
| 10 | 3,555 | Apache Hive: From MapReduce to Enterprise-grade Big Data Warehousing | 2019 | SIGMOD |