Apache Tez: A Unifying Framework for Modeling and Building Data Processing Applications
Summary: Tez is an open framework to build data-flow engines on YARN, enabling component reuse with a flexible data plane. It unifies blocks to curb fragmentation and enables dynamic partition pruning; Tez-backed Hive, Pig, Spark, Cascading beat native YARN on TPC-DS/H. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Bikas Saha (Hortonworks)
- 2. Hitesh Shah (Hortonworks)
- 3. Siddharth Seth (Hortonworks)
- 4. Gopal Vijayaraghavan (Hortonworks)
- 5. Arun Murthy (Hortonworks)
- 6. Carlo Curino (Microsoft)
BibTeX Citation
@inproceedings{saha_sigmod15,
title = {{Apache Tez: A Unifying Framework for Modeling and Building Data Processing Applications}},
author = {Saha, Bikas and Shah, Hitesh and Seth, Siddharth and Vijayaraghavan, Gopal and Murthy, Arun and Curino, Carlo},
series = {{SIGMOD} '15},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2723372.2742790},
url = {https://dl.acm.org/doi/10.1145/2723372.2742790},
year = {2015}
}
Incoming Citations (Sorted by Pagerank)
Showing 9 of 9 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 379 | Apache Calcite: A Foundational Framework for Optimized Query Processing Over Heterogeneous Data Sources | 2018 | SIGMOD | 0.00019514689 |
| 1,829 | DBEst: Revisiting Approximate Query Processing Engines with Machine Learning Models | 2019 | SIGMOD | 9.5510333e-05 |
| 3,446 | Apache Hive: From MapReduce to Enterprise-grade Big Data Warehousing | 2019 | SIGMOD | 7.2942885e-05 |
| 5,005 | Using Cloud Functions as Accelerator for Elastic Data Analytics | 2023 | SIGMOD | 6.3177105e-05 |
| 6,523 | REEF: Retainable Evaluator Execution Framework | 2015 | SIGMOD | 5.7568314e-05 |
| 11,456 | IcedTea: Efficient and Responsive Time-Travel Debugging in Dataflow Systems | 2025 | VLDB | 4.9793485e-05 |
| 12,031 | Fangorn: Adaptive Execution Framework for Heterogeneous Workloads on Shared Clusters | 2021 | VLDB | 4.9793485e-05 |
| 12,185 | Integration of Large-Scale Data Processing Systems and Traditional Parallel Database Technology | 2019 | VLDB | 4.9793485e-05 |
| 12,438 | Tutorial: SQL-on-Hadoop Systems | 2015 | VLDB | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 6 | Pig Latin: A Not-So-Foreign Language for Data Processing | 2008 | SIGMOD | 0.001052036 |
| 31 | Hive - A Warehousing Solution Over a Map-Reduce Framework | 2009 | VLDB | 0.00049839909 |
| 49 | Dremel: Interactive Analysis of Web-Scale Datasets | 2010 | VLDB | 0.00043160717 |
| 327 | Impala: A Modern, Open-Source SQL Engine for Hadoop | 2015 | CIDR | 0.0002095191 |
| 3,556 | WANalytics: Analytics for a Geo-Distributed Data-Intensive World | 2015 | CIDR | 7.2084485e-05 |
| 7,176 | REEF: Retainable Evaluator Execution Framework | 2013 | VLDB | 5.5934672e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 2,282 | A Platform for Scalable One-Pass Analytics using MapReduce | 2011 | SIGMOD |
| 2 | 2,689 | Major Technical Advancements in Apache Hive | 2014 | SIGMOD |
| 3 | 6,523 | REEF: Retainable Evaluator Execution Framework | 2015 | SIGMOD |
| 4 | 651 | Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience | 2009 | VLDB |
| 5 | 2,053 | tf.data: A Machine Learning Data Processing Framework | 2021 | VLDB |
| 6 | 876 | Tenzing: A SQL Implementation On The MapReduce Framework | 2011 | VLDB |
| 7 | 673 | Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) | 2010 | VLDB |
| 8 | 1,884 | SQL-on-Hadoop: Full Circle Back to Shared-Nothing Database Architectures | 2014 | VLDB |
| 9 | 379 | Apache Calcite: A Foundational Framework for Optimized Query Processing Over Heterogeneous Data Sources | 2018 | SIGMOD |
| 10 | 3,446 | Apache Hive: From MapReduce to Enterprise-grade Big Data Warehousing | 2019 | SIGMOD |