DBScholar

Back to papers

RHEEM: Enabling Cross-Platform Data Processing - May The Big Data Be With You! -

Summary: Rheem enables cross-platform data processing by decoupling applications from execution platforms and partitioning tasks. A cost-based optimizer selects platforms and maps subtasks, with an executor orchestrating multi-platform workflows for lower cost. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h57b4174d6ca0e37c
Venue
VLDB
Year
2018
Pagerank
7.3304477e-05
Overall Rank
3,402 | 77.13%
DOI
10.14778/3236187.3236195

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{agrawal_vldb18,
        title = {{RHEEM: Enabling Cross-Platform Data Processing - May The Big Data Be With You! -}},
        author = {Agrawal, Divy and Chawla, Sanjay and Contreras-Rojas, Bertty and Elmagarmid, Ahmed and Idris, Yasser and Kaoudi, Zoi and Kruse, Sebastian and Lucas, Ji and Mansour, Essam and Ouzzani, Mourad and Papotti, Paolo and Quiané-Ruiz, Jorge-Arnulfo and Tang, Nan and Thirumuruganathan, Saravanan and Troudi, Anis},
        journal = {PVLDB},
        series = {{VLDB} '18},
        volume = {11},
        number = {11},
        pages = {1414--1427},
        doi = {10.14778/3236187.3236195},
        url = {https://doi.org/10.14778/3236187.3236195},
        year = {2018}
}

Incoming Citations (Sorted by Pagerank)

Showing 21 of 21 citing papers.

Rank Citing Paper Year Venue Pagerank
3,743 Logical and Physical Optimizations for SQL Query Execution over Large Language Models 2025 SIGMOD 7.0586112e-05
5,091 Towards Scalable Hybrid Stores: Constraint-Based Rewriting to the Rescue 2019 SIGMOD 6.2797352e-05
5,230 Babelfish: Efficient Execution of Polyglot Queries 2022 VLDB 6.2189001e-05
6,141 Expand your Training Limits! Generating Training Data for ML-based Data Management 2021 SIGMOD 5.8733296e-05
6,343 ESTOCADA: Towards Scalable Polystore Systems 2020 VLDB 5.8092399e-05
6,835 Skeena: Efficient and Consistent Cross-Engine Transactions 2022 SIGMOD 5.6684174e-05
7,208 Dataset Relationship Management 2019 CIDR 5.584745e-05
7,864 Blueprinting the Cloud: Unifying and Automatically Optimizing Cloud Data Infrastructures with BRAD 2024 VLDB 5.4367919e-05
9,034 The Power of Nested Parallelism in Big Data Processing – Hitting Three Flies with One Slap – 2021 SIGMOD 5.2334993e-05
9,088 Hyperspace: The Indexing Subsystem of Azure Synapse 2021 VLDB 5.2283159e-05
9,115 Check Out the Big Brain on BRAD: Simplifying Cloud Data Processing with Learned Automated Data Meshes 2023 VLDB 5.2271833e-05
9,316 On-Demand State Separation for Cloud Data Warehousing 2022 VLDB 5.1958405e-05
9,533 Farm Your ML-based Query Optimizer's Food! - Human-Guided Training Data Generation - 2022 CIDR 5.1636858e-05
9,800 Apache Wayang in Action: Enabling Data Systems Integration via a Unified Data Analytics Framework 2025 SIGMOD 5.1257999e-05
9,921 Polyglot Data Management: State of the Art & Open Challenges 2022 VLDB 5.1103839e-05
9,922 Unified Data Analytics: State-of-the-art and Open Problems 2022 VLDB 5.1103839e-05
10,305 Fast and Scalable Data Transfer Across Data Systems 2025 SIGMOD 5.0400722e-05
10,733 APEROL: Adaptive Parallel Edge-to-cloud Runtime Optimization for Layered Workflow Execution 2026 VLDB 4.9793485e-05
11,258 Accio: Bolt-on Query Federation 2025 VLDB 4.9793485e-05
11,714 QaaD (Query-as-a-Data): Scalable Execution of Massive Number of Small Queries in Spark 2023 SIGMOD 4.9793485e-05
12,005 In the Land of Data Streams where Synopses are Missing, One Framework to Bring Them All 2021 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 22 of 22 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
6 Pig Latin: A Not-So-Foreign Language for Data Processing 2008 SIGMOD 0.001052036
15 How Good Are Query Optimizers, Really? 2016 VLDB 0.00061066921
43 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.00045546775
186 Federated Database Systems for Managing Distributed, Heterogeneous, and Autonomous Databases 1991 VLDB 0.00025951043
415 SystemML: Declarative Machine Learning on Spark 2016 VLDB 0.0001865959
481 Robust Query Processing through Progressive Optimization 2004 SIGMOD 0.00017603972
697 NADEEF: A Commodity Data Cleaning System 2013 SIGMOD 0.00014694048
972 The Data Civilizer System 2017 CIDR 0.00012763234
1,193 Weld: A Common Runtime for High Performance Data Analytics 2017 CIDR 0.0001158809
1,751 Split Query Processing in Polybase 2013 SIGMOD 9.7291888e-05
2,195 Opening the Black Boxes in Data Flow Optimization 2012 VLDB 8.8781177e-05
2,423 BigDansing: A System for Big Data Cleansing 2015 SIGMOD 8.4877894e-05
2,492 A Demonstration of the BigDAWG Polystore System 2015 VLDB 8.3899912e-05
3,258 How to Fit when No One Size Fits 2013 CIDR 7.4873069e-05
3,295 Lightning Fast and Space Efficient Inequality Joins 2015 VLDB 7.448168e-05
3,339 Optimizing Analytic Data Flows for Multiple Execution Engines 2012 SIGMOD 7.4080114e-05
3,685 The Myria Big Data Management and Analytics System and Cloud Service 2017 CIDR 7.0988339e-05
3,840 MISO: Souping Up Big Data Query Processing with a Multistore System 2014 SIGMOD 6.9920793e-05
3,862 Husky: Towards a More Efficient and Expressive Distributed Computing Framework 2016 VLDB 6.9652009e-05
5,700 A Demo of the Data Civilizer System 2017 SIGMOD 6.0317753e-05
7,142 A Cost-based Optimizer for Gradient Descent Optimization 2017 SIGMOD 5.600563e-05
10,161 Rheem: Enabling Multi-Platform Task Execution 2016 SIGMOD 5.0688288e-05
Previous Page 1 / 1 Next

Semantically Similar Papers