Back to papers
Distributed Numerical and Machine Learning Computations via Two-Phase Execution of Aggregated Join Trees
Summary: Two-phase execution for numerical/ML workloads expressed as aggregated join trees (joins then aggregation). Pilot run collects lineage to enable record-level planning before execution; experiments show this relational two-phase approach as an effective platform for distributed ML.
(summarized by gpt-5-nano on Feb 09 2026)
- Paper ID
- 12314
- Venue
- VLDB
- Year
- 2021
- Pagerank
- 4.2951758e-05
- Overall Rank
- 9,705 | 32.55%
- DOI
-
10.14778/3450980.3450991
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 17 of 17 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 1 |
Access Path Selection in a Relational Database Management System |
1979 |
SIGMOD |
0.0040465394 |
| 8 |
Optimal Aggregation Algorithms for Middleware [Extended Abstract] |
2001 |
PODS |
0.0015436578 |
| 216 |
Ripple Joins for Online Aggregation |
1999 |
SIGMOD |
0.00033560137 |
| 300 |
Trio: A System for Data, Uncertainty, and Lineage |
2006 |
VLDB |
0.00028503742 |
| 315 |
Distributed Query Processing In A Relational Data Base System |
1978 |
SIGMOD |
0.00028022592 |
| 454 |
An Overview of Query Optimization in Relational Systems |
1998 |
PODS |
0.00022796106 |
| 806 |
Exploiting Inter-Operation Parallelism in XPRS |
1992 |
SIGMOD |
0.00016428214 |
| 2,286 |
SMOKE: Fine-grained Lineage at Interactive Speed |
2018 |
VLDB |
9.102574e-05 |
| 2,518 |
Track Join: Distributed Joins with Minimal Network Traffic |
2014 |
SIGMOD |
8.6052941e-05 |
| 3,122 |
DB4ML – An In-Memory Database Kernel with Machine Learning Support |
2020 |
SIGMOD |
7.5284233e-05 |
| 3,961 |
MLog: Towards Declarative In-Database Machine Learning |
2017 |
VLDB |
6.5824022e-05 |
| 4,052 |
Optimizing Nested Queries with Parameter Sort Orders |
2005 |
VLDB |
6.4923e-05 |
| 5,012 |
Dynamically Optimizing Queries over Large Scale Data Platforms |
2014 |
SIGMOD |
5.7543101e-05 |
| 5,836 |
Tensor Relational Algebra for Distributed Machine Learning System Design |
2021 |
VLDB |
5.3079723e-05 |
| 6,747 |
DistME: A Fast and Elastic Distributed Matrix Computation Engine using GPUs |
2019 |
SIGMOD |
4.9369478e-05 |
| 9,002 |
Chasing Similarity: Distribution-aware Aggregation Scheduling |
2019 |
VLDB |
4.4077753e-05 |
| 9,337 |
PlinyCompute: A Platform for High-Performance, Distributed, Data-Intensive Tool Development |
2018 |
SIGMOD |
4.351469e-05 |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 2,599 |
Design and Evaluation of Parallel Pipelined Join Algorithms |
1987 |
SIGMOD |
8.4655075e-05 |
| 4,320 |
Parallel Processing of Recursive Queries in Distributed Architectures |
1989 |
VLDB |
6.2822407e-05 |
| 4,406 |
Declarative Recursive Computation on an RDBMS |
2019 |
VLDB |
6.2044305e-05 |
| 8,777 |
Accelerate Distributed Joins with Predicate Transfer |
2025 |
SIGMOD |
4.4492064e-05 |
| 6,194 |
Automatic Optimization of Matrix Implementations for Distributed Machine Learning and Linear Algebra |
2021 |
SIGMOD |
5.1592984e-05 |
| 5,836 |
Tensor Relational Algebra for Distributed Machine Learning System Design |
2021 |
VLDB |
5.3079723e-05 |
| 1,938 |
From Theory to Practice: Efficient Join Query Evaluation in a Parallel Database System |
2015 |
SIGMOD |
0.00010025547 |
| 11,898 |
Let's Rethink Join Optimization in Distributed Systems |
2015 |
CIDR |
4.1905499e-05 |
| 9,582 |
Sharing Aggregate Computation for Distributed Queries |
2007 |
SIGMOD |
4.3185789e-05 |
| 4,128 |
Advanced Join Strategies for Large-Scale Distributed Computation |
2014 |
VLDB |
6.4214449e-05 |