Pixida: Optimizing Data Parallel Jobs in Wide-Area Data Analytics
Summary: PIXIDA schedules wide-area data-parallel jobs to minimize traffic over constrained, volatile cross-DC links. Its SILO abstraction enables a job-specific graph-partitioning formulation and algorithm, integrated with Spark, reducing cross-DC traffic up to 9×. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Konstantinos Kloudas (University of Lisbon)
- 2. Margarida Mamede (NOVA University Lisbon)
- 3. Nuno Preguica (NOVA University Lisbon)
- 4. Rodrigo Rodrigues (University of Lisbon)
BibTeX Citation
@article{kloudas_vldb16,
title = {{Pixida: Optimizing Data Parallel Jobs in Wide-Area Data Analytics}},
author = {Kloudas, Konstantinos and Mamede, Margarida and Preguica, Nuno and Rodrigues, Rodrigo},
journal = {PVLDB},
series = {{VLDB} '16},
volume = {9},
number = {2},
pages = {72--83},
doi = {10.14778/2850578.2850581},
url = {https://doi.org/10.14778/2850578.2850581},
year = {2016}
}
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 5,568 | Yugong: Geo-Distributed Data and Job Placement at Scale | 2019 | VLDB | 6.0813683e-05 |
| 11,845 | AutoMon: Automatic Distributed Monitoring for Arbitrary Multivariate Functions | 2022 | SIGMOD | 4.9793485e-05 |
| 12,345 | The Challenges of Global-scale Data Management | 2016 | SIGMOD | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 218 | MillWheel: Fault-Tolerant Stream Processing at Internet Scale | 2013 | VLDB | 0.00024390324 |
| 1,175 | Photon: Fault-tolerant and Scalable Joining of Continuous Data Streams | 2013 | SIGMOD | 0.00011649474 |
| 1,622 | Mesa: Geo-Replicated, Near Real-Time, Scalable Data Warehousing | 2014 | VLDB | 0.00010051462 |
| 3,556 | WANalytics: Analytics for a Geo-Distributed Data-Intensive World | 2015 | CIDR | 7.2084485e-05 |
| 4,470 | Recurring Job Optimization in Scope | 2012 | SIGMOD | 6.5853561e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 12,220 | DITA: A Distributed In-Memory Trajectory Analytics System | 2018 | SIGMOD |
| 2 | 3,895 | Hyper Dimension Shuffle: Efficient Data Repartition at Petabyte Scale in SCOPE | 2019 | VLDB |
| 3 | 1,249 | Distributed Evaluation of Subgraph Queries Using Worst-case Optimal Low-Memory Dataflows | 2018 | VLDB |
| 4 | 5,058 | Fine-Grained Modeling and Optimization for Intelligent Resource Management in Big Data Processing | 2022 | VLDB |
| 5 | 8,601 | Towards Resource Efficiency: Practical Insights into Large-Scale Spark Workloads at ByteDance | 2024 | VLDB |
| 6 | 7,356 | Parallelism-Optimizing Data Placement for Faster Data-Parallel Computations | 2023 | VLDB |
| 7 | 11,714 | QaaD (Query-as-a-Data): Scalable Execution of Massive Number of Small Queries in Spark | 2023 | SIGMOD |
| 8 | 12,298 | Runtime Optimization of Join Location in Parallel Data Management Systems | 2017 | VLDB |
| 9 | 1,745 | Selecting Subexpressions to Materialize at Datacenter Scale | 2018 | VLDB |
| 10 | 2,196 | DITA: Distributed In-Memory Trajectory Analytics | 2018 | SIGMOD |