Pixida: Optimizing Data Parallel Jobs in Wide-Area Data Analytics
Summary: PIXIDA schedules wide-area data-parallel jobs to minimize traffic over constrained, volatile cross-DC links. Its SILO abstraction enables a job-specific graph-partitioning formulation and algorithm, integrated with Spark, reducing cross-DC traffic up to 9×. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Konstantinos Kloudas (University of Lisbon)
- 2. Margarida Mamede (NOVA University Lisbon)
- 3. Nuno Preguica (NOVA University Lisbon)
- 4. Rodrigo Rodrigues (University of Lisbon)
BibTeX Citation
@article{kloudas_vldb16,
title = {{Pixida: Optimizing Data Parallel Jobs in Wide-Area Data Analytics}},
author = {Kloudas, Konstantinos and Mamede, Margarida and Preguica, Nuno and Rodrigues, Rodrigo},
journal = {PVLDB},
series = {{VLDB} '16},
volume = {9},
number = {2},
pages = {72--83},
doi = {10.14778/2850578.2850581},
url = {https://doi.org/10.14778/2850578.2850581},
year = {2016}
}
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 5,582 | Yugong: Geo-Distributed Data and Job Placement at Scale | 2019 | VLDB | 6.1619082e-05 |
| 11,536 | AutoMon: Automatic Distributed Monitoring for Arbitrary Multivariate Functions | 2022 | SIGMOD | 5.093636e-05 |
| 12,050 | The Challenges of Global-scale Data Management | 2016 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 224 | MillWheel: Fault-Tolerant Stream Processing at Internet Scale | 2013 | VLDB | 0.00024130894 |
| 1,171 | Photon: Fault-tolerant and Scalable Joining of Continuous Data Streams | 2013 | SIGMOD | 0.00011826434 |
| 1,606 | Mesa: Geo-Replicated, Near Real-Time, Scalable Data Warehousing | 2014 | VLDB | 0.00010228576 |
| 3,494 | WANalytics: Analytics for a Geo-Distributed Data-Intensive World | 2015 | CIDR | 7.3642231e-05 |
| 4,400 | Recurring Job Optimization in Scope | 2012 | SIGMOD | 6.7239302e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 11,921 | DITA: A Distributed In-Memory Trajectory Analytics System | 2018 | SIGMOD |
| 2 | 3,964 | Hyper Dimension Shuffle: Efficient Data Repartition at Petabyte Scale in SCOPE | 2019 | VLDB |
| 3 | 1,246 | Distributed Evaluation of Subgraph Queries Using Worst-case Optimal Low-Memory Dataflows | 2018 | VLDB |
| 4 | 5,388 | Fine-Grained Modeling and Optimization for Intelligent Resource Management in Big Data Processing | 2022 | VLDB |
| 5 | 8,457 | Towards Resource Efficiency: Practical Insights into Large-Scale Spark Workloads at ByteDance | 2024 | VLDB |
| 6 | 8,091 | Parallelism-Optimizing Data Placement for Faster Data-Parallel Computations | 2023 | VLDB |
| 7 | 11,399 | QaaD (Query-as-a-Data): Scalable Execution of Massive Number of Small Queries in Spark | 2023 | SIGMOD |
| 8 | 12,001 | Runtime Optimization of Join Location in Parallel Data Management Systems | 2017 | VLDB |
| 9 | 1,765 | Selecting Subexpressions to Materialize at Datacenter Scale | 2018 | VLDB |
| 10 | 2,416 | DITA: Distributed In-Memory Trajectory Analytics | 2018 | SIGMOD |