DBScholar

Back to papers

Magnet: Push-based Shuffle Service for Large-scale Data Processing

Summary: Magnet is a push-based Spark shuffle service that merges fragmented intermediate data into large blocks and co-locates them with reducers, scaling to petabytes/day and thousands of nodes. It cuts LinkedIn production job runtimes nearly 30% while eliminating shuffle tuning. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
12405
Venue
VLDB
Year
2020
Pagerank
6.2346971e-05
Overall Rank
5,390 | 63.03%
DOI
10.14778/3415478.3415558

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{shen_vldb20,
        title = {{Magnet: Push-based Shuffle Service for Large-scale Data Processing}},
        author = {Shen, Min and Zhou, Ye and Singh, Chandni},
        journal = {PVLDB},
        series = {{VLDB} '20},
        volume = {13},
        number = {12},
        pages = {3382--3395},
        doi = {10.14778/3415478.3415558},
        url = {https://doi.org/10.14778/3415478.3415558},
        year = {2020}
}

Incoming Citations (Sorted by Pagerank)

Showing 5 of 5 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 3 of 3 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
24 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00054865648
923 Starfish: A Self-tuning System for Big Data Analytics 2011 CIDR 0.00013189886
3,964 Hyper Dimension Shuffle: Efficient Data Repartition at Petabyte Scale in SCOPE 2019 VLDB 6.9855158e-05
Previous Page 1 / 1 Next

Semantically Similar Papers