Lifetime-Based Memory Management for Distributed Data Processing Systems
Summary: Lifetime-aware allocation derives object lifetimes from UDFs and types, grouping similarly lived data into byte arrays for bulk reclamation. Deca integrates this transparently into Spark, sharply reducing GC, memory use, spilling, and execution time. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Lu Lu (Huazhong University of Science and Technology)
- 2. Xuanhua Shi (Huazhong University of Science and Technology)
- 3. Yongluan Zhou (University of Southern Denmark)
- 4. Xiong Zhang (Huazhong University of Science and Technology)
- 5. Hai Jin (Huazhong University of Science and Technology)
- 6. Cheng Pei (Huazhong University of Science and Technology)
- 7. Ligang He (University of Warwick)
- 8. Yuanzhen Geng (Huazhong University of Science and Technology)
BibTeX Citation
@article{lu_vldb16,
title = {{Lifetime-Based Memory Management for Distributed Data Processing Systems}},
author = {Lu, Lu and Shi, Xuanhua and Zhou, Yongluan and Zhang, Xiong and Jin, Hai and Pei, Cheng and He, Ligang and Geng, Yuanzhen},
journal = {PVLDB},
series = {{VLDB} '16},
volume = {9},
number = {12},
pages = {936--947},
doi = {10.14778/2994509.2994522},
url = {https://doi.org/10.14778/2994509.2994522},
year = {2016}
}
Incoming Citations (Sorted by Pagerank)
Showing 5 of 5 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 7,821 | Concurrent Log-Structured Memory for Many-Core Key-Value Stores | 2018 | VLDB | 5.5373067e-05 |
| 8,001 | Pangea: Monolithic Distributed Storage for Data Analytics | 2019 | VLDB | 5.508791e-05 |
| 9,572 | PlinyCompute: A Platform for High-Performance, Distributed, Data-Intensive Tool Development | 2018 | SIGMOD | 5.2528121e-05 |
| 10,076 | Chukonu: A Fully-Featured High-Performance Big Data Framework that Integrates a Native Compute Engine into Spark | 2022 | VLDB | 5.1613298e-05 |
| 11,889 | An Experimental Evaluation of Garbage Collectors on Big Data Applications | 2019 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 4 of 4 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 24 | Spark SQL: Relational Data Processing in Spark | 2015 | SIGMOD | 0.00054865648 |
| 2,265 | A Platform for Scalable One-Pass Analytics using MapReduce | 2011 | SIGMOD | 8.8398946e-05 |
| 3,719 | M3R: Increased Performance for In-Memory Hadoop Jobs | 2012 | VLDB | 7.1740814e-05 |
| 4,208 | Clash of the Titans: MapReduce vs. Spark for Large Scale Data Analytics | 2015 | VLDB | 6.8319812e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 4,480 | LocationSpark: A Distributed In-Memory Data Management System for Big Spatial Data | 2016 | VLDB |
| 2 | 8,457 | Towards Resource Efficiency: Practical Insights into Large-Scale Spark Workloads at ByteDance | 2024 | VLDB |
| 3 | 1,580 | Compaction management in distributed key-value datastores | 2015 | VLDB |
| 4 | 4,049 | Resource Elasticity for Large-Scale Machine Learning | 2015 | SIGMOD |
| 5 | 2,610 | Maintaining Time-Decaying Stream Aggregates | 2003 | PODS |
| 6 | 12,001 | Runtime Optimization of Join Location in Parallel Data Management Systems | 2017 | VLDB |
| 7 | 9,655 | [Demo] Low-latency Spark Queries on Updatable Data | 2019 | SIGMOD |
| 8 | 2,681 | Exploiting Matrix Dependency for Efficient Distributed Matrix Computation | 2015 | SIGMOD |
| 9 | 7,105 | Efficient In-memory Data Management: An Analysis | 2014 | VLDB |
| 10 | 11,889 | An Experimental Evaluation of Garbage Collectors on Big Data Applications | 2019 | VLDB |