ParaX: Boosting Deep Learning for Big Data Analytics on Many-Core CPUs
Summary: ParaX removes per-layer barriers in CPU deep learning by mapping each instance to a core and overlapping memory- and compute-intensive layers. A NUMA-aware gradient server reduces synchronization overhead, delivering 1.73–2.93× throughput gains. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Lujia Yin (National University of Defense Technology)
- 2. Yiming Zhang (National University of Defense Technology)
- 3. Zhaoning Zhang (National University of Defense Technology)
- 4. Yuxing Peng (National University of Defense Technology)
- 5. Peng Zhao (Intel)
BibTeX Citation
@article{yin_vldb21,
title = {{ParaX: Boosting Deep Learning for Big Data Analytics on Many-Core CPUs}},
author = {Yin, Lujia and Zhang, Yiming and Zhang, Zhaoning and Peng, Yuxing and Zhao, Peng},
journal = {PVLDB},
series = {{VLDB} '21},
volume = {14},
number = {6},
pages = {864--877},
doi = {10.14778/3447689.3447692},
url = {https://doi.org/10.14778/3447689.3447692},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 4,273 | User-Defined Operators: Efficiently Integrating Custom Algorithms into Modern Databases | 2022 | VLDB | 6.7909197e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 3 of 3 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,150 | DimmWitted: A Study of Main-Memory Statistical Analytics | 2014 | VLDB | 0.00011943462 |
| 1,250 | Data Management in Machine Learning: Challenges, Techniques, and Systems | 2017 | SIGMOD | 0.00011485301 |
| 4,859 | Machine Learning for Big Data | 2013 | SIGMOD | 6.4750538e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 9,371 | Towards an Optimized GROUP BY Abstraction for Large-Scale Machine Learning | 2021 | VLDB |
| 2 | 3,782 | CROSSBOW: Scaling Deep Learning with Small Batch Sizes on Multi-GPU Servers | 2019 | VLDB |
| 3 | 4,067 | Distributed Deep Learning on Data Systems: A Comparative Analysis of Approaches | 2021 | VLDB |
| 4 | 3,303 | A Distributed Multi-GPU System for Fast Graph Processing | 2018 | VLDB |
| 5 | 8,905 | TensorSocket: Shared Data Loading for Deep Learning Training | 2026 | SIGMOD |
| 6 | 9,729 | Scalable Graph Convolutional Network Training on Distributed-Memory Systems | 2023 | VLDB |
| 7 | 9,414 | Model-Parallel Model Selection for Deep Learning Systems | 2021 | SIGMOD |
| 8 | 8,321 | FusionFlow: Accelerating Data Preprocessing for Machine Learning with CPU-GPU Cooperation | 2024 | VLDB |
| 9 | 3,764 | FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data Pipeline | 2023 | VLDB |
| 10 | 1,079 | Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML | 2014 | VLDB |