DBScholar

Back to papers

GoldMiner: Elastic Scaling of Training Data Pre-Processing Pipelines for Deep Learning

Summary: GoldMiner decouples data pre-processing from model training with stateless data workers that elastically pool cluster resources. By automatically extracting stateless pre-processing from pipelines, it scales across nodes, delivering up to 12.1x faster jobs and up to 2.5x better GPU utilization in large clusters. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
6758
Venue
SIGMOD
Year
2023
Pagerank
6.2840712e-05
Overall Rank
5,282 | 63.77%
DOI
10.1145/3589773

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{zhao_sigmod23,
        title = {{GoldMiner: Elastic Scaling of Training Data Pre-Processing Pipelines for Deep Learning}},
        author = {Zhao, Hanyu and Yang, Zhi and Cheng, Yu and Tian, Chao and Ren, Shiru and Xiao, Wencong and Yuan, Man and Chen, Langshi and Liu, Kaibo and Zhang, Yang and Li, Yong and Lin, Wei},
        series = {{SIGMOD} '23},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3589773},
        url = {https://dl.acm.org/doi/10.1145/3589773},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 5 of 5 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 3 of 3 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers