Data-Juicer: A One-Stop Data Processing System for Large Language Models
Summary: Data-Juicer: a one-stop data-processing system for constructing diverse data recipes to train and evaluate LLMs. Unique data-management features—fine-grained pipelines with 50+ operators, heterogeneous sources, visual auto-evaluation, and distributed-LLM integration—achieving up to 7.45% average gains across 16 benchmarks and 17.5% higher GPT-4 win-rate. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Daoyuan Chen (Alibaba)
- 2. Yilun Huang (Alibaba)
- 3. Zhijian Ma (Alibaba)
- 4. Hesen Chen (Alibaba)
- 5. Xuchen Pan (Alibaba)
- 6. Ce Ge (Alibaba)
- 7. Dawei Gao (Alibaba)
- 8. Yuexiang Xie (Alibaba)
- 9. Zhaoyang Liu (Alibaba)
- 10. Jinyang Gao (Alibaba)
- 11. Yaliang Li (Alibaba)
- 12. Bolin Ding (Alibaba)
- 13. Jingren Zhou (Alibaba)
BibTeX Citation
@inproceedings{chen_sigmod24,
title = {{Data-Juicer: A One-Stop Data Processing System for Large Language Models}},
author = {Chen, Daoyuan and Huang, Yilun and Ma, Zhijian and Chen, Hesen and Pan, Xuchen and Ge, Ce and Gao, Dawei and Xie, Yuexiang and Liu, Zhaoyang and Gao, Jinyang and Li, Yaliang and Ding, Bolin and Zhou, Jingren},
series = {{SIGMOD} '24},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3626246.3653385},
url = {https://dl.acm.org/doi/10.1145/3626246.3653385},
year = {2024}
}
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 7,809 | Can Large Language Models Be Query Optimizer for Relational Databases? | 2026 | SIGMOD | 5.5399022e-05 |
| 10,472 | Mixtera: A Data Plane for Foundation Model Training | 2026 | SIGMOD | 5.093636e-05 |
| 10,614 | LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning | 2026 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 0 of 0 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 6,545 | DTT: An Example-Driven Tabular Transformer for Joinability by Leveraging Large Language Models | 2024 | SIGMOD |
| 2 | 8,906 | Unveiling Challenges for LLMs in Enterprise Data Engineering | 2026 | VLDB |
| 3 | 10,736 | Sentence to Model: Cost-Effective Data Collection LLM Agent | 2025 | SIGMOD |
| 4 | 7,439 | AOP: Automated and Interactive LLM Pipeline Orchestration for Answering Complex Queries | 2025 | CIDR |
| 5 | 10,733 | ScaleLLM: A Technique for Scalable LLM-augmented Data Systems | 2025 | SIGMOD |
| 6 | 10,882 | CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines | 2025 | VLDB |
| 7 | 13,303 | Demonstrating CatDB: LLM-based Generation of Data-centric ML Pipelines | 2025 | SIGMOD |
| 8 | 6,101 | LLM for Data Management | 2024 | VLDB |
| 9 | 713 | Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes | 2024 | VLDB |
| 10 | 10,614 | LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning | 2026 | VLDB |