Data-efficient Online Training for Direct Alignment in LLMs
Summary: DOTA makes online preference alignment data-efficient by using theoretically grounded Preference Perplexity for gradient-based example selection. Its iterative end-to-end strategy avoids generating responses for all candidates, cutting data-generation cost 3× without sacrificing downstream performance. (summarized by gpt-5.6-luna on Aug 28 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Chi Zhang (Beijing Institute of Technology)
- 2. Jiacheng Wang (Beijing Institute of Technology)
- 3. Kun He (Renmin University of China)
- 4. Chengliang Chai (Beijing Institute of Technology)
- 5. Yunpeng Zhang (Beijing Institute of Technology)
- 6. Yuping Wang (Beijing Institute of Technology)
- 7. Xu Zhou (Hunan University)
- 8. Linan Zheng (University of Arizona)
- 9. Lijun Wu (Shanghai AI Laboratory)
- 10. Conghui He (Shanghai AI Laboratory)
- 11. Lei Cao (Massachusetts Institute of Technology)
BibTeX Citation
@article{zhang_vldb26,
title = {{Data-efficient Online Training for Direct Alignment in LLMs}},
author = {Zhang, Chi and Wang, Jiacheng and He, Kun and Chai, Chengliang and Zhang, Yunpeng and Wang, Yuping and Zhou, Xu and Zheng, Linan and Wu, Lijun and He, Conghui and Cao, Lei},
journal = {PVLDB},
series = {{VLDB} '26},
volume = {19},
number = {10},
pages = {2741--2754},
doi = {10.14778/3828612.3828628},
url = {https://doi.org/10.14778/3828612.3828628},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 5,149 | LLM for Data Management | 2024 | VLDB |
| 2 | 8,888 | Optimized Batch Prompting for Cost-effective LLMs | 2025 | VLDB |
| 3 | 10,792 | Replacing Multi-Step Assembly of Data Preparation Pipelines with One-Step LLM Pipeline Generation for Table QA | 2026 | VLDB |
| 4 | 10,869 | DeepPrep: An LLM-Powered Agentic System for Autonomous Data Preparation | 2026 | VLDB |
| 5 | 8,316 | ThriftLLM: On Cost-Effective Selection of Large Language Models for Classification Queries | 2025 | VLDB |
| 6 | 10,435 | DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization | 2026 | SIGMOD |
| 7 | 10,711 | BRIEF: Bi-level Coreset Selection for Efficient Instruction Tuning in LLMs | 2026 | VLDB |
| 8 | 8,997 | Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees | 2026 | SIGMOD |
| 9 | 7,090 | LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning | 2026 | VLDB |
| 10 | 11,061 | LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning | 2026 | VLDB |