DBScholar

Back to papers

Data-efficient Online Training for Direct Alignment in LLMs

Summary: DOTA makes online preference alignment data-efficient by using theoretically grounded Preference Perplexity for gradient-based example selection. Its iterative end-to-end strategy avoids generating responses for all candidates, cutting data-generation cost 3× without sacrificing downstream performance. (summarized by gpt-5.6-luna on Aug 28 2026)

Paper ID
h6858fd8dcee581eb
Venue
VLDB
Year
2026
Pagerank
4.9793485e-05
Overall Rank
10,821 | 27.25%
DOI
10.14778/3828612.3828628

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

BibTeX Citation

@article{zhang_vldb26,
        title = {{Data-efficient Online Training for Direct Alignment in LLMs}},
        author = {Zhang, Chi and Wang, Jiacheng and He, Kun and Chai, Chengliang and Zhang, Yunpeng and Wang, Yuping and Zhou, Xu and Zheng, Linan and Wu, Lijun and He, Conghui and Cao, Lei},
        journal = {PVLDB},
        series = {{VLDB} '26},
        volume = {19},
        number = {10},
        pages = {2741--2754},
        doi = {10.14778/3828612.3828628},
        url = {https://doi.org/10.14778/3828612.3828628},
        year = {2026}
}

Incoming Citations (Sorted by Pagerank)

Showing 0 of 0 citing papers.

Rank Citing Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 9 of 9 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers