DBScholar

Back to papers

PrepBench: How Far Are We from Natural-Language-Driven Data Preparation?

Summary: PrepBench benchmarks NL-driven data preparation beyond code generation, testing interactive intent disambiguation, prep-code synthesis, and code-to-workflow translation on realistic, multi-step tasks. Results show current LLM agents remain far from reliably realizing this paradigm. (summarized by gpt-5.6-luna on Aug 28 2026)

Paper ID
h9c72b2aa51ef3f76
Venue
VLDB
Year
2026
Pagerank
4.9793485e-05
Overall Rank
10,831 | 27.18%
DOI
10.14778/3828612.3828638

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

BibTeX Citation

@article{xu_vldb26,
        title = {{PrepBench: How Far Are We from Natural-Language-Driven Data Preparation?}},
        author = {Xu, Jingzhe and Wang, Rui and Wang, Jiannan and Li, Guoliang},
        journal = {PVLDB},
        series = {{VLDB} '26},
        volume = {19},
        number = {10},
        pages = {2866--2879},
        doi = {10.14778/3828612.3828638},
        url = {https://doi.org/10.14778/3828612.3828638},
        year = {2026}
}

Incoming Citations (Sorted by Pagerank)

Showing 0 of 0 citing papers.

Rank Citing Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 11 of 11 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers