PrepBench: How Far Are We from Natural-Language-Driven Data Preparation?
Summary: PrepBench benchmarks NL-driven data preparation beyond code generation, testing interactive intent disambiguation, prep-code synthesis, and code-to-workflow translation on realistic, multi-step tasks. Results show current LLM agents remain far from reliably realizing this paradigm. (summarized by gpt-5.6-luna on Aug 28 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Jingzhe Xu (Tsinghua University)
- 2. Rui Wang (Tsinghua University)
- 3. Jiannan Wang (Tsinghua University)
- 4. Guoliang Li (Tsinghua University)
BibTeX Citation
@article{xu_vldb26,
title = {{PrepBench: How Far Are We from Natural-Language-Driven Data Preparation?}},
author = {Xu, Jingzhe and Wang, Rui and Wang, Jiannan and Li, Guoliang},
journal = {PVLDB},
series = {{VLDB} '26},
volume = {19},
number = {10},
pages = {2866--2879},
doi = {10.14778/3828612.3828638},
url = {https://doi.org/10.14778/3828612.3828638},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 11 of 11 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next