AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent Framework
Summary: AutoPrep: an LLM multi-agent framework for question-aware table prep in TQA, decomposing tasks (column derivation/filtering, value normalization) across Planner/Programmer/Executor agents. Uses Chain-of-Clauses reasoning and tool-augmented codegen to produce executable plans, improving SOTA on real TQA benchmarks. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Meihao Fan (Renmin University of China)
- 2. Ju Fan (Renmin University of China)
- 3. Nan Tang (Hong Kong University of Science and Technology)
- 4. Lei Cao (University of Arizona)
- 5. Guoliang Li (Tsinghua University)
- 6. Xiaoyong Du (Renmin University of China)
BibTeX Citation
@article{fan_vldb25,
title = {{AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent Framework}},
author = {Fan, Meihao and Fan, Ju and Tang, Nan and Cao, Lei and Li, Guoliang and Du, Xiaoyong},
journal = {PVLDB},
series = {{VLDB} '25},
volume = {18},
number = {10},
pages = {3504--3517},
doi = {10.14778/3748191.3748211},
url = {https://doi.org/10.14778/3748191.3748211},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 10,285 | Reward-SQL: Boosting Text-to-SQL via Stepwise Execution-Aware Reasoning and Process-Supervised Rewards | 2026 | SIGMOD | 5.093636e-05 |
| 10,304 | VecBench: A Controllable Benchmark for Filtered Vector Search: [Experiments & Analysis] | 2026 | SIGMOD | 5.093636e-05 |
| 10,537 | TACO: A Benchmark for Open-Domain Text-to-SQL with Ambiguous and Cross-Database Queries | 2026 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 10 of 10 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next