DBScholar

Back to papers

Weak-to-Strong Prompts with Lightweight-to-Powerful LLMs for High-Accuracy, Low-Cost, and Explainable Data Transformation

Summary: MegaTran turns underspecified data-transformation requests into structured prompts with a lightweight LLM, then uses a powerful LLM for code generation. Checklist-based reflection and LazyRAG improve correctness, explainability, and cost efficiency, yielding 2.2–26.1% accuracy gains. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
hddef9211b78f72e8
Venue
VLDB
Year
2025
Pagerank
4.9769913e-05
Overall Rank
11,278 | 24.20%
DOI
10.14778/3742728.3742734
PDF
Download (CC BY-NC-ND 4.0)

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

BibTeX Citation

@article{li_vldb25,
        title = {{Weak-to-Strong Prompts with Lightweight-to-Powerful LLMs for High-Accuracy, Low-Cost, and Explainable Data Transformation}},
        author = {Li, Changlun and Yang, Chenyu and Luo, Yuyu and Fan, Ju and Tang, Nan},
        journal = {PVLDB},
        series = {{VLDB} '25},
        volume = {18},
        number = {8},
        pages = {2371--2384},
        doi = {10.14778/3742728.3742734},
        url = {https://doi.org/10.14778/3742728.3742734},
        year = {2025}
}

Incoming Citations (Sorted by Pagerank)

Showing 3 of 3 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 20 of 20 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
95 Potter's Wheel: An Interactive Data Cleaning System 2001 VLDB 0.00034367518
329 Can Foundation Models Wrangle Your Data? 2023 VLDB 0.00020867521
697 NADEEF: A Commodity Data Cleaning System 2013 SIGMOD 0.00014687805
972 The Data Civilizer System 2017 CIDR 0.00012757732
1,095 BlinkFill: Semi-supervised Programming By Example for Syntactic String Transformations 2016 VLDB 0.00012048043
1,115 Foofah: Transforming Data By Example 2017 SIGMOD 0.00011958625
1,493 Synthesizing Entity Matching Rules by Examples 2018 VLDB 0.00010500948
1,929 CHORUS: Foundation Models for Unified Data Discovery and Exploration 2024 VLDB 9.3551286e-05
1,993 RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation 2021 VLDB 9.2348951e-05
2,744 Transform-Data-by-Example (TDE): An Extensible Search Engine for Data Transformations 2018 VLDB 8.064168e-05
2,967 Towards Dependable Data Repairing with Fixing Rules 2014 SIGMOD 7.8016942e-05
3,473 Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data Integration 2023 SIGMOD 7.2697311e-05
3,489 Combining Small Language Models and Large Language Models for Zero-Shot NL2SQL 2024 VLDB 7.2594757e-05
4,426 Auto-Transform: Learning-to-Transform by Patterns 2020 VLDB 6.601896e-05
5,100 Explaining Dataset Changes for Semantic Data Versioning with Explain-Da-V 2023 VLDB 6.271266e-05
5,642 DataPrep.EDA: Task-Centric Exploratory Data Analysis for Statistical Modeling in Python 2021 SIGMOD 6.0519777e-05
5,686 DTT: An Example-Driven Tabular Transformer for Joinability by Leveraging Large Language Models 2024 SIGMOD 6.0347166e-05
7,293 Data Civilizer 2.0: A Holistic Framework for Data Preparation and Analytics 2019 VLDB 5.5616488e-05
7,701 Trinity: An Extensible Synthesis Framework for Data Science 2019 VLDB 5.4730936e-05
7,854 CoClean: Collaborative Data Cleaning 2020 SIGMOD 5.4380042e-05
Previous Page 1 / 1 Next

Semantically Similar Papers