DBScholar

Back to papers

Weak-to-Strong Prompts with Lightweight-to-Powerful LLMs for High-Accuracy, Low-Cost, and Explainable Data Transformation

Summary: MegaTran turns underspecified data-transformation requests into structured prompts with a lightweight LLM, then uses a powerful LLM for code generation. Checklist-based reflection and LazyRAG improve correctness, explainability, and cost efficiency, yielding 2.2–26.1% accuracy gains. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
14073
Venue
VLDB
Year
2025
Pagerank
5.093636e-05
Overall Rank
10,867 | 25.45%
DOI
10.14778/3742728.3742734

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

BibTeX Citation

@article{li_vldb25,
        title = {{Weak-to-Strong Prompts with Lightweight-to-Powerful LLMs for High-Accuracy, Low-Cost, and Explainable Data Transformation}},
        author = {Li, Changlun and Yang, Chenyu and Luo, Yuyu and Fan, Ju and Tang, Nan},
        journal = {PVLDB},
        series = {{VLDB} '25},
        volume = {18},
        number = {8},
        pages = {2371--2384},
        doi = {10.14778/3742728.3742734},
        url = {https://doi.org/10.14778/3742728.3742734},
        year = {2025}
}

Incoming Citations (Sorted by Pagerank)

Showing 1 of 1 citing papers.

Rank Citing Paper Year Venue Pagerank
10,587 LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning 2026 VLDB 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 20 of 20 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
94 Potter's Wheel: An Interactive Data Cleaning System 2001 VLDB 0.00034616103
420 Can Foundation Models Wrangle Your Data? 2023 VLDB 0.00018789852
725 NADEEF: A Commodity Data Cleaning System 2013 SIGMOD 0.00014617251
963 The Data Civilizer System 2017 CIDR 0.00012935145
1,160 BlinkFill: Semi-supervised Programming By Example for Syntactic String Transformations 2016 VLDB 0.00011883163
1,165 Foofah: Transforming Data By Example 2017 SIGMOD 0.00011860616
1,465 Synthesizing Entity Matching Rules by Examples 2018 VLDB 0.00010689571
2,019 RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation 2021 VLDB 9.2983994e-05
2,242 CHORUS: Foundation Models for Unified Data Discovery and Exploration 2024 VLDB 8.8823802e-05
2,978 Transform-Data-by-Example (TDE): An Extensible Search Engine for Data Transformations 2018 VLDB 7.9030989e-05
2,981 Towards Dependable Data Repairing with Fixing Rules 2014 SIGMOD 7.8960114e-05
3,436 Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data Integration 2023 SIGMOD 7.4157897e-05
3,787 Combining Small Language Models and Large Language Models for Zero-Shot NL2SQL 2024 VLDB 7.1249098e-05
4,412 Auto-Transform: Learning-to-Transform by Patterns 2020 VLDB 6.7168614e-05
4,990 Explaining Dataset Changes for Semantic Data Versioning with Explain-Da-V 2023 VLDB 6.4105738e-05
5,620 DataPrep.EDA: Task-Centric Exploratory Data Analysis for Statistical Modeling in Python 2021 SIGMOD 6.1482864e-05
6,545 DTT: An Example-Driven Tabular Transformer for Joinability by Leveraging Large Language Models 2024 SIGMOD 5.8454882e-05
7,552 Trinity: An Extensible Synthesis Framework for Data Science 2019 VLDB 5.6005243e-05
8,242 Data Civilizer 2.0: A Holistic Framework for Data Preparation and Analytics 2019 VLDB 5.459288e-05
9,653 CoClean: Collaborative Data Cleaning 2020 SIGMOD 5.2425585e-05
Previous Page 1 / 1 Next

Semantically Similar Papers