RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation
Summary: RPT: denoising tuple-to-tuple autoencoder; Transformer encoder-decoder unifies BERT and GPT. Pre-trained, it enables data cleaning, auto-completion, and normalization and annotation, plus few-shot and collaborative ER/IE. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Nan Tang (Hamad Bin Khalifa University; Qatar Computing Research Institute)
- 2. Ju Fan (Renmin University of China)
- 3. Fangyi Li (Renmin University of China)
- 4. Jianhong Tu (Renmin University of China)
- 5. Xiaoyong Du (Renmin University of China)
- 6. Guoliang Li (Tsinghua University)
- 7. Sam Madden (Massachusetts Institute of Technology)
- 8. Mourad Ouzzani (Hamad Bin Khalifa University; Qatar Computing Research Institute)
BibTeX Citation
@article{tang_vldb21,
title = {{RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation}},
author = {Tang, Nan and Fan, Ju and Li, Fangyi and Tu, Jianhong and Du, Xiaoyong and Li, Guoliang and Madden, Sam and Ouzzani, Mourad},
journal = {PVLDB},
series = {{VLDB} '21},
volume = {14},
number = {8},
pages = {1254--1261},
doi = {10.14778/3457390.3457391},
url = {https://doi.org/10.14778/3457390.3457391},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 27 of 27 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 24 of 24 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 10,831 | PrepBench: How Far Are We from Natural-Language-Driven Data Preparation? | 2026 | VLDB |
| 2 | 9,956 | Graph Transformers for Query Plan Representation: Potentials and Challenges | 2025 | VLDB |
| 3 | 1,978 | Table-GPT: Table Fine-tuned GPT for Diverse Table Tasks | 2024 | SIGMOD |
| 4 | 10,418 | BAT: Target-Instance-Free Data Preparation Synthesis via LLM-Driven Tree Search | 2026 | SIGMOD |
| 5 | 11,061 | LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning | 2026 | VLDB |
| 6 | 4,630 | Transformers for Tabular Data Representation: A Tutorial on Models and Applications | 2022 | VLDB |
| 7 | 7,270 | AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent Framework | 2025 | VLDB |
| 8 | 4,424 | Auto-Transform: Learning-to-Transform by Patterns | 2020 | VLDB |
| 9 | 5,687 | DTT: An Example-Driven Tabular Transformer for Joinability by Leveraging Large Language Models | 2024 | SIGMOD |
| 10 | 10,869 | DeepPrep: An LLM-Powered Agentic System for Autonomous Data Preparation | 2026 | VLDB |