| 420 |
Can Foundation Models Wrangle Your Data? |
2023 |
VLDB |
0.00018824506 |
| 1,916 |
Annotating Columns with Pre-trained Language Models |
2022 |
SIGMOD |
9.5789978e-05 |
| 1,997 |
RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation |
2021 |
VLDB |
9.4175365e-05 |
| 2,210 |
Table-GPT: Table Fine-tuned GPT for Diverse Table Tasks |
2024 |
SIGMOD |
9.0173497e-05 |
| 2,216 |
Semantics-aware Dataset Discovery from Data Lakes with Contextualized Column-based Representation Learning |
2023 |
VLDB |
9.0031577e-05 |
| 2,351 |
SANTOS: Relationship-based Semantic Table Union Search |
2023 |
SIGMOD |
8.7846455e-05 |
| 2,364 |
Chorus: Foundation Models for Unified Data Discovery and Exploration |
2024 |
VLDB |
8.7622416e-05 |
| 2,888 |
GitTables: A Large-Scale Corpus of Relational Tables |
2023 |
SIGMOD |
8.0622723e-05 |
| 3,063 |
DeepJoin: Joinable Table Discovery with Pre-trained Language Models |
2023 |
VLDB |
7.8644242e-05 |
| 3,541 |
How Large Language Models Will Disrupt Data Management |
2023 |
VLDB |
7.387031e-05 |
| 3,644 |
Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data Integration |
2023 |
SIGMOD |
7.3007456e-05 |
| 4,076 |
Integrating Data Lake Tables |
2023 |
VLDB |
6.9824526e-05 |
| 4,455 |
ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models |
2024 |
VLDB |
6.7496546e-05 |
| 4,626 |
PreQR: Pre-training Representation for SQL Understanding |
2022 |
SIGMOD |
6.6622779e-05 |
| 4,704 |
Leva: Boosting Machine Learning Performance with Relational Embedding Data Augmentation |
2022 |
SIGMOD |
6.6170966e-05 |
| 4,796 |
Transformers for Tabular Data Representation: A Tutorial on Models and Applications |
2022 |
VLDB |
6.5714621e-05 |
| 5,113 |
Logical and Physical Optimizations for SQL Query Execution over Large Language Models |
2025 |
SIGMOD |
6.4232517e-05 |
| 5,793 |
Observatory: Characterizing Embeddings of Relational Tables |
2024 |
VLDB |
6.1462772e-05 |
| 5,805 |
Auto-Tables: Synthesizing Multi-Step Transformations to Relationalize Tables without Using Examples |
2023 |
VLDB |
6.1423731e-05 |
| 5,834 |
Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial Networks |
2023 |
VLDB |
6.132242e-05 |
| 6,284 |
A Unified Transferable Model for ML-Enhanced DBMS |
2022 |
CIDR |
5.9902184e-05 |
| 6,487 |
DTT: An Example-Driven Tabular Transformer for Joinability by Leveraging Large Language Models |
2024 |
SIGMOD |
5.9225411e-05 |
| 7,434 |
Cross Modal Data Discovery over Structured and Unstructured Data Lakes |
2023 |
VLDB |
5.6805318e-05 |
| 8,027 |
WarpGate: A Semantic Join Discovery System for Cloud Data Warehouses |
2023 |
CIDR |
5.5619731e-05 |
| 8,334 |
RECA: Related Tables Enhanced Column Semantic Type Annotation Framework |
2023 |
VLDB |
5.5056757e-05 |
| 8,727 |
Towards Foundation Database Models |
2025 |
CIDR |
5.4404112e-05 |
| 8,735 |
ANN Softmax: Acceleration of Extreme Classification Training |
2022 |
VLDB |
5.4393764e-05 |
| 8,736 |
Watchog: A Light-weight Contrastive Learning based Framework for Column Annotation |
2023 |
SIGMOD |
5.4393613e-05 |
| 8,775 |
Unveiling Challenges for LLMs in Enterprise Data Engineering |
2026 |
VLDB |
5.4311509e-05 |
| 8,781 |
CtxPipe: Context-aware Data Preparation Pipeline Construction for Machine Learning |
2024 |
SIGMOD |
5.4311509e-05 |
| 8,939 |
Making Table Understanding Work in Practice |
2022 |
CIDR |
5.4076395e-05 |
| 9,042 |
VerifAI: Verified Generative AI |
2024 |
CIDR |
5.3900354e-05 |
| 9,300 |
GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language Models |
2024 |
SIGMOD |
5.3503577e-05 |
| 9,400 |
TabulaX: Leveraging Large Language Models for Multi-Class Table Transformations |
2025 |
VLDB |
5.3341661e-05 |
| 9,475 |
Data Imputation with Limited Data Redundancy Using Data Lakes |
2025 |
VLDB |
5.3246578e-05 |
| 9,767 |
Data Augmentation for ML-driven Data Preparation and Integration |
2021 |
VLDB |
5.2759752e-05 |
| 9,960 |
QueryArtisan: Generating Data Manipulation Codes for Ad-hoc Analysis in Data Lakes |
2025 |
VLDB |
5.2142386e-05 |
| 10,059 |
Burr: A Benchmark for Ontology Learning from Relational Databases |
2026 |
SIGMOD |
5.1725247e-05 |
| 10,109 |
Retrieve-and-Verify: A Table Context Selection Framework for Accurate Column Annotations |
2026 |
SIGMOD |
5.1725247e-05 |
| 10,142 |
AutoDDG: Automated Dataset Description Generation using Large Language Models |
2026 |
SIGMOD |
5.1725247e-05 |
| 10,197 |
Qualitative Join Discovery in Data Lakes using Examples |
2026 |
SIGMOD |
5.1725247e-05 |
| 10,268 |
OpenSQL: Data-Efficient Text-to-SQL for Open-Source LLMs via Synthesized Intermediate Supervision |
2026 |
VLDB |
5.1725247e-05 |
| 10,507 |
PLM4NDV: Minimizing Data Access for Number of Distinct Values Estimation with Pre-trained Language Models |
2025 |
SIGMOD |
5.1725247e-05 |
| 10,519 |
Table Overlap Estimation through Graph Embeddings |
2025 |
SIGMOD |
5.1725247e-05 |
| 10,597 |
Birdie: Natural Language-Driven Table Discovery Using Differentiable Search Index |
2025 |
VLDB |
5.1725247e-05 |
| 10,759 |
Cents: A Flexible and Cost-Effective Framework for LLM-Based Table Understanding |
2025 |
VLDB |
5.1725247e-05 |
| 10,760 |
OmniMatch: Joinability Discovery in Data Products |
2025 |
VLDB |
5.1725247e-05 |
| 10,848 |
Panel on Neural Relational Data: Tabular Foundation Models, LLMs... or both? |
2025 |
VLDB |
5.1725247e-05 |
| 10,954 |
Determining the Largest Overlap between Tables |
2024 |
SIGMOD |
5.1725247e-05 |
| 10,976 |
Unstructured Data Fusion for Schema and Data Extraction |
2024 |
SIGMOD |
5.1725247e-05 |