DBScholar

Back to papers

Magneto: Combining Small and Large Language Models for Schema Matching

Summary: Magneto uses a two-phase retrieval+reranking pipeline that leverages cheap small LMs to generate candidates and powerful LLMs to rerank, trading minimal compute for high matching accuracy. Novel contributions: LLM-synthesized self-supervised SLM fine-tuning, effective reranking prompts, and a challenging biomedical benchmark. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
14099
Venue
VLDB
Year
2025
Pagerank
6.279939e-05
Overall Rank
5,293 | 63.69%
DOI
10.14778/3742728.3742757

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{liu_vldb25,
        title = {{Magneto: Combining Small and Large Language Models for Schema Matching}},
        author = {Liu, Yurong and Pena, Eduardo H. M. and Santos, Aécio and Wu, Eden and Freire, Juliana},
        journal = {PVLDB},
        series = {{VLDB} '25},
        volume = {18},
        number = {8},
        pages = {2681--2694},
        doi = {10.14778/3742728.3742757},
        url = {https://doi.org/10.14778/3742728.3742757},
        year = {2025}
}

Incoming Citations (Sorted by Pagerank)

Showing 7 of 7 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 20 of 20 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
297 Generic Schema Matching with Cupid 2001 VLDB 0.00022157284
390 COMA - A system for flexible combination of schema matching approaches 2002 VLDB 0.00019382486
420 Can Foundation Models Wrangle Your Data? 2023 VLDB 0.00018789852
493 Data Integration for the Relational Web 2009 VLDB 0.00017558709
1,158 Schema and Ontology Matching with COMA++ 2005 SIGMOD 0.00011899424
1,402 Creating Embeddings of Heterogeneous Relational Datasets for Data Integration Tasks 2020 SIGMOD 0.00010888094
1,923 Annotating Columns with Pre-trained Language Models 2022 SIGMOD 9.4789109e-05
2,099 Table-GPT: Table Fine-tuned GPT for Diverse Table Tasks 2024 SIGMOD 9.1682353e-05
2,192 Semantics-aware Dataset Discovery from Data Lakes with Contextualized Column-based Representation Learning 2023 VLDB 8.9767241e-05
2,216 Open Data Integration 2018 VLDB 8.9374127e-05
2,242 CHORUS: Foundation Models for Unified Data Discovery and Exploration 2024 VLDB 8.8823802e-05
2,315 SANTOS: Relationship-based Semantic Table Union Search 2023 SIGMOD 8.7605101e-05
3,073 DeepJoin: Joinable Table Discovery with Pre-trained Language Models 2023 VLDB 7.785842e-05
3,436 Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data Integration 2023 SIGMOD 7.4157897e-05
3,456 Automatic Discovery of Attributes in Relational Databases 2011 SIGMOD 7.3998323e-05
4,515 ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models 2024 VLDB 6.6492389e-05
4,836 Transformers for Tabular Data Representation: A Tutorial on Models and Applications 2022 VLDB 6.4848269e-05
5,694 Top-K Generation of Integrated Schemas Based on Directed and Weighted Correspondences 2009 SIGMOD 6.1183376e-05
7,325 ADnEV: Cross-Domain Schema Matching using Deep Similarity Matrix Adjustment and Evaluation 2020 VLDB 5.6443833e-05
8,252 Sampling Methods for Inner Product Sketching 2024 VLDB 5.4574671e-05
Previous Page 1 / 1 Next

Semantically Similar Papers