DBScholar

Back to papers

Magneto: Combining Small and Large Language Models for Schema Matching

Summary: Magneto uses a two-phase retrieval+reranking pipeline that leverages cheap small LMs to generate candidates and powerful LLMs to rerank, trading minimal compute for high matching accuracy. Novel contributions: LLM-synthesized self-supervised SLM fine-tuning, effective reranking prompts, and a challenging biomedical benchmark. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
he45ee1c73b7c425b
Venue
VLDB
Year
2025
Pagerank
6.7270169e-05
Overall Rank
4,218 | 71.65%
DOI
10.14778/3742728.3742757

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{liu_vldb25,
        title = {{Magneto: Combining Small and Large Language Models for Schema Matching}},
        author = {Liu, Yurong and Pena, Eduardo H. M. and Santos, Aécio and Wu, Eden and Freire, Juliana},
        journal = {PVLDB},
        series = {{VLDB} '25},
        volume = {18},
        number = {8},
        pages = {2681--2694},
        doi = {10.14778/3742728.3742757},
        url = {https://doi.org/10.14778/3742728.3742757},
        year = {2025}
}

Incoming Citations (Sorted by Pagerank)

Showing 12 of 12 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 20 of 20 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
296 Generic Schema Matching with Cupid 2001 VLDB 0.00021867512
329 Can Foundation Models Wrangle Your Data? 2023 VLDB 0.00020858443
407 COMA - A system for flexible combination of schema matching approaches 2002 VLDB 0.00019026291
484 Data Integration for the Relational Web 2009 VLDB 0.00017548541
1,146 Schema and Ontology Matching with COMA++ 2005 SIGMOD 0.00011815973
1,391 Creating Embeddings of Heterogeneous Relational Datasets for Data Integration Tasks 2020 SIGMOD 0.00010816237
1,780 Annotating Columns with Pre-trained Language Models 2022 SIGMOD 9.6560923e-05
1,852 Semantics-aware Dataset Discovery from Data Lakes with Contextualized Column-based Representation Learning 2023 VLDB 9.5004635e-05
1,932 CHORUS: Foundation Models for Unified Data Discovery and Exploration 2024 VLDB 9.34643e-05
1,978 Table-GPT: Table Fine-tuned GPT for Diverse Table Tasks 2024 SIGMOD 9.2730152e-05
1,988 SANTOS: Relationship-based Semantic Table Union Search 2023 SIGMOD 9.2482448e-05
2,158 Open Data Integration 2018 VLDB 8.941016e-05
2,382 DeepJoin: Joinable Table Discovery with Pre-trained Language Models 2023 VLDB 8.5458532e-05
3,470 Automatic Discovery of Attributes in Relational Databases 2011 SIGMOD 7.2765721e-05
3,473 Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data Integration 2023 SIGMOD 7.2728706e-05
4,589 ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models 2024 VLDB 6.5144711e-05
4,630 Transformers for Tabular Data Representation: A Tutorial on Models and Applications 2022 VLDB 6.4953891e-05
5,765 Top-K Generation of Integrated Schemas Based on Directed and Weighted Correspondences 2009 SIGMOD 6.0026982e-05
6,883 ADnEV: Cross-Domain Schema Matching using Deep Similarity Matrix Adjustment and Evaluation 2020 VLDB 5.6565418e-05
8,421 Sampling Methods for Inner Product Sketching 2024 VLDB 5.3350162e-05
Previous Page 1 / 1 Next

Semantically Similar Papers