Magneto: Combining Small and Large Language Models for Schema Matching
Summary: Magneto uses a two-phase retrieval+reranking pipeline that leverages cheap small LMs to generate candidates and powerful LLMs to rerank, trading minimal compute for high matching accuracy. Novel contributions: LLM-synthesized self-supervised SLM fine-tuning, effective reranking prompts, and a challenging biomedical benchmark. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Yurong Liu (New York University)
- 2. Eduardo H. M. Pena (Federal University of Technology Paraná; New York University)
- 3. Aécio Santos (New York University)
- 4. Eden Wu (New York University)
- 5. Juliana Freire (New York University)
BibTeX Citation
@article{liu_vldb25,
title = {{Magneto: Combining Small and Large Language Models for Schema Matching}},
author = {Liu, Yurong and Pena, Eduardo H. M. and Santos, Aécio and Wu, Eden and Freire, Juliana},
journal = {PVLDB},
series = {{VLDB} '25},
volume = {18},
number = {8},
pages = {2681--2694},
doi = {10.14778/3742728.3742757},
url = {https://doi.org/10.14778/3742728.3742757},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 12 of 12 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 20 of 20 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next