Leva: Boosting Machine Learning Performance with Relational Embedding Data Augmentation
Summary: Leva constructs a relational embedding by graphifying the database and learning vectors that summarize the entire data. Downstream supervision filters noisy graph signals, reducing cross-relational feature engineering and data-discovery burden, and boosting ML performance on classification/regression tasks. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Zixuan Zhao (University of Chicago)
- 2. Raul Castro Fernandez (University of Chicago)
BibTeX Citation
@inproceedings{zhao_sigmod22,
title = {{Leva: Boosting Machine Learning Performance with Relational Embedding Data Augmentation}},
author = {Zhao, Zixuan and Fernandez, Raul Castro},
series = {{SIGMOD} '22},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3514221.3517891},
url = {https://dl.acm.org/doi/10.1145/3514221.3517891},
year = {2022}
}
Incoming Citations (Sorted by Pagerank)
Showing 9 of 9 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,192 | Semantics-aware Dataset Discovery from Data Lakes with Contextualized Column-based Representation Learning | 2023 | VLDB | 8.9767241e-05 |
| 3,536 | How Large Language Models Will Disrupt Data Management | 2023 | VLDB | 7.3297343e-05 |
| 5,291 | DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular Data | 2023 | SIGMOD | 6.2801343e-05 |
| 5,485 | Pneuma: Leveraging LLMs for Tabular Data Representation and Retrieval in an End-to-End System | 2025 | SIGMOD | 6.2013279e-05 |
| 7,220 | Solo: Data Discovery Using Natural Language Questions Via A Self-Supervised Approach | 2023 | SIGMOD | 5.6679948e-05 |
| 8,860 | Watchog: A Light-weight Contrastive Learning based Framework for Column Annotation | 2023 | SIGMOD | 5.3564029e-05 |
| 10,987 | OmniMatch: Joinability Discovery in Data Products | 2025 | VLDB | 5.093636e-05 |
| 11,186 | Unstructured Data Fusion for Schema and Data Extraction | 2024 | SIGMOD | 5.093636e-05 |
| 11,262 | Enriching Relations with Additional Attributes for ER | 2024 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 10 of 10 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 397 | TURL: Table Understanding through Representation Learning | 2021 | VLDB | 0.00019278189 |
| 489 | Distributed Representations of Tuples for Entity Resolution | 2018 | VLDB | 0.0001761456 |
| 764 | To Join or Not to Join? Thinking Twice about Joins before Feature Selection | 2016 | SIGMOD | 0.00014226652 |
| 1,121 | ARDA: Automatic Relational Data Augmentation for Machine Learning | 2020 | VLDB | 0.00012093059 |
| 1,402 | Creating Embeddings of Heterogeneous Relational Datasets for Data Integration Tasks | 2020 | SIGMOD | 0.00010888094 |
| 1,447 | Auctus: A Dataset Search Engine for Data Discovery and Augmentation | 2021 | VLDB | 0.00010760327 |
| 1,495 | LSH Ensemble: Internet-Scale Domain Search | 2016 | VLDB | 0.00010571481 |
| 3,228 | Correlation Sketches for Approximate Join-Correlation Queries | 2021 | SIGMOD | 7.6215176e-05 |
| 3,681 | Are Key-Foreign Key Joins Safe to Avoid when Learning High-Capacity Classifiers? | 2018 | VLDB | 7.2037388e-05 |
| 7,880 | Learning Over Dirty Data Without Cleaning | 2020 | SIGMOD | 5.5244204e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 10,557 | Database Views as Explanations for Relational Deep Learning | 2026 | VLDB |
| 2 | 2,865 | End-to-end Optimization of Machine Learning Prediction Queries | 2022 | SIGMOD |
| 3 | 9,470 | Powering In-Database Dynamic Model Slicing for Structured Data Analytics | 2024 | VLDB |
| 4 | 9,155 | Adda: Towards Efficient in-Database Feature Generation via LLM-based Agents | 2025 | SIGMOD |
| 5 | 1,235 | Towards Linear Algebra over Normalized Data | 2017 | VLDB |
| 6 | 10,042 | Scalable and Usable Relational Learning With Automatic Language Bias | 2021 | SIGMOD |
| 7 | 10,757 | Data Enhancement for Binary Classification of Relational Data | 2025 | SIGMOD |
| 8 | 9,962 | Structure-Aware Machine Learning over Multi-Relational Databases | 2021 | SIGMOD |
| 9 | 1,121 | ARDA: Automatic Relational Data Augmentation for Machine Learning | 2020 | VLDB |
| 10 | 1,402 | Creating Embeddings of Heterogeneous Relational Datasets for Data Integration Tasks | 2020 | SIGMOD |