Adda: Towards Efficient in-Database Feature Generation via LLM-based Agents
Summary: Adda enables in-database feature generation via LLM-based agents for ML analytics; natural-language tasks generate SQL-ready feature code compiled as UDFs. On 14 datasets, 5 ML tasks: up to 33.2% AUC gains and 100x latency vs Madlib. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Kuan Lu (Zhejiang University)
- 2. Zhihui Yang (Zhejiang University)
- 3. Sai Wu (Zhejiang University)
- 4. Ruichen Xia (Zhejiang University)
- 5. Dongxiang Zhang (Zhejiang University)
- 6. Gang Chen (Zhejiang University)
BibTeX Citation
@inproceedings{lu_sigmod25,
title = {{Adda: Towards Efficient in-Database Feature Generation via LLM-based Agents}},
author = {Lu, Kuan and Yang, Zhihui and Wu, Sai and Xia, Ruichen and Zhang, Dongxiang and Chen, Gang},
series = {{SIGMOD} '25},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3725262},
url = {https://dl.acm.org/doi/10.1145/3725262},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 2 of 2 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 10,232 | EncoderForge: Generating Efficient SQL for Encoders in Machine Learning Inference Pipelines | 2026 | SIGMOD | 5.093636e-05 |
| 10,432 | Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management | 2026 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 10 of 10 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 34 | The Design Of Postgres | 1986 | SIGMOD | 0.00049302774 |
| 106 | The MADlib Analytics Library or MAD Skills, the SQL | 2012 | VLDB | 0.00033539462 |
| 420 | Can Foundation Models Wrangle Your Data? | 2023 | VLDB | 0.00018789852 |
| 2,019 | RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation | 2021 | VLDB | 9.2983994e-05 |
| 2,347 | Vertica-ML: Distributed Machine Learning in Vertica Database | 2020 | SIGMOD | 8.7157552e-05 |
| 2,786 | DB4ML – An In-Memory Database Kernel with Machine Learning Support | 2020 | SIGMOD | 8.1207221e-05 |
| 2,865 | End-to-end Optimization of Machine Learning Prediction Queries | 2022 | SIGMOD | 8.0180243e-05 |
| 4,114 | Optimizing Machine Learning Inference Queries with Correlative Proxy Models | 2022 | VLDB | 6.8941194e-05 |
| 4,475 | UlTraMan: A Unified Platform for Big Trajectory Data Management and Analytics | 2018 | VLDB | 6.6803264e-05 |
| 8,197 | SMARTFEAT: Efficient Feature Construction through Feature-Level Foundation Model Interactions | 2024 | CIDR | 5.4691464e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 4,886 | Chat2Data: An Interactive Data Analysis System with RAG, Vector Databases and LLMs | 2024 | VLDB |
| 2 | 4,308 | AutoTQA: Towards Autonomous Tabular Question Answering through Multi-Agent Large Language Models | 2024 | VLDB |
| 3 | 11,013 | Towards Automated Cross-domain Exploratory Data Analysis through Large Language Models | 2025 | VLDB |
| 4 | 8,906 | Unveiling Challenges for LLMs in Enterprise Data Engineering | 2026 | VLDB |
| 5 | 713 | Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes | 2024 | VLDB |
| 6 | 1,121 | ARDA: Automatic Relational Data Augmentation for Machine Learning | 2020 | VLDB |
| 7 | 4,053 | Database-Agnostic Workload Management | 2019 | CIDR |
| 8 | 10,431 | AutoDDG: Automated Dataset Description Generation using Large Language Models | 2026 | SIGMOD |
| 9 | 3,757 | Panda: Performance Debugging for Databases using LLM Agents | 2024 | CIDR |
| 10 | 6,101 | LLM for Data Management | 2024 | VLDB |