SMARTFEAT: Efficient Feature Construction through Feature-Level Foundation Model Interactions
Summary: SMARTFEAT uses foundation models to synthesize informative new features via feature-level FM interactions, guided by an intelligent operator selector to avoid exhaustive operator/feature combinations. A function generator emits efficient dataframe/lambda transformations (not per-row FM calls), enabling scalable, cost- and latency-efficient feature construction for large datasets. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Yin Lin (University of Michigan)
- 2. Bolin Ding (Alibaba)
- 3. H. V. Jagadish (University of Michigan)
- 4. Jingren Zhou (Alibaba)
BibTeX Citation
@inproceedings{lin_cidr24,
address = {Amsterdam, Netherlands},
series = {{CIDR} '24},
title = {{SMARTFEAT: Efficient Feature Construction through Feature-Level Foundation Model Interactions}},
booktitle = {Proceedings of the {Conference} on {Innovative} {Data} {Systems} {Research}},
author = {Lin, Yin and Ding, Bolin and Jagadish, H. V. and Zhou, Jingren},
year = {2024}
}
Incoming Citations (Sorted by Pagerank)
Showing 2 of 2 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 9,155 | Adda: Towards Efficient in-Database Feature Generation via LLM-based Agents | 2025 | SIGMOD | 5.3122817e-05 |
| 10,882 | CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines | 2025 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 4 of 4 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 141 | Deep Entity Matching with Pre-Trained Language Models | 2021 | VLDB | 0.0002964847 |
| 420 | Can Foundation Models Wrangle Your Data? | 2023 | VLDB | 0.00018789852 |
| 1,351 | Detecting Data Errors: Where are we and what needs to be done? | 2016 | VLDB | 0.00011064851 |
| 2,019 | RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation | 2021 | VLDB | 9.2983994e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 420 | Can Foundation Models Wrangle Your Data? | 2023 | VLDB |
| 2 | 2,179 | Enabling and Optimizing Non-linear Feature Interactions in Factorized Linear Algebra | 2019 | SIGMOD |
| 3 | 12,739 | SMART: A Tool for Semantic-Driven Creation of Complex XML Mappings | 2005 | SIGMOD |
| 4 | 9,155 | Adda: Towards Efficient in-Database Feature Generation via LLM-based Agents | 2025 | SIGMOD |
| 5 | 5,804 | A Relational Framework for Classifier Engineering | 2017 | PODS |
| 6 | 8,854 | Towards Foundation Database Models | 2025 | CIDR |
| 7 | 11,744 | CAFE: Constraint-Aware Feature Extraction from Large Databases | 2020 | CIDR |
| 8 | 11,177 | FeatureLTE: Learning to Estimate Feature Importance | 2024 | SIGMOD |
| 9 | 13,302 | DANTE: Hybrid AI System for Context-Aware Interpretable Feature Engineering | 2025 | SIGMOD |
| 10 | 5,297 | An Integrated Development Environment for Faster Feature Engineering | 2014 | VLDB |