CAFE: Constraint-Aware Feature Extraction from Large Databases
Summary: CAFE extracts features from large DBs while enforcing high-level constraints (consistency, interpretability, fairness) by mapping them to low-level pruning strategies and using an inverted index to find candidate columns. An optimizer-like planner uses sample-based estimates, models strategy dependencies, and orders pruning to maximize downstream ML accuracy while bounding runtime and feature quality. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Mahdi Esmailoghli (Technical University of Berlin)
- 2. Ziawasch Abedjan (Technical University of Berlin)
BibTeX Citation
@inproceedings{esmailoghli_cidr20,
address = {Amsterdam, Netherlands},
series = {{CIDR} '20},
title = {{CAFE: Constraint-Aware Feature Extraction from Large Databases}},
booktitle = {Proceedings of the {Conference} on {Innovative} {Data} {Systems} {Research}},
author = {Esmailoghli, Mahdi and Abedjan, Ziawasch},
year = {2020}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 3,905 | Automated Feature Engineering for Algorithmic Fairness | 2021 | VLDB | 7.029145e-05 |
| 5,479 | MATE: Multi-Attribute Table Extraction | 2022 | VLDB | 6.204351e-05 |
| 11,062 | Data Discovery in Data Lakes: Operations, Indexes, Systems | 2025 | VLDB | 5.093636e-05 |
| 11,674 | Enforcing Constraints for Machine Learning Systems via Declarative Feature Selection: An Experimental Study | 2021 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 2 of 2 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 367 | InfoGather: Entity Augmentation and Attribute Discovery By Holistic Matching with Web Tables | 2012 | SIGMOD | 0.00019979463 |
| 963 | The Data Civilizer System | 2017 | CIDR | 0.00012935145 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 295 | Accelerating Machine Learning Inference with Probabilistic Predicates | 2018 | SIGMOD |
| 2 | 4,969 | Optimization of Constrained Frequent Set Queries with 2-variable Constraints | 1999 | SIGMOD |
| 3 | 4,114 | Optimizing Machine Learning Inference Queries with Correlative Proxy Models | 2022 | VLDB |
| 4 | 5,804 | A Relational Framework for Classifier Engineering | 2017 | PODS |
| 5 | 9,353 | On Efficient Approximate Queries over Machine Learning Models | 2023 | VLDB |
| 6 | 11,674 | Enforcing Constraints for Machine Learning Systems via Declarative Feature Selection: An Experimental Study | 2021 | SIGMOD |
| 7 | 1,535 | Exploratory Mining and Pruning Optimizations of Constrained Association Rules | 1998 | SIGMOD |
| 8 | 465 | An End-to-End Learning-based Cost Estimator | 2020 | VLDB |
| 9 | 7,112 | Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning | 2023 | VLDB |
| 10 | 9,556 | CAFE: Towards Compact, Adaptive, and Fast Embedding for Large-scale Recommendation Models | 2024 | SIGMOD |