Steered Training Data Generation for Learned Semantic Type Detection
Summary: STEER adapts learned semantic type extraction to unseen data lakes via a data-programming labeling framework. Steered-Labeling generates labeled data for numeric and non-numeric columns to fine-tune learned models, boosting performance across four data lakes. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Sven Langenecker (Duale Hochschule Baden-Württemberg; Läpple AG; Technical University of Darmstadt)
- 2. Christoph Sturm (Duale Hochschule Baden-Württemberg)
- 3. Christian Schalles (Duale Hochschule Baden-Württemberg)
- 4. Carsten Binnig (German National Research Center for Information Technology; Technical University of Darmstadt)
BibTeX Citation
@inproceedings{langenecker_sigmod23,
title = {{Steered Training Data Generation for Learned Semantic Type Detection}},
author = {Langenecker, Sven and Sturm, Christoph and Schalles, Christian and Binnig, Carsten},
series = {{SIGMOD} '23},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3589786},
url = {https://dl.acm.org/doi/10.1145/3589786},
year = {2023}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 205 | Snorkel: Rapid Training Data Creation with Weak Supervision | 2018 | VLDB | 0.00025235185 |
| 397 | TURL: Table Understanding through Representation Learning | 2021 | VLDB | 0.00019278189 |
| 1,094 | Snuba: Automating Weak Supervision to Label Training Data | 2019 | VLDB | 0.00012214617 |
| 1,923 | Annotating Columns with Pre-trained Language Models | 2022 | SIGMOD | 9.4789109e-05 |
| 2,223 | Sato: Contextual Semantic Type Detection in Tables | 2020 | VLDB | 8.9189986e-05 |
| 2,790 | GitTables: A Large-Scale Corpus of Relational Tables | 2023 | SIGMOD | 8.1200509e-05 |
| 3,456 | Automatic Discovery of Attributes in Relational Databases | 2011 | SIGMOD | 7.3998323e-05 |
| 7,331 | Witan: Unsupervised Labelling Function Generation for Assisted Data Programming | 2022 | VLDB | 5.6435171e-05 |
| 9,067 | Making Table Understanding Work in Practice | 2022 | CIDR | 5.3251649e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 11,590 | Automated Relational Data Explanation using External Semantic Knowledge | 2022 | VLDB |
| 2 | 11,232 | SeLeP: Learning Based Semantic Prefetching for Exploratory Database Workloads | 2024 | VLDB |
| 3 | 4,929 | AutoSteer: Learned Query Optimization for Any SQL Database | 2023 | VLDB |
| 4 | 6,436 | Synthesizing Type-Detection Logic for Rich Semantic Data Types using Open-source Code | 2018 | SIGMOD |
| 5 | 5,125 | Data-Driven Domain Discovery for Structured Datasets | 2020 | VLDB |
| 6 | 2,223 | Sato: Contextual Semantic Type Detection in Tables | 2020 | VLDB |
| 7 | 10,785 | Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables | 2025 | SIGMOD |
| 8 | 5,051 | Towards Benchmarking Feature Type Inference for AutoML Platforms | 2021 | SIGMOD |
| 9 | 4,515 | ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models | 2024 | VLDB |
| 10 | 2,192 | Semantics-aware Dataset Discovery from Data Lakes with Contextualized Column-based Representation Learning | 2023 | VLDB |