Steered Training Data Generation for Learned Semantic Type Detection
Summary: STEER adapts learned semantic type extraction to unseen data lakes via a data-programming labeling framework. Steered-Labeling generates labeled data for numeric and non-numeric columns to fine-tune learned models, boosting performance across four data lakes. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Sven Langenecker (Duale Hochschule Baden-Württemberg; Läpple AG; Technical University of Darmstadt)
- 2. Christoph Sturm (Duale Hochschule Baden-Württemberg)
- 3. Christian Schalles (Duale Hochschule Baden-Württemberg)
- 4. Carsten Binnig (German National Research Center for Information Technology; Technical University of Darmstadt)
BibTeX Citation
@inproceedings{langenecker_sigmod23,
title = {{Steered Training Data Generation for Learned Semantic Type Detection}},
author = {Langenecker, Sven and Sturm, Christoph and Schalles, Christian and Binnig, Carsten},
series = {{SIGMOD} '23},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3589786},
url = {https://dl.acm.org/doi/10.1145/3589786},
year = {2023}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 205 | Snorkel: Rapid Training Data Creation with Weak Supervision | 2018 | VLDB | 0.00025181304 |
| 377 | TURL: Table Understanding through Representation Learning | 2021 | VLDB | 0.00019570264 |
| 1,120 | Snuba: Automating Weak Supervision to Label Training Data | 2019 | VLDB | 0.00011946047 |
| 1,780 | Annotating Columns with Pre-trained Language Models | 2022 | SIGMOD | 9.6560923e-05 |
| 2,092 | Sato: Contextual Semantic Type Detection in Tables | 2020 | VLDB | 9.0626928e-05 |
| 2,521 | GitTables: A Large-Scale Corpus of Relational Tables | 2023 | SIGMOD | 8.3479333e-05 |
| 3,470 | Automatic Discovery of Attributes in Relational Databases | 2011 | SIGMOD | 7.2765721e-05 |
| 7,473 | Witan: Unsupervised Labelling Function Generation for Assisted Data Programming | 2022 | VLDB | 5.5168918e-05 |
| 9,244 | Making Table Understanding Work in Practice | 2022 | CIDR | 5.2056825e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 11,898 | Automated Relational Data Explanation using External Semantic Knowledge | 2022 | VLDB |
| 2 | 11,568 | SeLeP: Learning Based Semantic Prefetching for Exploratory Database Workloads | 2024 | VLDB |
| 3 | 6,488 | Synthesizing Type-Detection Logic for Rich Semantic Data Types using Open-source Code | 2018 | SIGMOD |
| 4 | 4,683 | AutoSteer: Learned Query Optimization for Any SQL Database | 2023 | VLDB |
| 5 | 5,106 | Data-Driven Domain Discovery for Structured Datasets | 2020 | VLDB |
| 6 | 2,092 | Sato: Contextual Semantic Type Detection in Tables | 2020 | VLDB |
| 7 | 9,229 | Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables | 2025 | SIGMOD |
| 8 | 5,100 | Towards Benchmarking Feature Type Inference for AutoML Platforms | 2021 | SIGMOD |
| 9 | 4,589 | ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models | 2024 | VLDB |
| 10 | 1,852 | Semantics-aware Dataset Discovery from Data Lakes with Contextualized Column-based Representation Learning | 2023 | VLDB |