DBScholar

Back to papers

Snuba: Automating Weak Supervision to Label Training Data

Summary: Snuba automates weak supervision by generating task-specific labeling heuristics from a small labeled set to label a large unlabeled corpus. It grows coverage iteratively with a statistical termination guarantee, finishing under five minutes and beating handcrafted rules by 9.74 F1 and semi-supervised baselines by 14.35 F1. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hf201c174eb06c565
Venue
VLDB
Year
2019
Pagerank
0.000119406
Overall Rank
1,120 | 92.48%
DOI
10.14778/3291264.3291268
PDF
Download (CC BY-NC-ND 4.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{varma_vldb19,
        title = {{Snuba: Automating Weak Supervision to Label Training Data}},
        author = {Varma, Paroma and Ré, Christopher},
        journal = {PVLDB},
        series = {{VLDB} '19},
        volume = {12},
        number = {3},
        pages = {223--236},
        doi = {10.14778/3291264.3291268},
        url = {https://doi.org/10.14778/3291264.3291268},
        year = {2019}
}

Incoming Citations (Sorted by Pagerank)

Showing 26 of 26 citing papers.

Rank Citing Paper Year Venue Pagerank
496 Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes 2024 VLDB 0.00017318538
3,627 The Role of Massively Multi-Task and Weak Supervision in Software 2.0 2019 CIDR 7.1492314e-05
3,956 Smile: A System to Support Machine Learning on EEG Data at Scale 2019 VLDB 6.8992907e-05
4,544 GOGGLES: Automatic Image Labeling with Affinity Coding 2020 SIGMOD 6.5439123e-05
5,078 Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond 2021 SIGMOD 6.2832055e-05
5,102 Towards Benchmarking Feature Type Inference for AutoML Platforms 2021 SIGMOD 6.2705617e-05
5,179 Explainable AI: Foundations, Applications, Opportunities for Data Management Research 2022 SIGMOD 6.2382214e-05
5,853 Adaptive Rule Discovery for Labeling Text Data 2021 SIGMOD 5.9690904e-05
5,948 VOCAL: Video Organization and Interactive Compositional AnaLytics 2022 CIDR 5.934521e-05
6,666 UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads 2022 VLDB 5.7144587e-05
6,940 Inspector Gadget: A Data Programming-based Labeling System for Industrial Images 2021 VLDB 5.637052e-05
7,479 Witan: Unsupervised Labelling Function Generation for Assisted Data Programming 2022 VLDB 5.5142801e-05
7,881 CHEF: A Cheap and Fast Pipeline for Iteratively Cleaning Label Uncertainties 2021 VLDB 5.4322729e-05
8,672 CrowdGame: A Game-Based Crowdsourcing System for Cost-Effective Data Labeling 2019 SIGMOD 5.2901608e-05
8,698 Nemo: Guiding and Contextualizing Weak Supervision for Interactive Data Programming 2022 VLDB 5.2880532e-05
8,764 Exploratory Training: When Annotators Learn About Data 2023 SIGMOD 5.2815275e-05
9,044 LANCET: Labeling Complex Data at Scale 2021 VLDB 5.2308178e-05
9,713 Ground Truth Inference for Weakly Supervised Entity Matching 2023 SIGMOD 5.1350726e-05
10,048 The Image Calculator: 10x Faster Image-AI Inference by Replacing JPEG with Self-designing Storage Format 2024 SIGMOD 5.0896901e-05
10,116 Data Augmentation for ML-driven Data Preparation and Integration 2021 VLDB 5.0765311e-05
10,217 CORAL: Collaborative Automatic Labeling System based on Large Language Models 2024 VLDB 5.0572653e-05
10,954 Morphing-based Compression for Data-centric ML Pipelines 2026 VLDB 4.9769913e-05
11,183 A Cost-Effective LLM-based Approach to Identify Wildlife Trafficking in Online Marketplaces 2025 SIGMOD 4.9769913e-05
11,227 WeShap: Weak Supervision Source Evaluation with Shapley Values 2025 VLDB 4.9769913e-05
11,726 Steered Training Data Generation for Learned Semantic Type Detection 2023 SIGMOD 4.9769913e-05
12,131 Leveraging Organizational Resources to Adapt Models to New Data Modalities 2020 VLDB 4.9769913e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 5 of 5 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
205 Snorkel: Rapid Training Data Creation with Weak Supervision 2018 VLDB 0.00025171314
428 Big Data Integration 2013 VLDB 0.00018457189
1,284 Fusing Data with Correlations 2014 SIGMOD 0.00011197034
3,063 Fonduer: Knowledge Base Construction from Richly Formatted Data 2018 SIGMOD 7.689108e-05
3,793 SLiMFast: Guaranteed Results for Data Fusion and Source Reliability 2017 SIGMOD 7.0156889e-05
Previous Page 1 / 1 Next

Semantically Similar Papers