DBScholar

Back to papers

Snuba: Automating Weak Supervision to Label Training Data

Summary: Snuba automates weak supervision by generating task-specific labeling heuristics from a small labeled set to label a large unlabeled corpus. It grows coverage iteratively with a statistical termination guarantee, finishing under five minutes and beating handcrafted rules by 9.74 F1 and semi-supervised baselines by 14.35 F1. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
12129
Venue
VLDB
Year
2019
Pagerank
0.00012214617
Overall Rank
1,094 | 92.50%
DOI
10.14778/3291264.3291268

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{varma_vldb19,
        title = {{Snuba: Automating Weak Supervision to Label Training Data}},
        author = {Varma, Paroma and Ré, Christopher},
        journal = {PVLDB},
        series = {{VLDB} '19},
        volume = {12},
        number = {3},
        pages = {223--236},
        doi = {10.14778/3291264.3291268},
        url = {https://doi.org/10.14778/3291264.3291268},
        year = {2019}
}

Incoming Citations (Sorted by Pagerank)

Showing 26 of 26 citing papers.

Rank Citing Paper Year Venue Pagerank
713 Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes 2024 VLDB 0.00014672521
3,600 The Role of Massively Multi-Task and Weak Supervision in Software 2.0 2019 CIDR 7.2709969e-05
3,869 Smile: A System to Support Machine Learning on EEG Data at Scale 2019 VLDB 7.0609879e-05
4,451 GOGGLES: Automatic Image Labeling with Affinity Coding 2020 SIGMOD 6.6952549e-05
4,966 Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond 2021 SIGMOD 6.4225454e-05
5,051 Towards Benchmarking Feature Type Inference for AutoML Platforms 2021 SIGMOD 6.385354e-05
5,054 Explainable AI: Foundations, Applications, Opportunities for Data Management Research 2022 SIGMOD 6.3843089e-05
5,722 Adaptive Rule Discovery for Labeling Text Data 2021 SIGMOD 6.1089867e-05
5,966 VOCAL: Video Organization and Interactive Compositional AnaLytics 2022 CIDR 6.0255527e-05
6,538 UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads 2022 VLDB 5.8477764e-05
6,816 Inspector Gadget: A Data Programming-based Labeling System for Industrial Images 2021 VLDB 5.7649828e-05
7,331 Witan: Unsupervised Labelling Function Generation for Assisted Data Programming 2022 VLDB 5.6435171e-05
7,868 CHEF: A Cheap and Fast Pipeline for Iteratively Cleaning Label Uncertainties 2021 VLDB 5.5277527e-05
8,495 CrowdGame: A Game-Based Crowdsourcing System for Cost-Effective Data Labeling 2019 SIGMOD 5.4139128e-05
8,523 Nemo: Guiding and Contextualizing Weak Supervision for Interactive Data Programming 2022 VLDB 5.4119882e-05
8,594 Exploratory Training: When Annotators Learn About Data 2023 SIGMOD 5.4053095e-05
8,876 LANCET: Labeling Complex Data at Scale 2021 VLDB 5.3534114e-05
9,560 Ground Truth Inference for Weakly Supervised Entity Matching 2023 SIGMOD 5.2528121e-05
9,881 The Image Calculator: 10x Faster Image-AI Inference by Replacing JPEG with Self-designing Storage Format 2024 SIGMOD 5.2040783e-05
9,930 Data Augmentation for ML-driven Data Preparation and Integration 2021 VLDB 5.1955087e-05
10,021 CORAL: Collaborative Automatic Labeling System based on Large Language Models 2024 VLDB 5.1757914e-05
10,589 Morphing-based Compression for Data-centric ML Pipelines 2026 VLDB 5.093636e-05
10,746 A Cost-Effective LLM-based Approach to Identify Wildlife Trafficking in Online Marketplaces 2025 SIGMOD 5.093636e-05
10,805 WeShap: Weak Supervision Source Evaluation with Shapley Values 2025 VLDB 5.093636e-05
11,406 Steered Training Data Generation for Learned Semantic Type Detection 2023 SIGMOD 5.093636e-05
11,824 Leveraging Organizational Resources to Adapt Models to New Data Modalities 2020 VLDB 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 5 of 5 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
205 Snorkel: Rapid Training Data Creation with Weak Supervision 2018 VLDB 0.00025235185
427 Big Data Integration 2013 VLDB 0.00018661543
1,273 Fusing Data with Correlations 2014 SIGMOD 0.00011384191
3,192 Fonduer: Knowledge Base Construction from Richly Formatted Data 2018 SIGMOD 7.65035e-05
3,715 SLiMFast: Guaranteed Results for Data Fusion and Source Reliability 2017 SIGMOD 7.1763559e-05
Previous Page 1 / 1 Next

Semantically Similar Papers