DBScholar

Back to papers

Snuba: Automating Weak Supervision to Label Training Data

Summary: Snuba automates weak supervision by generating task-specific labeling heuristics from a small labeled set to label a large unlabeled corpus. It grows coverage iteratively with a statistical termination guarantee, finishing under five minutes and beating handcrafted rules by 9.74 F1 and semi-supervised baselines by 14.35 F1. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hf201c174eb06c565
Venue
VLDB
Year
2019
Pagerank
0.00011946047
Overall Rank
1,120 | 92.48%
DOI
10.14778/3291264.3291268

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{varma_vldb19,
        title = {{Snuba: Automating Weak Supervision to Label Training Data}},
        author = {Varma, Paroma and Ré, Christopher},
        journal = {PVLDB},
        series = {{VLDB} '19},
        volume = {12},
        number = {3},
        pages = {223--236},
        doi = {10.14778/3291264.3291268},
        url = {https://doi.org/10.14778/3291264.3291268},
        year = {2019}
}

Incoming Citations (Sorted by Pagerank)

Showing 26 of 26 citing papers.

Rank Citing Paper Year Venue Pagerank
501 Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes 2024 VLDB 0.00017267905
3,627 The Role of Massively Multi-Task and Weak Supervision in Software 2.0 2019 CIDR 7.1514557e-05
3,955 Smile: A System to Support Machine Learning on EEG Data at Scale 2019 VLDB 6.9025583e-05
4,543 GOGGLES: Automatic Image Labeling with Affinity Coding 2020 SIGMOD 6.5470116e-05
5,076 Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond 2021 SIGMOD 6.2860582e-05
5,100 Towards Benchmarking Feature Type Inference for AutoML Platforms 2021 SIGMOD 6.27349e-05
5,178 Explainable AI: Foundations, Applications, Opportunities for Data Management Research 2022 SIGMOD 6.2411759e-05
5,850 Adaptive Rule Discovery for Labeling Text Data 2021 SIGMOD 5.9719174e-05
5,970 VOCAL: Video Organization and Interactive Compositional AnaLytics 2022 CIDR 5.9316412e-05
6,662 UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads 2022 VLDB 5.7171651e-05
6,937 Inspector Gadget: A Data Programming-based Labeling System for Industrial Images 2021 VLDB 5.6397218e-05
7,473 Witan: Unsupervised Labelling Function Generation for Assisted Data Programming 2022 VLDB 5.5168918e-05
7,876 CHEF: A Cheap and Fast Pipeline for Iteratively Cleaning Label Uncertainties 2021 VLDB 5.4348457e-05
8,664 CrowdGame: A Game-Based Crowdsourcing System for Cost-Effective Data Labeling 2019 SIGMOD 5.2926663e-05
8,690 Nemo: Guiding and Contextualizing Weak Supervision for Interactive Data Programming 2022 VLDB 5.2905577e-05
8,756 Exploratory Training: When Annotators Learn About Data 2023 SIGMOD 5.2840289e-05
9,036 LANCET: Labeling Complex Data at Scale 2021 VLDB 5.2332952e-05
9,708 Ground Truth Inference for Weakly Supervised Entity Matching 2023 SIGMOD 5.1374628e-05
10,043 The Image Calculator: 10x Faster Image-AI Inference by Replacing JPEG with Self-designing Storage Format 2024 SIGMOD 5.0921006e-05
10,112 Data Augmentation for ML-driven Data Preparation and Integration 2021 VLDB 5.0789354e-05
10,211 CORAL: Collaborative Automatic Labeling System based on Large Language Models 2024 VLDB 5.0596605e-05
10,945 Morphing-based Compression for Data-centric ML Pipelines 2026 VLDB 4.9793485e-05
11,174 A Cost-Effective LLM-based Approach to Identify Wildlife Trafficking in Online Marketplaces 2025 SIGMOD 4.9793485e-05
11,218 WeShap: Weak Supervision Source Evaluation with Shapley Values 2025 VLDB 4.9793485e-05
11,720 Steered Training Data Generation for Learned Semantic Type Detection 2023 SIGMOD 4.9793485e-05
12,125 Leveraging Organizational Resources to Adapt Models to New Data Modalities 2020 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 5 of 5 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
205 Snorkel: Rapid Training Data Creation with Weak Supervision 2018 VLDB 0.00025181304
427 Big Data Integration 2013 VLDB 0.00018465558
1,284 Fusing Data with Correlations 2014 SIGMOD 0.00011202115
3,061 Fonduer: Knowledge Base Construction from Richly Formatted Data 2018 SIGMOD 7.6927483e-05
3,791 SLiMFast: Guaranteed Results for Data Fusion and Source Reliability 2017 SIGMOD 7.0189865e-05
Previous Page 1 / 1 Next

Semantically Similar Papers