DBScholar

Back to papers

Snorkel: Rapid Training Data Creation with Weak Supervision

Summary: Snorkel enables rapid ML training from weak supervision via labeling functions with unknown accuracies. End-to-end data programming denoises labels without ground truth, with a tradeoff optimizer, showing speedups and accuracy gains over hand labeling. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
11929
Venue
VLDB
Year
2018
Pagerank
0.00025235185
Overall Rank
205 | 98.60%
DOI
10.14778/3157794.3157797

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{ratner_vldb18,
        title = {{Snorkel: Rapid Training Data Creation with Weak Supervision}},
        author = {Ratner, Alexander and Bach, Stephen H. and Ehrenberg, Henry and Fries, Jason and Wu, Sen and Ré, Christopher},
        journal = {PVLDB},
        series = {{VLDB} '18},
        volume = {11},
        number = {3},
        pages = {269--282},
        doi = {10.14778/3157794.3157797},
        url = {https://doi.org/10.14778/3157794.3157797},
        year = {2018}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 70 citing papers.

Rank Citing Paper Year Venue Pagerank
176 Deep Learning for Entity Matching: A Design Space Exploration 2018 SIGMOD 0.00027191081
713 Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes 2024 VLDB 0.00014672521
946 HoloDetect: Few-Shot Learning for Error Detection 2019 SIGMOD 0.00013054126
1,094 Snuba: Automating Weak Supervision to Label Training Data 2019 VLDB 0.00012214617
1,569 HELIX: Holistic Optimization for Accelerating Iterative Machine Learning 2019 VLDB 0.00010335423
2,058 Automatically Generating Data Exploration Sessions Using Deep Reinforcement Learning 2020 SIGMOD 9.2480248e-05
2,079 DBPal: A Fully Pluggable NL2SQL Training Pipeline 2020 SIGMOD 9.2060425e-05
2,273 SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model Debugging 2021 SIGMOD 8.8230899e-05
3,192 Fonduer: Knowledge Base Construction from Richly Formatted Data 2018 SIGMOD 7.65035e-05
3,272 VolcanoML: Speeding up End-to-End AutoML via Scalable Search Space Decomposition 2021 VLDB 7.5775321e-05
3,589 Ember: No-Code Context Enrichment via Similarity-Based Keyless Joins 2022 VLDB 7.2812353e-05
3,600 The Role of Massively Multi-Task and Weak Supervision in Software 2.0 2019 CIDR 7.2709969e-05
3,869 Smile: A System to Support Machine Learning on EEG Data at Scale 2019 VLDB 7.0609879e-05
3,947 Overton: A Data System for Monitoring and Improving Machine-Learned Products 2020 CIDR 7.0040437e-05
3,961 MB2: Decomposed Behavior Modeling for Self-Driving Database Management Systems 2021 SIGMOD 6.987575e-05
4,062 AutoOD: Automatic Outlier Detection 2023 SIGMOD 6.9309994e-05
4,100 Data Integration and Machine Learning: A Natural Synergy 2018 SIGMOD 6.9016092e-05
4,451 GOGGLES: Automatic Image Labeling with Affinity Coding 2020 SIGMOD 6.6952549e-05
4,621 spade: Synthesizing Data Quality Assertions for Large Language Model Pipelines 2024 VLDB 6.6024412e-05
4,658 OmniFair: A Declarative System for Model-Agnostic Group Fairness in Machine Learning 2021 SIGMOD 6.5817368e-05
4,808 Selective Data Acquisition in the Wild for Model Charging 2022 VLDB 6.4991553e-05
4,924 ODIN: Automated Drift Detection and Recovery in Video Analytics 2020 VLDB 6.4429959e-05
4,966 Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond 2021 SIGMOD 6.4225454e-05
5,051 Towards Benchmarking Feature Type Inference for AutoML Platforms 2021 SIGMOD 6.385354e-05
5,054 Explainable AI: Foundations, Applications, Opportunities for Data Management Research 2022 SIGMOD 6.3843089e-05
5,058 Snorkel DryBell: A Case Study in Deploying Weak Supervision at Industrial Scale 2019 SIGMOD 6.3815523e-05
5,391 MDedup: Duplicate Detection with Matching Dependencies 2020 VLDB 6.2343031e-05
5,548 Mining an "Anti-Knowledge Base" from Wikipedia Updates with Applications to Fact Checking and Beyond 2020 VLDB 6.1790655e-05
5,722 Adaptive Rule Discovery for Labeling Text Data 2021 SIGMOD 6.1089867e-05
5,925 Managing ML Pipelines: Feature Stores and the Coming Wave of Embedding Ecosystems 2021 VLDB 6.0397888e-05
5,966 VOCAL: Video Organization and Interactive Compositional AnaLytics 2022 CIDR 6.0255527e-05
6,024 Expand your Training Limits! Generating Training Data for ML-based Data Management 2021 SIGMOD 6.0031118e-05
6,041 Demonstration of Panda: A Weakly Supervised Entity Matching System 2021 VLDB 5.9976574e-05
6,074 Automatic Data Acquisition for Deep Learning 2021 VLDB 5.9860133e-05
6,105 Optimizing In-memory Database Engine for AI-powered On-line Decision Augmentation Using Persistent Memory 2021 VLDB 5.9743349e-05
6,431 Finding Label and Model Errors in Perception Data With Learned Observation Assertions 2022 SIGMOD 5.8802622e-05
6,538 UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads 2022 VLDB 5.8477764e-05
6,854 iFlipper: Label Flipping for Individual Fairness 2023 SIGMOD 5.753578e-05
6,908 Cost-Effective Data Annotation using Game-Based Crowdsourcing 2019 VLDB 5.7415834e-05
7,331 Witan: Unsupervised Labelling Function Generation for Assisted Data Programming 2022 VLDB 5.6435171e-05
7,424 Data Integration and Machine Learning: A Natural Synergy 2018 VLDB 5.6214566e-05
7,572 Cross Modal Data Discovery over Structured and Unstructured Data Lakes 2023 VLDB 5.5938952e-05
7,656 Nautilus: An Optimized System for Deep Transfer Learning over Evolving Training Datasets 2022 SIGMOD 5.5740571e-05
7,868 CHEF: A Cheap and Fast Pipeline for Iteratively Cleaning Label Uncertainties 2021 VLDB 5.5277527e-05
8,043 Falcon: Fair Active Learning using Multi-armed Bandits 2024 VLDB 5.5013766e-05
8,244 SHiFT: An Efficient, Flexible Search Engine for Transfer Learning 2023 VLDB 5.4587712e-05
8,495 CrowdGame: A Game-Based Crowdsourcing System for Cost-Effective Data Labeling 2019 SIGMOD 5.4139128e-05
8,523 Nemo: Guiding and Contextualizing Weak Supervision for Interactive Data Programming 2022 VLDB 5.4119882e-05
8,876 LANCET: Labeling Complex Data at Scale 2021 VLDB 5.3534114e-05
9,290 Hyper-Tune: Towards Efficient Hyper-parameter Tuning at Scale 2022 VLDB 5.2910774e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 5 of 5 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
112 HoloClean: Holistic Data Repairs with Probabilistic Inference 2017 VLDB 0.00032801121
427 Big Data Integration 2013 VLDB 0.00018661543
504 A Bayesian Approach to Discovering Truth from Conflicting Sources for Data Integration 2012 VLDB 0.00017300628
1,273 Fusing Data with Correlations 2014 SIGMOD 0.00011384191
3,715 SLiMFast: Guaranteed Results for Data Fusion and Source Reliability 2017 SIGMOD 7.1763559e-05
Previous Page 1 / 1 Next

Semantically Similar Papers