Database Paper Browser

Back to papers

Snorkel: Rapid Training Data Creation with Weak Supervision

Summary: Snorkel enables rapid ML training from weak supervision via labeling functions with unknown accuracies. End-to-end data programming denoises labels without ground truth, with a tradeoff optimizer, showing speedups and accuracy gains over hand labeling. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
11742
Venue
VLDB
Year
2018
Pagerank
0.00030532082
Overall Rank
252 | 98.26%
DOI
10.14778/3157794.3157797

Incoming Non-self Citations Over Time

Authors

Incoming Citations (Sorted by Pagerank)

Showing 50 of 70 citing papers.

Rank Citing Paper Year Venue Pagerank
293 Deep Learning for Entity Matching: A Design Space Exploration 2018 SIGMOD 0.00028661817
1,088 Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes 2024 VLDB 0.00014158762
1,218 Snuba: Automating Weak Supervision to Label Training Data 2019 VLDB 0.00013221309
1,340 HoloDetect: Few-Shot Learning for Error Detection 2019 SIGMOD 0.00012492795
1,666 HELIX: Holistic Optimization for Accelerating Iterative Machine Learning 2019 VLDB 0.00010955907
1,942 SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model Debugging 2021 SIGMOD 0.00010010569
1,992 Automatically Generating Data Exploration Sessions Using Deep Reinforcement Learning 2020 SIGMOD 9.8415851e-05
2,325 DBPal: A Fully Pluggable NL2SQL Training Pipeline 2020 SIGMOD 9.0277894e-05
2,831 Smile: A System to Support Machine Learning on EEG Data at Scale 2019 VLDB 8.0485807e-05
2,845 VolcanoML: Speeding up End-to-End AutoML via Scalable Search Space Decomposition 2021 VLDB 8.0301674e-05
2,960 The Role of Massively Multi-Task and Weak Supervision in Software 2.0 2019 CIDR 7.8103118e-05
3,305 Fonduer: Knowledge Base Construction from Richly Formatted Data 2018 SIGMOD 7.2417724e-05
3,513 spade: Synthesizing Data Quality Assertions for Large Language Model Pipelines 2024 VLDB 7.0203909e-05
3,765 Ember: No-Code Context Enrichment via Similarity-Based Keyless Joins 2022 VLDB 6.7760748e-05
4,197 Overton: A Data System for Monitoring and Improving Machine-Learned Products 2020 CIDR 6.3625568e-05
4,455 AutoOD: Automatic Outlier Detection 2023 SIGMOD 6.1644904e-05
4,471 GOGGLES: Automatic Image Labeling with Affinity Coding 2020 SIGMOD 6.1496765e-05
4,587 MB2: Decomposed Behavior Modeling for Self-Driving Database Management Systems 2021 SIGMOD 6.0594195e-05
4,596 Data Integration and Machine Learning: A Natural Synergy 2018 SIGMOD 6.0540725e-05
4,749 ODIN: Automated Drift Detection and Recovery in Video Analytics 2020 VLDB 5.9428569e-05
4,866 OmniFair: A Declarative System for Model-Agnostic Group Fairness in Machine Learning 2021 SIGMOD 5.8620848e-05
4,875 Explainable AI: Foundations, Applications, Opportunities for Data Management Research 2022 SIGMOD 5.8552996e-05
5,244 Towards Benchmarking Feature Type Inference for AutoML Platforms 2021 SIGMOD 5.6021738e-05
5,257 Snorkel DryBell: A Case Study in Deploying Weak Supervision at Industrial Scale 2019 SIGMOD 5.5975788e-05
5,358 Adaptive Rule Discovery for Labeling Text Data 2021 SIGMOD 5.5507062e-05
5,386 Selective Data Acquisition in the Wild for Model Charging 2022 VLDB 5.5346315e-05
5,420 Mining an "Anti-Knowledge Base" from Wikipedia Updates with Applications to Fact Checking and Beyond 2020 VLDB 5.5154502e-05
5,872 Demonstration of Panda: A Weakly Supervised Entity Matching System 2021 VLDB 5.2908178e-05
5,965 Automatic Data Acquisition for Deep Learning 2021 VLDB 5.2476363e-05
5,974 Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond 2021 SIGMOD 5.2458154e-05
6,047 MDedup: Duplicate Detection with Matching Dependencies 2020 VLDB 5.2355891e-05
6,122 VOCAL: Video Organization and Interactive Compositional AnaLytics 2022 CIDR 5.1966758e-05
6,139 Finding Label and Model Errors in Perception Data With Learned Observation Assertions 2022 SIGMOD 5.1893488e-05
6,225 Managing ML Pipelines: Feature Stores and the Coming Wave of Embedding Ecosystems 2021 VLDB 5.1425119e-05
6,242 Optimizing In-memory Database Engine for AI-powered On-line Decision Augmentation Using Persistent Memory 2021 VLDB 5.1351431e-05
6,507 Expand your Training Limits! Generating Training Data for ML-based Data Management 2021 SIGMOD 5.0273414e-05
6,873 Cost-Effective Data Annotation using Game-Based Crowdsourcing 2019 VLDB 4.8963037e-05
7,241 Data Integration and Machine Learning: A Natural Synergy 2018 VLDB 4.7867685e-05
7,284 Witan: Unsupervised Labelling Function Generation for Assisted Data Programming 2022 VLDB 4.7716465e-05
7,642 Cross Modal Data Discovery over Structured and Unstructured Data Lakes 2023 VLDB 4.6856127e-05
7,656 Nautilus: An Optimized System for Deep Transfer Learning over Evolving Training Datasets 2022 SIGMOD 4.6826896e-05
7,798 CHEF: A Cheap and Fast Pipeline for Iteratively Cleaning Label Uncertainties 2021 VLDB 4.6438053e-05
8,058 iFlipper: Label Flipping for Individual Fairness 2023 SIGMOD 4.5903348e-05
8,183 SHiFT: An Efficient, Flexible Search Engine for Transfer Learning 2023 VLDB 4.5615358e-05
8,285 Nemo: Guiding and Contextualizing Weak Supervision for Interactive Data Programming 2022 VLDB 4.5392079e-05
8,337 CrowdGame: A Game-Based Crowdsourcing System for Cost-Effective Data Labeling 2019 SIGMOD 4.5385651e-05
8,515 UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads 2022 VLDB 4.4901466e-05
8,713 LANCET: Labeling Complex Data at Scale 2021 VLDB 4.4577046e-05
9,196 Hyper-Tune: Towards Efficient Hyper-parameter Tuning at Scale 2022 VLDB 4.3723457e-05
9,259 Improving Information Extraction from Visually Rich Documents using Visual Span Representations 2021 VLDB 4.3648789e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 5 of 5 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
192 HoloClean: Holistic Data Repairs with Probabilistic Inference 2017 VLDB 0.00035692958
372 A Bayesian Approach to Discovering Truth from Conflicting Sources for Data Integration 2012 VLDB 0.00025371138
394 Big Data Integration 2013 VLDB 0.0002447017
906 Fusing Data with Correlations 2014 SIGMOD 0.00015420344
3,900 SLiMFast: Guaranteed Results for Data Fusion and Source Reliability 2017 SIGMOD 6.649432e-05
Previous Page 1 / 1 Next

Semantically Similar Papers