Database Paper Browser

Back to papers

Snorkel: Rapid Training Data Creation with Weak Supervision

Summary: Snorkel enables rapid ML training from weak supervision via labeling functions with unknown accuracies. End-to-end data programming denoises labels without ground truth, with a tradeoff optimizer, showing speedups and accuracy gains over hand labeling. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
11742
Venue
VLDB
Year
2018
Pagerank
0.00025523636
Overall Rank
204 | 98.59%
DOI
10.14778/3157794.3157797

Incoming Non-self Citations Over Time

Authors

Incoming Citations (Sorted by Pagerank)

Showing 50 of 70 citing papers.

Rank Citing Paper Year Venue Pagerank
181 Deep Learning for Entity Matching: A Design Space Exploration 2018 SIGMOD 0.00026923442
906 Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes 2024 VLDB 0.00013351599
926 HoloDetect: Few-Shot Learning for Error Detection 2019 SIGMOD 0.00013229887
1,080 Snuba: Automating Weak Supervision to Label Training Data 2019 VLDB 0.0001239664
1,536 HELIX: Holistic Optimization for Accelerating Iterative Machine Learning 2019 VLDB 0.00010490973
2,019 Automatically Generating Data Exploration Sessions Using Deep Reinforcement Learning 2020 SIGMOD 9.3828128e-05
2,133 DBPal: A Fully Pluggable NL2SQL Training Pipeline 2020 SIGMOD 9.1819568e-05
2,243 SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model Debugging 2021 SIGMOD 8.9475888e-05
3,144 Fonduer: Knowledge Base Construction from Richly Formatted Data 2018 SIGMOD 7.7688364e-05
3,234 VolcanoML: Speeding up End-to-End AutoML via Scalable Search Space Decomposition 2021 VLDB 7.6850504e-05
3,549 The Role of Massively Multi-Task and Weak Supervision in Software 2.0 2019 CIDR 7.3816787e-05
3,554 Ember: No-Code Context Enrichment via Similarity-Based Keyless Joins 2022 VLDB 7.3770107e-05
3,800 Smile: A System to Support Machine Learning on EEG Data at Scale 2019 VLDB 7.1680016e-05
3,881 Overton: A Data System for Monitoring and Improving Machine-Learned Products 2020 CIDR 7.1066569e-05
3,911 MB2: Decomposed Behavior Modeling for Self-Driving Database Management Systems 2021 SIGMOD 7.0870659e-05
4,059 Data Integration and Machine Learning: A Natural Synergy 2018 SIGMOD 6.9975008e-05
4,221 AutoOD: Automatic Outlier Detection 2023 SIGMOD 6.8862117e-05
4,381 GOGGLES: Automatic Image Labeling with Affinity Coding 2020 SIGMOD 6.7988732e-05
4,545 spade: Synthesizing Data Quality Assertions for Large Language Model Pipelines 2024 VLDB 6.704524e-05
4,597 OmniFair: A Declarative System for Model-Agnostic Group Fairness in Machine Learning 2021 SIGMOD 6.6836728e-05
4,805 Selective Data Acquisition in the Wild for Model Charging 2022 VLDB 6.5663072e-05
4,854 ODIN: Automated Drift Detection and Recovery in Video Analytics 2020 VLDB 6.5424395e-05
5,003 Snorkel DryBell: A Case Study in Deploying Weak Supervision at Industrial Scale 2019 SIGMOD 6.4732621e-05
5,014 Towards Benchmarking Feature Type Inference for AutoML Platforms 2021 SIGMOD 6.46838e-05
5,225 Explainable AI: Foundations, Applications, Opportunities for Data Management Research 2022 SIGMOD 6.3782447e-05
5,352 MDedup: Duplicate Detection with Matching Dependencies 2020 VLDB 6.3223357e-05
5,367 Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond 2021 SIGMOD 6.3159594e-05
5,471 Mining an "Anti-Knowledge Base" from Wikipedia Updates with Applications to Fact Checking and Beyond 2020 VLDB 6.2747651e-05
5,637 Adaptive Rule Discovery for Labeling Text Data 2021 SIGMOD 6.2036009e-05
5,852 Managing ML Pipelines: Feature Stores and the Coming Wave of Embedding Ecosystems 2021 VLDB 6.1270605e-05
5,928 VOCAL: Video Organization and Interactive Compositional AnaLytics 2022 CIDR 6.1004421e-05
5,964 Demonstration of Panda: A Weakly Supervised Entity Matching System 2021 VLDB 6.0876865e-05
5,965 Expand your Training Limits! Generating Training Data for ML-based Data Management 2021 SIGMOD 6.0876627e-05
6,017 Automatic Data Acquisition for Deep Learning 2021 VLDB 6.0676337e-05
6,029 Optimizing In-memory Database Engine for AI-powered On-line Decision Augmentation Using Persistent Memory 2021 VLDB 6.061589e-05
6,337 Finding Label and Model Errors in Perception Data With Learned Observation Assertions 2022 SIGMOD 5.9713243e-05
6,797 Cost-Effective Data Annotation using Game-Based Crowdsourcing 2019 VLDB 5.8305073e-05
7,207 Witan: Unsupervised Labelling Function Generation for Assisted Data Programming 2022 VLDB 5.7309223e-05
7,293 Data Integration and Machine Learning: A Natural Synergy 2018 VLDB 5.7085008e-05
7,434 Cross Modal Data Discovery over Structured and Unstructured Data Lakes 2023 VLDB 5.6805318e-05
7,476 Nautilus: An Optimized System for Deep Transfer Learning over Evolving Training Datasets 2022 SIGMOD 5.6710868e-05
7,741 CHEF: A Cheap and Fast Pipeline for Iteratively Cleaning Label Uncertainties 2021 VLDB 5.6133649e-05
7,764 iFlipper: Label Flipping for Individual Fairness 2023 SIGMOD 5.607573e-05
8,108 SHiFT: An Efficient, Flexible Search Engine for Transfer Learning 2023 VLDB 5.543315e-05
8,371 CrowdGame: A Game-Based Crowdsourcing System for Cost-Effective Data Labeling 2019 SIGMOD 5.4977619e-05
8,397 Nemo: Guiding and Contextualizing Weak Supervision for Interactive Data Programming 2022 VLDB 5.4958075e-05
8,441 UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads 2022 VLDB 5.4955352e-05
8,750 LANCET: Labeling Complex Data at Scale 2021 VLDB 5.4363235e-05
9,149 Hyper-Tune: Towards Efficient Hyper-parameter Tuning at Scale 2022 VLDB 5.3729482e-05
9,267 Improving Information Extraction from Visually Rich Documents using Visual Span Representations 2021 VLDB 5.3572577e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 5 of 5 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
117 HoloClean: Holistic Data Repairs with Probabilistic Inference 2017 VLDB 0.00032346273
417 Big Data Integration 2013 VLDB 0.00018903682
499 A Bayesian Approach to Discovering Truth from Conflicting Sources for Data Integration 2012 VLDB 0.00017523126
1,257 Fusing Data with Correlations 2014 SIGMOD 0.00011543418
3,665 SLiMFast: Guaranteed Results for Data Fusion and Source Reliability 2017 SIGMOD 7.278909e-05
Previous Page 1 / 1 Next

Semantically Similar Papers