DBScholar

Back to papers

Snorkel: Rapid Training Data Creation with Weak Supervision

Summary: Snorkel enables rapid ML training from weak supervision via labeling functions with unknown accuracies. End-to-end data programming denoises labels without ground truth, with a tradeoff optimizer, showing speedups and accuracy gains over hand labeling. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h88c20434faae0223
Venue
VLDB
Year
2018
Pagerank
0.00025181304
Overall Rank
205 | 98.63%
DOI
10.14778/3157794.3157797

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{ratner_vldb18,
        title = {{Snorkel: Rapid Training Data Creation with Weak Supervision}},
        author = {Ratner, Alexander and Bach, Stephen H. and Ehrenberg, Henry and Fries, Jason and Wu, Sen and Ré, Christopher},
        journal = {PVLDB},
        series = {{VLDB} '18},
        volume = {11},
        number = {3},
        pages = {269--282},
        doi = {10.14778/3157794.3157797},
        url = {https://doi.org/10.14778/3157794.3157797},
        year = {2018}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 72 citing papers.

Rank Citing Paper Year Venue Pagerank
158 Deep Learning for Entity Matching: A Design Space Exploration 2018 SIGMOD 0.00028046388
501 Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes 2024 VLDB 0.00017267905
883 HoloDetect: Few-Shot Learning for Error Detection 2019 SIGMOD 0.00013268059
1,120 Snuba: Automating Weak Supervision to Label Training Data 2019 VLDB 0.00011946047
1,568 HELIX: Holistic Optimization for Accelerating Iterative Machine Learning 2019 VLDB 0.0001021302
2,011 DBPal: A Fully Pluggable NL2SQL Training Pipeline 2020 SIGMOD 9.1896169e-05
2,048 Automatically Generating Data Exploration Sessions Using Deep Reinforcement Learning 2020 SIGMOD 9.1229916e-05
2,326 SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model Debugging 2021 SIGMOD 8.6309237e-05
3,061 Fonduer: Knowledge Base Construction from Richly Formatted Data 2018 SIGMOD 7.6927483e-05
3,331 VolcanoML: Speeding up End-to-End AutoML via Scalable Search Space Decomposition 2021 VLDB 7.4166095e-05
3,521 Ember: No-Code Context Enrichment via Similarity-Based Keyless Joins 2022 VLDB 7.2351481e-05
3,627 The Role of Massively Multi-Task and Weak Supervision in Software 2.0 2019 CIDR 7.1514557e-05
3,955 Smile: A System to Support Machine Learning on EEG Data at Scale 2019 VLDB 6.9025583e-05
3,965 MB2: Decomposed Behavior Modeling for Self-Driving Database Management Systems 2021 SIGMOD 6.8918628e-05
3,997 Overton: A Data System for Monitoring and Improving Machine-Learned Products 2020 CIDR 6.8655222e-05
4,103 AutoOD: Automatic Outlier Detection 2023 SIGMOD 6.8066074e-05
4,171 Data Integration and Machine Learning: A Natural Synergy 2018 SIGMOD 6.7608137e-05
4,543 GOGGLES: Automatic Image Labeling with Affinity Coding 2020 SIGMOD 6.5470116e-05
4,675 spade: Synthesizing Data Quality Assertions for Large Language Model Pipelines 2024 VLDB 6.474947e-05
4,719 OmniFair: A Declarative System for Model-Agnostic Group Fairness in Machine Learning 2021 SIGMOD 6.4548076e-05
4,867 Selective Data Acquisition in the Wild for Model Charging 2022 VLDB 6.3749346e-05
5,043 ODIN: Automated Drift Detection and Recovery in Video Analytics 2020 VLDB 6.3002257e-05
5,076 Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond 2021 SIGMOD 6.2860582e-05
5,100 Towards Benchmarking Feature Type Inference for AutoML Platforms 2021 SIGMOD 6.27349e-05
5,178 Explainable AI: Foundations, Applications, Opportunities for Data Management Research 2022 SIGMOD 6.2411759e-05
5,181 Snorkel DryBell: A Case Study in Deploying Weak Supervision at Industrial Scale 2019 SIGMOD 6.2406557e-05
5,506 Managing ML Pipelines: Feature Stores and the Coming Wave of Embedding Ecosystems 2021 VLDB 6.1016716e-05
5,522 MDedup: Duplicate Detection with Matching Dependencies 2020 VLDB 6.0956156e-05
5,679 Mining an "Anti-Knowledge Base" from Wikipedia Updates with Applications to Fact Checking and Beyond 2020 VLDB 6.0404239e-05
5,850 Adaptive Rule Discovery for Labeling Text Data 2021 SIGMOD 5.9719174e-05
5,970 VOCAL: Video Organization and Interactive Compositional AnaLytics 2022 CIDR 5.9316412e-05
6,139 iFlipper: Label Flipping for Individual Fairness 2023 SIGMOD 5.8734506e-05
6,141 Expand your Training Limits! Generating Training Data for ML-based Data Management 2021 SIGMOD 5.8733296e-05
6,156 Automatic Data Acquisition for Deep Learning 2021 VLDB 5.8680826e-05
6,165 Demonstration of Panda: A Weakly Supervised Entity Matching System 2021 VLDB 5.8635301e-05
6,236 Optimizing In-memory Database Engine for AI-powered On-line Decision Augmentation Using Persistent Memory 2021 VLDB 5.8409336e-05
6,552 Finding Label and Model Errors in Perception Data With Learned Observation Assertions 2022 SIGMOD 5.750057e-05
6,662 UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads 2022 VLDB 5.7171651e-05
7,052 Cost-Effective Data Annotation using Game-Based Crowdsourcing 2019 VLDB 5.6127577e-05
7,090 LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning 2026 VLDB 5.601767e-05
7,197 Cross Modal Data Discovery over Structured and Unstructured Data Lakes 2023 VLDB 5.5890662e-05
7,415 Data Integration and Machine Learning: A Natural Synergy 2018 VLDB 5.5332662e-05
7,473 Witan: Unsupervised Labelling Function Generation for Assisted Data Programming 2022 VLDB 5.5168918e-05
7,809 Nautilus: An Optimized System for Deep Transfer Learning over Evolving Training Datasets 2022 SIGMOD 5.4490255e-05
7,876 CHEF: A Cheap and Fast Pipeline for Iteratively Cleaning Label Uncertainties 2021 VLDB 5.4348457e-05
8,176 Improving Information Extraction from Visually Rich Documents using Visual Span Representations 2021 VLDB 5.3835163e-05
8,205 Falcon: Fair Active Learning using Multi-armed Bandits 2024 VLDB 5.3781556e-05
8,277 Hyper-Tune: Towards Efficient Hyper-parameter Tuning at Scale 2022 VLDB 5.3639084e-05
8,411 SHiFT: An Efficient, Flexible Search Engine for Transfer Learning 2023 VLDB 5.336291e-05
8,664 CrowdGame: A Game-Based Crowdsourcing System for Cost-Effective Data Labeling 2019 SIGMOD 5.2926663e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 5 of 5 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
104 HoloClean: Holistic Data Repairs with Probabilistic Inference 2017 VLDB 0.00033690989
427 Big Data Integration 2013 VLDB 0.00018465558
504 A Bayesian Approach to Discovering Truth from Conflicting Sources for Data Integration 2012 VLDB 0.00017164466
1,284 Fusing Data with Correlations 2014 SIGMOD 0.00011202115
3,791 SLiMFast: Guaranteed Results for Data Fusion and Source Reliability 2017 SIGMOD 7.0189865e-05
Previous Page 1 / 1 Next

Semantically Similar Papers