DBScholar

Back to papers

Snorkel: Rapid Training Data Creation with Weak Supervision

Summary: Snorkel enables rapid ML training from weak supervision via labeling functions with unknown accuracies. End-to-end data programming denoises labels without ground truth, with a tradeoff optimizer, showing speedups and accuracy gains over hand labeling. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h88c20434faae0223
Venue
VLDB
Year
2018
Pagerank
0.00025171314
Overall Rank
205 | 98.63%
DOI
10.14778/3157794.3157797

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{ratner_vldb18,
        title = {{Snorkel: Rapid Training Data Creation with Weak Supervision}},
        author = {Ratner, Alexander and Bach, Stephen H. and Ehrenberg, Henry and Fries, Jason and Wu, Sen and Ré, Christopher},
        journal = {PVLDB},
        series = {{VLDB} '18},
        volume = {11},
        number = {3},
        pages = {269--282},
        doi = {10.14778/3157794.3157797},
        url = {https://doi.org/10.14778/3157794.3157797},
        year = {2018}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 72 citing papers.

Rank Citing Paper Year Venue Pagerank
158 Deep Learning for Entity Matching: A Design Space Exploration 2018 SIGMOD 0.00028038831
496 Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes 2024 VLDB 0.00017318538
884 HoloDetect: Few-Shot Learning for Error Detection 2019 SIGMOD 0.00013263269
1,120 Snuba: Automating Weak Supervision to Label Training Data 2019 VLDB 0.000119406
1,568 HELIX: Holistic Optimization for Accelerating Iterative Machine Learning 2019 VLDB 0.00010208225
2,014 DBPal: A Fully Pluggable NL2SQL Training Pipeline 2020 SIGMOD 9.185301e-05
2,050 Automatically Generating Data Exploration Sessions Using Deep Reinforcement Learning 2020 SIGMOD 9.1186729e-05
2,329 SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model Debugging 2021 SIGMOD 8.6268411e-05
3,063 Fonduer: Knowledge Base Construction from Richly Formatted Data 2018 SIGMOD 7.689108e-05
3,332 VolcanoML: Speeding up End-to-End AutoML via Scalable Search Space Decomposition 2021 VLDB 7.4130985e-05
3,521 Ember: No-Code Context Enrichment via Similarity-Based Keyless Joins 2022 VLDB 7.2320843e-05
3,627 The Role of Massively Multi-Task and Weak Supervision in Software 2.0 2019 CIDR 7.1492314e-05
3,956 Smile: A System to Support Machine Learning on EEG Data at Scale 2019 VLDB 6.8992907e-05
3,964 MB2: Decomposed Behavior Modeling for Self-Driving Database Management Systems 2021 SIGMOD 6.889374e-05
3,998 Overton: A Data System for Monitoring and Improving Machine-Learned Products 2020 CIDR 6.862274e-05
4,105 AutoOD: Automatic Outlier Detection 2023 SIGMOD 6.8033852e-05
4,171 Data Integration and Machine Learning: A Natural Synergy 2018 SIGMOD 6.7576159e-05
4,544 GOGGLES: Automatic Image Labeling with Affinity Coding 2020 SIGMOD 6.5439123e-05
4,678 spade: Synthesizing Data Quality Assertions for Large Language Model Pipelines 2024 VLDB 6.4720352e-05
4,721 OmniFair: A Declarative System for Model-Agnostic Group Fairness in Machine Learning 2021 SIGMOD 6.451752e-05
4,869 Selective Data Acquisition in the Wild for Model Charging 2022 VLDB 6.3719187e-05
5,046 ODIN: Automated Drift Detection and Recovery in Video Analytics 2020 VLDB 6.2972614e-05
5,078 Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond 2021 SIGMOD 6.2832055e-05
5,102 Towards Benchmarking Feature Type Inference for AutoML Platforms 2021 SIGMOD 6.2705617e-05
5,179 Explainable AI: Foundations, Applications, Opportunities for Data Management Research 2022 SIGMOD 6.2382214e-05
5,182 Snorkel DryBell: A Case Study in Deploying Weak Supervision at Industrial Scale 2019 SIGMOD 6.2377015e-05
5,509 Managing ML Pipelines: Feature Stores and the Coming Wave of Embedding Ecosystems 2021 VLDB 6.0987831e-05
5,525 MDedup: Duplicate Detection with Matching Dependencies 2020 VLDB 6.0927371e-05
5,680 Mining an "Anti-Knowledge Base" from Wikipedia Updates with Applications to Fact Checking and Beyond 2020 VLDB 6.0375644e-05
5,853 Adaptive Rule Discovery for Labeling Text Data 2021 SIGMOD 5.9690904e-05
5,948 VOCAL: Video Organization and Interactive Compositional AnaLytics 2022 CIDR 5.934521e-05
6,141 Expand your Training Limits! Generating Training Data for ML-based Data Management 2021 SIGMOD 5.8708409e-05
6,142 iFlipper: Label Flipping for Individual Fairness 2023 SIGMOD 5.8706702e-05
6,159 Automatic Data Acquisition for Deep Learning 2021 VLDB 5.8653048e-05
6,167 Demonstration of Panda: A Weakly Supervised Entity Matching System 2021 VLDB 5.8607544e-05
6,239 Optimizing In-memory Database Engine for AI-powered On-line Decision Augmentation Using Persistent Memory 2021 VLDB 5.8381685e-05
6,554 Finding Label and Model Errors in Perception Data With Learned Observation Assertions 2022 SIGMOD 5.7473435e-05
6,666 UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads 2022 VLDB 5.7144587e-05
7,054 Cost-Effective Data Annotation using Game-Based Crowdsourcing 2019 VLDB 5.6101007e-05
7,092 LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning 2026 VLDB 5.5991152e-05
7,199 Cross Modal Data Discovery over Structured and Unstructured Data Lakes 2023 VLDB 5.5864204e-05
7,418 Data Integration and Machine Learning: A Natural Synergy 2018 VLDB 5.5306468e-05
7,479 Witan: Unsupervised Labelling Function Generation for Assisted Data Programming 2022 VLDB 5.5142801e-05
7,816 Nautilus: An Optimized System for Deep Transfer Learning over Evolving Training Datasets 2022 SIGMOD 5.446446e-05
7,881 CHEF: A Cheap and Fast Pipeline for Iteratively Cleaning Label Uncertainties 2021 VLDB 5.4322729e-05
8,183 Improving Information Extraction from Visually Rich Documents using Visual Span Representations 2021 VLDB 5.3809679e-05
8,213 Falcon: Fair Active Learning using Multi-armed Bandits 2024 VLDB 5.375648e-05
8,283 Hyper-Tune: Towards Efficient Hyper-parameter Tuning at Scale 2022 VLDB 5.3613692e-05
8,418 SHiFT: An Efficient, Flexible Search Engine for Transfer Learning 2023 VLDB 5.3337649e-05
8,672 CrowdGame: A Game-Based Crowdsourcing System for Cost-Effective Data Labeling 2019 SIGMOD 5.2901608e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 5 of 5 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
104 HoloClean: Holistic Data Repairs with Probabilistic Inference 2017 VLDB 0.00033676943
428 Big Data Integration 2013 VLDB 0.00018457189
504 A Bayesian Approach to Discovering Truth from Conflicting Sources for Data Integration 2012 VLDB 0.00017156765
1,284 Fusing Data with Correlations 2014 SIGMOD 0.00011197034
3,793 SLiMFast: Guaranteed Results for Data Fusion and Source Reliability 2017 SIGMOD 7.0156889e-05
Previous Page 1 / 1 Next

Semantically Similar Papers