Snorkel: Rapid Training Data Creation with Weak Supervision
Summary: Snorkel enables rapid ML training from weak supervision via labeling functions with unknown accuracies. End-to-end data programming denoises labels without ground truth, with a tradeoff optimizer, showing speedups and accuracy gains over hand labeling. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Alexander Ratner (Stanford University)
- 2. Stephen H. Bach (Stanford University)
- 3. Henry Ehrenberg (Stanford University)
- 4. Jason Fries (Stanford University)
- 5. Sen Wu (Stanford University)
- 6. Christopher Ré (Stanford University)
BibTeX Citation
@article{ratner_vldb18,
title = {{Snorkel: Rapid Training Data Creation with Weak Supervision}},
author = {Ratner, Alexander and Bach, Stephen H. and Ehrenberg, Henry and Fries, Jason and Wu, Sen and Ré, Christopher},
journal = {PVLDB},
series = {{VLDB} '18},
volume = {11},
number = {3},
pages = {269--282},
doi = {10.14778/3157794.3157797},
url = {https://doi.org/10.14778/3157794.3157797},
year = {2018}
}
Incoming Citations (Sorted by Pagerank)
Showing 50 of 70 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 112 | HoloClean: Holistic Data Repairs with Probabilistic Inference | 2017 | VLDB | 0.00032801121 |
| 427 | Big Data Integration | 2013 | VLDB | 0.00018661543 |
| 504 | A Bayesian Approach to Discovering Truth from Conflicting Sources for Data Integration | 2012 | VLDB | 0.00017300628 |
| 1,273 | Fusing Data with Correlations | 2014 | SIGMOD | 0.00011384191 |
| 3,715 | SLiMFast: Guaranteed Results for Data Fusion and Source Reliability | 2017 | SIGMOD | 7.1763559e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 8,594 | Exploratory Training: When Annotators Learn About Data | 2023 | SIGMOD |
| 2 | 9,560 | Ground Truth Inference for Weakly Supervised Entity Matching | 2023 | SIGMOD |
| 3 | 6,074 | Automatic Data Acquisition for Deep Learning | 2021 | VLDB |
| 4 | 5,722 | Adaptive Rule Discovery for Labeling Text Data | 2021 | SIGMOD |
| 5 | 8,523 | Nemo: Guiding and Contextualizing Weak Supervision for Interactive Data Programming | 2022 | VLDB |
| 6 | 6,816 | Inspector Gadget: A Data Programming-based Labeling System for Industrial Images | 2021 | VLDB |
| 7 | 3,600 | The Role of Massively Multi-Task and Weak Supervision in Software 2.0 | 2019 | CIDR |
| 8 | 5,058 | Snorkel DryBell: A Case Study in Deploying Weak Supervision at Industrial Scale | 2019 | SIGMOD |
| 9 | 1,094 | Snuba: Automating Weak Supervision to Label Training Data | 2019 | VLDB |
| 10 | 3,701 | Snorkel: Fast Training Set Generation for Information Extraction | 2017 | SIGMOD |