Snorkel DryBell: A Case Study in Deploying Weak Supervision at Industrial Scale
Summary: Industrial-scale weak supervision via Snorkel DryBell uses organizational knowledge as labeling signals. Template-based ingestion, cross-feature serving, and sampling-free execution scale to millions of points, yielding near hand-labeled accuracy with ~52% uplift. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Stephen H. Bach (Brown University)
- 2. Daniel Rodriguez (Google)
- 3. Yintao Liu (Google)
- 4. Chong Luo (Google)
- 5. Haidong Shao (Google)
- 6. Cassandra Xia (Google)
- 7. Souvik Sen (Google)
- 8. Alex Ratner (Stanford University)
- 9. Braden Hancock (Stanford University)
- 10. Houman Alborzi (Google)
- 11. Rahul Kuchhal (Google)
- 12. Chris Ré (Stanford University)
- 13. Rob Malkin (Google)
BibTeX Citation
@inproceedings{bach_sigmod19,
title = {{Snorkel DryBell: A Case Study in Deploying Weak Supervision at Industrial Scale}},
author = {Bach, Stephen H. and Rodriguez, Daniel and Liu, Yintao and Luo, Chong and Shao, Haidong and Xia, Cassandra and Sen, Souvik and Ratner, Alex and Hancock, Braden and Alborzi, Houman and Kuchhal, Rahul and Ré, Chris and Malkin, Rob},
series = {{SIGMOD} '19},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3299869.3314036},
url = {https://dl.acm.org/doi/10.1145/3299869.3314036},
year = {2019}
}
Incoming Citations (Sorted by Pagerank)
Showing 9 of 9 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,273 | SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model Debugging | 2021 | SIGMOD | 8.8230899e-05 |
| 3,600 | The Role of Massively Multi-Task and Weak Supervision in Software 2.0 | 2019 | CIDR | 7.2709969e-05 |
| 3,947 | Overton: A Data System for Monitoring and Improving Machine-Learned Products | 2020 | CIDR | 7.0040437e-05 |
| 6,816 | Inspector Gadget: A Data Programming-based Labeling System for Industrial Images | 2021 | VLDB | 5.7649828e-05 |
| 7,868 | CHEF: A Cheap and Fast Pipeline for Iteratively Cleaning Label Uncertainties | 2021 | VLDB | 5.5277527e-05 |
| 8,523 | Nemo: Guiding and Contextualizing Weak Supervision for Interactive Data Programming | 2022 | VLDB | 5.4119882e-05 |
| 10,805 | WeShap: Weak Supervision Source Evaluation with Shapley Values | 2025 | VLDB | 5.093636e-05 |
| 11,430 | VersaMatch: Ontology Matching with Weak Supervision | 2023 | VLDB | 5.093636e-05 |
| 11,824 | Leveraging Organizational Resources to Adapt Models to New Data Modalities | 2020 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 112 | HoloClean: Holistic Data Repairs with Probabilistic Inference | 2017 | VLDB | 0.00032801121 |
| 205 | Snorkel: Rapid Training Data Creation with Weak Supervision | 2018 | VLDB | 0.00025235185 |
| 427 | Big Data Integration | 2013 | VLDB | 0.00018661543 |
| 504 | A Bayesian Approach to Discovering Truth from Conflicting Sources for Data Integration | 2012 | VLDB | 0.00017300628 |
| 3,715 | SLiMFast: Guaranteed Results for Data Fusion and Source Reliability | 2017 | SIGMOD | 7.1763559e-05 |
| 4,410 | Extracting Databases from Dark Data with DeepDive | 2016 | SIGMOD | 6.717496e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 6,816 | Inspector Gadget: A Data Programming-based Labeling System for Industrial Images | 2021 | VLDB |
| 2 | 9,560 | Ground Truth Inference for Weakly Supervised Entity Matching | 2023 | SIGMOD |
| 3 | 5,722 | Adaptive Rule Discovery for Labeling Text Data | 2021 | SIGMOD |
| 4 | 2,657 | Production Machine Learning Pipelines: Empirical Analysis and Optimization Opportunities | 2021 | SIGMOD |
| 5 | 9,245 | Towards Observability for Production Machine Learning Pipelines | 2022 | VLDB |
| 6 | 11,824 | Leveraging Organizational Resources to Adapt Models to New Data Modalities | 2020 | VLDB |
| 7 | 1,094 | Snuba: Automating Weak Supervision to Label Training Data | 2019 | VLDB |
| 8 | 3,701 | Snorkel: Fast Training Set Generation for Information Extraction | 2017 | SIGMOD |
| 9 | 3,600 | The Role of Massively Multi-Task and Weak Supervision in Software 2.0 | 2019 | CIDR |
| 10 | 205 | Snorkel: Rapid Training Data Creation with Weak Supervision | 2018 | VLDB |