The Role of Massively Multi-Task and Weak Supervision in Software 2.0
Summary: Vision: program Software 2.0 by labeling—declarative weak supervision aggregated via unsupervised label models to cheaply generate training data. Introduce massively multitask central models to amortize labeling across many tasks and validate via Snorkel deployments (ad fraud, diagnostics). (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Alexander Ratner (Stanford University)
- 2. Braden Hancock (Stanford University)
- 3. Christopher Ré (Stanford University)
BibTeX Citation
@inproceedings{ratner_cidr19,
address = {Amsterdam, Netherlands},
series = {{CIDR} '19},
title = {{The Role of Massively Multi-Task and Weak Supervision in Software 2.0}},
booktitle = {Proceedings of the {Conference} on {Innovative} {Data} {Systems} {Research}},
author = {Ratner, Alexander and Hancock, Braden and Ré, Christopher},
year = {2019}
}
Incoming Citations (Sorted by Pagerank)
Showing 5 of 5 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,702 | Panorama: A Data System for Unbounded Vocabulary Querying over Video | 2020 | VLDB | 8.2342712e-05 |
| 3,947 | Overton: A Data System for Monitoring and Improving Machine-Learned Products | 2020 | CIDR | 7.0040437e-05 |
| 4,382 | Slice Tuner: A Selective Data Acquisition Framework for Accurate and Fair Machine Learning Models | 2021 | SIGMOD | 6.7331832e-05 |
| 9,647 | Rock: Cleaning Data by Embedding ML in Logic Rules | 2024 | SIGMOD | 5.2430158e-05 |
| 11,740 | Migrating a Privacy-Safe Information Extraction System to a Software 2.0 Design | 2020 | CIDR | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 4 of 4 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 112 | HoloClean: Holistic Data Repairs with Probabilistic Inference | 2017 | VLDB | 0.00032801121 |
| 205 | Snorkel: Rapid Training Data Creation with Weak Supervision | 2018 | VLDB | 0.00025235185 |
| 1,094 | Snuba: Automating Weak Supervision to Label Training Data | 2019 | VLDB | 0.00012214617 |
| 5,058 | Snorkel DryBell: A Case Study in Deploying Weak Supervision at Industrial Scale | 2019 | SIGMOD | 6.3815523e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 8,523 | Nemo: Guiding and Contextualizing Weak Supervision for Interactive Data Programming | 2022 | VLDB |
| 2 | 3,701 | Snorkel: Fast Training Set Generation for Information Extraction | 2017 | SIGMOD |
| 3 | 11,824 | Leveraging Organizational Resources to Adapt Models to New Data Modalities | 2020 | VLDB |
| 4 | 6,816 | Inspector Gadget: A Data Programming-based Labeling System for Industrial Images | 2021 | VLDB |
| 5 | 7,609 | Ease.ml/ci and Ease.ml/meter in Action: Towards Data Management for Statistical Generalization | 2019 | VLDB |
| 6 | 6,315 | Data Collection and Quality Challenges for Deep Learning | 2020 | VLDB |
| 7 | 9,245 | Towards Observability for Production Machine Learning Pipelines | 2022 | VLDB |
| 8 | 11,740 | Migrating a Privacy-Safe Information Extraction System to a Software 2.0 Design | 2020 | CIDR |
| 9 | 205 | Snorkel: Rapid Training Data Creation with Weak Supervision | 2018 | VLDB |
| 10 | 5,058 | Snorkel DryBell: A Case Study in Deploying Weak Supervision at Industrial Scale | 2019 | SIGMOD |