DBScholar

Back to papers

Snorkel: Rapid Training Data Creation with Weak Supervision

Summary: Snorkel enables rapid ML training from weak supervision via labeling functions with unknown accuracies. End-to-end data programming denoises labels without ground truth, with a tradeoff optimizer, showing speedups and accuracy gains over hand labeling. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h88c20434faae0223
Venue
VLDB
Year
2018
Pagerank
0.00025181304
Overall Rank
205 | 98.63%
DOI
10.14778/3157794.3157797

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{ratner_vldb18,
        title = {{Snorkel: Rapid Training Data Creation with Weak Supervision}},
        author = {Ratner, Alexander and Bach, Stephen H. and Ehrenberg, Henry and Fries, Jason and Wu, Sen and Ré, Christopher},
        journal = {PVLDB},
        series = {{VLDB} '18},
        volume = {11},
        number = {3},
        pages = {269--282},
        doi = {10.14778/3157794.3157797},
        url = {https://doi.org/10.14778/3157794.3157797},
        year = {2018}
}

Incoming Citations (Sorted by Pagerank)

Showing 22 of 72 citing papers.

Rank Citing Paper Year Venue Pagerank
8,690 Nemo: Guiding and Contextualizing Weak Supervision for Interactive Data Programming 2022 VLDB 5.2905577e-05
9,036 LANCET: Labeling Complex Data at Scale 2021 VLDB 5.2332952e-05
9,581 Glean: Structured Extractions from Templatic Documents 2021 VLDB 5.1571823e-05
9,708 Ground Truth Inference for Weakly Supervised Entity Matching 2023 SIGMOD 5.1374628e-05
9,766 Bootleg: Chasing the Tail with Self-Supervised Named Entity Disambiguation 2021 CIDR 5.1344318e-05
10,043 The Image Calculator: 10x Faster Image-AI Inference by Replacing JPEG with Self-designing Storage Format 2024 SIGMOD 5.0921006e-05
10,112 Data Augmentation for ML-driven Data Preparation and Integration 2021 VLDB 5.0789354e-05
10,171 Towards Autonomous, Hands-Free Data Exploration 2020 CIDR 5.0682654e-05
10,845 Featurized-Decomposition Join: Low-Cost Semantic Joins with Guarantees 2026 VLDB 4.9793485e-05
10,935 OmniTable: A Unified Wide-Table System for Petabyte-Scale LLM Data Curation and Exploration 2026 VLDB 4.9793485e-05
10,945 Morphing-based Compression for Data-centric ML Pipelines 2026 VLDB 4.9793485e-05
11,218 WeShap: Weak Supervision Source Evaluation with Shapley Values 2025 VLDB 4.9793485e-05
11,235 A Systematic Study on Early Stopping Metrics in HPO and the Implications of Uncertainty 2025 VLDB 4.9793485e-05
11,661 Generalizable Data Cleaning of Tabular Data in Latent Space 2024 VLDB 4.9793485e-05
11,720 Steered Training Data Generation for Learned Semantic Type Detection 2023 SIGMOD 4.9793485e-05
11,744 VersaMatch: Ontology Matching with Weak Supervision 2023 VLDB 4.9793485e-05
11,915 Machine Programming: Turning Data into Programmer Productivity 2022 VLDB 4.9793485e-05
11,936 Ease.ML: A Lifecycle Management System for MLDev and MLOps 2021 CIDR 4.9793485e-05
12,025 An Extensible and Reusable Pipeline for Automated Utterance Paraphrases 2021 VLDB 4.9793485e-05
12,038 Quality of Sentiment Analysis Tools: The Reasons of Inconsistency 2021 VLDB 4.9793485e-05
12,060 Factorized Graph Representations for Semi-Supervised Learning from Sparse Data 2020 SIGMOD 4.9793485e-05
12,125 Leveraging Organizational Resources to Adapt Models to New Data Modalities 2020 VLDB 4.9793485e-05
Previous Page 2 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 5 of 5 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
104 HoloClean: Holistic Data Repairs with Probabilistic Inference 2017 VLDB 0.00033690989
427 Big Data Integration 2013 VLDB 0.00018465558
504 A Bayesian Approach to Discovering Truth from Conflicting Sources for Data Integration 2012 VLDB 0.00017164466
1,284 Fusing Data with Correlations 2014 SIGMOD 0.00011202115
3,791 SLiMFast: Guaranteed Results for Data Fusion and Source Reliability 2017 SIGMOD 7.0189865e-05
Previous Page 1 / 1 Next

Semantically Similar Papers