Adaptive Rule Discovery for Labeling Text Data
Summary: Weakly supervised labeling of text with feedback; Darwin auto-generates and refines rules from an initial cue and scales to 1M+ sentences. CFG-based labeling functions; yields ~40% more positives than Snuba with the same effort. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Sainyam Galhotra (University of Massachusetts Amherst)
- 2. Behzad Golshan (Megagon Labs)
- 3. Wang-Chiew Tan (Meta)
BibTeX Citation
@inproceedings{galhotra_sigmod21,
title = {{Adaptive Rule Discovery for Labeling Text Data}},
author = {Galhotra, Sainyam and Golshan, Behzad and Tan, Wang-Chiew},
series = {{SIGMOD} '21},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3448016.3457334},
url = {https://dl.acm.org/doi/10.1145/3448016.3457334},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 5,054 | Explainable AI: Foundations, Applications, Opportunities for Data Management Research | 2022 | SIGMOD | 6.3843089e-05 |
| 7,331 | Witan: Unsupervised Labelling Function Generation for Assisted Data Programming | 2022 | VLDB | 5.6435171e-05 |
| 8,523 | Nemo: Guiding and Contextualizing Weak Supervision for Interactive Data Programming | 2022 | VLDB | 5.4119882e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 8 of 8 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 205 | Snorkel: Rapid Training Data Creation with Weak Supervision | 2018 | VLDB | 0.00025235185 |
| 427 | Big Data Integration | 2013 | VLDB | 0.00018661543 |
| 1,094 | Snuba: Automating Weak Supervision to Label Training Data | 2019 | VLDB | 0.00012214617 |
| 3,715 | SLiMFast: Guaranteed Results for Data Fusion and Source Reliability | 2017 | SIGMOD | 7.1763559e-05 |
| 6,908 | Cost-Effective Data Annotation using Game-Based Crowdsourcing | 2019 | VLDB | 5.7415834e-05 |
| 8,160 | ICARUS: Minimizing Human Effort in Iterative Data Completion | 2018 | VLDB | 5.4754476e-05 |
| 8,592 | Robust Entity Resolution using Random Graphs | 2018 | SIGMOD | 5.4058393e-05 |
| 11,961 | Scalable Semantic Querying of Text | 2018 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 5,051 | Towards Benchmarking Feature Type Inference for AutoML Platforms | 2021 | SIGMOD |
| 2 | 11,980 | Building Structured Databases of Factual Knowledge from Massive Text Corpora | 2017 | SIGMOD |
| 3 | 4,752 | Automatic Rule Refinement for Information Extraction | 2010 | VLDB |
| 4 | 9,427 | Discovering Top-k Rules using Subjective and Objective Criteria | 2023 | SIGMOD |
| 5 | 205 | Snorkel: Rapid Training Data Creation with Weak Supervision | 2018 | VLDB |
| 6 | 3,701 | Snorkel: Fast Training Set Generation for Information Extraction | 2017 | SIGMOD |
| 7 | 11,737 | DBTagger: Multi-Task Learning for Keyword Mapping in NLIDBs Using Bi-Directional Recurrent Neural Networks | 2021 | VLDB |
| 8 | 9,560 | Ground Truth Inference for Weakly Supervised Entity Matching | 2023 | SIGMOD |
| 9 | 6,908 | Cost-Effective Data Annotation using Game-Based Crowdsourcing | 2019 | VLDB |
| 10 | 1,094 | Snuba: Automating Weak Supervision to Label Training Data | 2019 | VLDB |