Inspector Gadget: A Data Programming-based Labeling System for Industrial Images
Summary: Directly applying data programming to images without conversion for industrial labeling. Inspector Gadget fuses crowdsourcing, augmentation, and labeling functions to generate scalable weak labels for image classification, beating Snuba, GOGGLES, and self-learning CNN baselines without pretraining. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Geon Heo (Korea Advanced Institute of Science and Technology)
- 2. Yuji Roh (Korea Advanced Institute of Science and Technology)
- 3. Seonghyeon Hwang (Korea Advanced Institute of Science and Technology)
- 4. Dayun Lee (Korea Advanced Institute of Science and Technology)
- 5. Steven Euijong Whang (Korea Advanced Institute of Science and Technology)
BibTeX Citation
@article{heo_vldb21,
title = {{Inspector Gadget: A Data Programming-based Labeling System for Industrial Images}},
author = {Heo, Geon and Roh, Yuji and Hwang, Seonghyeon and Lee, Dayun and Whang, Steven Euijong},
journal = {PVLDB},
series = {{VLDB} '21},
volume = {14},
number = {1},
pages = {28--40},
doi = {10.14778/3421424.3421430},
url = {https://doi.org/10.14778/3421424.3421430},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 8,959 | TSM-Bench: Benchmarking Time Series Database Systems for Monitoring Applications | 2023 | VLDB | 5.344413e-05 |
| 10,746 | A Cost-Effective LLM-based Approach to Identify Wildlife Trafficking in Online Marketplaces | 2025 | SIGMOD | 5.0935618e-05 |
| 11,315 | SEER: An End-to-End Toolkit for Benchmarking Time Series Database Systems in Monitoring Applications | 2024 | VLDB | 5.0935618e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 251 | Crowdsourced Databases: Query Processing with People | 2011 | CIDR | 0.00023260775 |
| 439 | Corleone: Hands-Off Crowdsourcing for Entity Matching | 2014 | SIGMOD | 0.00018464644 |
| 1,094 | Snuba: Automating Weak Supervision to Label Training Data | 2019 | VLDB | 0.00012214439 |
| 3,701 | Snorkel: Fast Training Set Generation for Information Extraction | 2017 | SIGMOD | 7.1852164e-05 |
| 3,971 | CLAMShell: Speeding up Crowds for Low-latency Data Labeling | 2016 | VLDB | 6.9834246e-05 |
| 5,058 | Snorkel DryBell: A Case Study in Deploying Weak Supervision at Industrial Scale | 2019 | SIGMOD | 6.3814594e-05 |
| 8,495 | CrowdGame: A Game-Based Crowdsourcing System for Cost-Effective Data Labeling | 2019 | SIGMOD | 5.413834e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 6,251 | Lightweight Inspection of Data Preprocessing in Native Machine Learning Pipelines | 2021 | CIDR |
| 2 | 6,432 | Finding Label and Model Errors in Perception Data With Learned Observation Assertions | 2022 | SIGMOD |
| 3 | 8,523 | Nemo: Guiding and Contextualizing Weak Supervision for Interactive Data Programming | 2022 | VLDB |
| 4 | 8,876 | LANCET: Labeling Complex Data at Scale | 2021 | VLDB |
| 5 | 1,094 | Snuba: Automating Weak Supervision to Label Training Data | 2019 | VLDB |
| 6 | 6,075 | Automatic Data Acquisition for Deep Learning | 2021 | VLDB |
| 7 | 11,822 | Exploiting Domain Knowledge to address Multi-Class Imbalance and a Heterogeneous Feature Space in Classification Tasks for Manufacturing Data | 2020 | VLDB |
| 8 | 4,451 | GOGGLES: Automatic Image Labeling with Affinity Coding | 2020 | SIGMOD |
| 9 | 3,600 | The Role of Massively Multi-Task and Weak Supervision in Software 2.0 | 2019 | CIDR |
| 10 | 205 | Snorkel: Rapid Training Data Creation with Weak Supervision | 2018 | VLDB |