Migrating a Privacy-Safe Information Extraction System to a Software 2.0 Design
Summary: Case study converting Gmail's privacy-safe, production rule-based IE to Software 2.0: use rule outputs as training labels to build ML extractors that improve precision/recall, shrink codebase, and enable cross-language extraction. Discusses challenges in training-data generation/management, model evaluation, and necessary Software‑1.0 infrastructure to safely deploy ML extractors. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Ying Sheng (Google)
- 2. Nguyen Vo (Google)
- 3. James B. Wendt (Google)
- 4. Sandeep Tata (Google)
- 5. Marc Najork (Google)
BibTeX Citation
@inproceedings{sheng_cidr20,
address = {Amsterdam, Netherlands},
series = {{CIDR} '20},
title = {{Migrating a Privacy-Safe Information Extraction System to a Software 2.0 Design}},
booktitle = {Proceedings of the {Conference} on {Innovative} {Data} {Systems} {Research}},
author = {Sheng, Ying and Vo, Nguyen and Wendt, James B. and Tata, Sandeep and Najork, Marc},
year = {2020}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 10 of 10 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 43 | The Case for Learned Index Structures | 2018 | SIGMOD | 0.00046060254 |
| 65 | Freebase: A Collaboratively Created Graph Database For Structuring Human Knowledge | 2008 | SIGMOD | 0.00038697603 |
| 112 | HoloClean: Holistic Data Repairs with Probabilistic Inference | 2017 | VLDB | 0.00032801121 |
| 588 | Extracting Structured Data from Web Pages | 2003 | SIGMOD | 0.00016092668 |
| 3,340 | Automatic Wrappers for Large Scale Web Extraction | 2011 | VLDB | 7.5040045e-05 |
| 3,600 | The Role of Massively Multi-Task and Weak Supervision in Software 2.0 | 2019 | CIDR | 7.2709969e-05 |
| 3,701 | Snorkel: Fast Training Set Generation for Information Extraction | 2017 | SIGMOD | 7.185321e-05 |
| 5,745 | DIADEM: Thousands of Websites to a Single Database | 2014 | VLDB | 6.1018855e-05 |
| 6,320 | CERES: Distantly Supervised Relation Extraction from the Semi-Structured Web | 2018 | VLDB | 5.9150411e-05 |
| 11,868 | Online Template Induction for Machine-Generated Emails | 2019 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 2,604 | Brainwash: A Data System for Feature Engineering | 2013 | CIDR |
| 2 | 11,824 | Leveraging Organizational Resources to Adapt Models to New Data Modalities | 2020 | VLDB |
| 3 | 6,315 | Data Collection and Quality Challenges for Deep Learning | 2020 | VLDB |
| 4 | 9,399 | Glean: Structured Extractions from Templatic Documents | 2021 | VLDB |
| 5 | 4,752 | Automatic Rule Refinement for Information Extraction | 2010 | VLDB |
| 6 | 244 | On the Design and Quantification of Privacy Preserving Data Mining Algorithms | 2001 | PODS |
| 7 | 9,563 | Retrofitting GDPR Compliance onto Legacy Databases | 2022 | VLDB |
| 8 | 12,604 | Mining Patterns and Rules for Software Specification Discovery | 2008 | VLDB |
| 9 | 12,809 | Privacy in Data Systems | 2003 | PODS |
| 10 | 3,600 | The Role of Massively Multi-Task and Weak Supervision in Software 2.0 | 2019 | CIDR |