Database Paper Browser

Back to papers

Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond

Summary: Rotom: meta-learned data augmentation for entity matching, data cleaning, and text classification. Introduces InvDA (seq2seq) and a learned policy to combine DA operators, reducing hyperparameter search and boosting low-resource results, beating SOTA. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
6150
Venue
SIGMOD
Year
2021
Pagerank
5.2458154e-05
Overall Rank
5,974 | 58.49%
DOI
10.1145/3448016.3457258

Incoming Non-self Citations Over Time

Authors

Incoming Citations (Sorted by Pagerank)

Showing 11 of 11 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 30 of 30 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
192 HoloClean: Holistic Data Repairs with Probabilistic Inference 2017 VLDB 0.00035692958
219 Deep Entity Matching with Pre-Trained Language Models 2021 VLDB 0.00033354456
252 Snorkel: Rapid Training Data Creation with Weak Supervision 2018 VLDB 0.00030532082
268 A Cost-Based Model and Effective Heuristic for Repairing Constraints by Value Modification 2005 SIGMOD 0.00029739054
293 Deep Learning for Entity Matching: A Design Space Exploration 2018 SIGMOD 0.00028661817
507 On Active Learning of Record Matching Packages 2010 SIGMOD 0.00021474096
609 Goods: Organizing Google's Datasets 2016 SIGMOD 0.00019223217
705 Magellan: Toward Building Entity Matching Management Systems 2016 VLDB 0.00017779048
788 ActiveClean: Interactive Data Cleaning For Statistical Modeling 2016 VLDB 0.00016618698
799 Entity Resolution: Theory, Practice & Open Challenges 2012 VLDB 0.00016479804
901 To Join or Not to Join? Thinking Twice about Joins before Feature Selection 2016 SIGMOD 0.00015462938
1,218 Snuba: Automating Weak Supervision to Label Training Data 2019 VLDB 0.00013221309
1,340 HoloDetect: Few-Shot Learning for Error Detection 2019 SIGMOD 0.00012492795
1,462 ARDA: Automatic Relational Data Augmentation for Machine Learning 2020 VLDB 0.00011866333
1,534 Data Management in Machine Learning: Challenges, Techniques, and Systems 2017 SIGMOD 0.00011462072
1,544 KATARA: A Data Cleaning System Powered by Knowledge Bases and Crowdsourcing 2015 SIGMOD 0.00011438274
1,895 Baran: Effective Error Correction via a Unified Context Representation and Transfer Learning 2020 VLDB 0.00010174634
2,176 Falcon: Scaling Up Hands-Off Crowdsourced Entity Matching to Build Cloud Services 2017 SIGMOD 9.3729351e-05
2,758 A Comprehensive Benchmark Framework for Active Learning Methods in Entity Matching 2020 SIGMOD 8.1668285e-05
2,895 Sato: Contextual Semantic Type Detection in Tables 2020 VLDB 7.9539265e-05
2,968 Raha: A Configuration-Free Error Detection System 2019 SIGMOD 7.7964476e-05
3,071 CrowdFill: Collecting Structured Data from the Crowd 2014 SIGMOD 7.6120337e-05
3,767 Cleaning Crowdsourced Labels Using Oracles for Statistical Classification 2019 VLDB 6.7748725e-05
3,900 SLiMFast: Guaranteed Results for Data Fusion and Source Reliability 2017 SIGMOD 6.649432e-05
4,123 Are Key-Foreign Key Joins Safe to Avoid when Learning High-Capacity Classifiers? 2018 VLDB 6.4290005e-05
4,453 CLAMShell: Speeding up Crowds for Low-latency Data Labeling 2016 VLDB 6.1690121e-05
4,911 Temporal Rules Discovery for Web Data Cleaning 2016 VLDB 5.8349225e-05
6,047 MDedup: Duplicate Detection with Matching Dependencies 2020 VLDB 5.2355891e-05
7,013 Qualitative Data Cleaning 2016 VLDB 4.8576683e-05
7,116 Crowdsourced Data Management: Overview and Challenges 2017 SIGMOD 4.8219732e-05
Previous Page 1 / 1 Next

Semantically Similar Papers