| 192 |
HoloClean: Holistic Data Repairs with Probabilistic Inference |
2017 |
VLDB |
0.00035692958 |
| 219 |
Deep Entity Matching with Pre-Trained Language Models |
2021 |
VLDB |
0.00033354456 |
| 252 |
Snorkel: Rapid Training Data Creation with Weak Supervision |
2018 |
VLDB |
0.00030532082 |
| 268 |
A Cost-Based Model and Effective Heuristic for Repairing Constraints by Value Modification |
2005 |
SIGMOD |
0.00029739054 |
| 293 |
Deep Learning for Entity Matching: A Design Space Exploration |
2018 |
SIGMOD |
0.00028661817 |
| 507 |
On Active Learning of Record Matching Packages |
2010 |
SIGMOD |
0.00021474096 |
| 609 |
Goods: Organizing Google's Datasets |
2016 |
SIGMOD |
0.00019223217 |
| 705 |
Magellan: Toward Building Entity Matching Management Systems |
2016 |
VLDB |
0.00017779048 |
| 788 |
ActiveClean: Interactive Data Cleaning For Statistical Modeling |
2016 |
VLDB |
0.00016618698 |
| 799 |
Entity Resolution: Theory, Practice & Open Challenges |
2012 |
VLDB |
0.00016479804 |
| 901 |
To Join or Not to Join? Thinking Twice about Joins before Feature Selection |
2016 |
SIGMOD |
0.00015462938 |
| 1,218 |
Snuba: Automating Weak Supervision to Label Training Data |
2019 |
VLDB |
0.00013221309 |
| 1,340 |
HoloDetect: Few-Shot Learning for Error Detection |
2019 |
SIGMOD |
0.00012492795 |
| 1,462 |
ARDA: Automatic Relational Data Augmentation for Machine Learning |
2020 |
VLDB |
0.00011866333 |
| 1,534 |
Data Management in Machine Learning: Challenges, Techniques, and Systems |
2017 |
SIGMOD |
0.00011462072 |
| 1,544 |
KATARA: A Data Cleaning System Powered by Knowledge Bases and Crowdsourcing |
2015 |
SIGMOD |
0.00011438274 |
| 1,895 |
Baran: Effective Error Correction via a Unified Context Representation and Transfer Learning |
2020 |
VLDB |
0.00010174634 |
| 2,176 |
Falcon: Scaling Up Hands-Off Crowdsourced Entity Matching to Build Cloud Services |
2017 |
SIGMOD |
9.3729351e-05 |
| 2,758 |
A Comprehensive Benchmark Framework for Active Learning Methods in Entity Matching |
2020 |
SIGMOD |
8.1668285e-05 |
| 2,895 |
Sato: Contextual Semantic Type Detection in Tables |
2020 |
VLDB |
7.9539265e-05 |
| 2,968 |
Raha: A Configuration-Free Error Detection System |
2019 |
SIGMOD |
7.7964476e-05 |
| 3,071 |
CrowdFill: Collecting Structured Data from the Crowd |
2014 |
SIGMOD |
7.6120337e-05 |
| 3,767 |
Cleaning Crowdsourced Labels Using Oracles for Statistical Classification |
2019 |
VLDB |
6.7748725e-05 |
| 3,900 |
SLiMFast: Guaranteed Results for Data Fusion and Source Reliability |
2017 |
SIGMOD |
6.649432e-05 |
| 4,123 |
Are Key-Foreign Key Joins Safe to Avoid when Learning High-Capacity Classifiers? |
2018 |
VLDB |
6.4290005e-05 |
| 4,453 |
CLAMShell: Speeding up Crowds for Low-latency Data Labeling |
2016 |
VLDB |
6.1690121e-05 |
| 4,911 |
Temporal Rules Discovery for Web Data Cleaning |
2016 |
VLDB |
5.8349225e-05 |
| 6,047 |
MDedup: Duplicate Detection with Matching Dependencies |
2020 |
VLDB |
5.2355891e-05 |
| 7,013 |
Qualitative Data Cleaning |
2016 |
VLDB |
4.8576683e-05 |
| 7,116 |
Crowdsourced Data Management: Overview and Challenges |
2017 |
SIGMOD |
4.8219732e-05 |