Deep Learning for Entity Matching: A Design Space Exploration
Summary: Explores deep learning for entity matching, defines a DL design space (SIF, RNN, Attention, Hybrid), and maps NLP-style methods to EM. Empirically, DL lags on structured EM vs Magellan but excels on textual and dirty EM, guiding DL use for noisy data and outlining directions. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Sidharth Mudgal (University of Wisconsin)
- 2. Han Li (University of Wisconsin)
- 3. Theodoros Rekatsinas (University of Wisconsin)
- 4. AnHai Doan (University of Wisconsin)
- 5. Youngchoon Park (Johnson Controls)
- 6. Ganesh Krishnan (Walmart Labs)
- 7. Rohit Deep (Walmart Labs)
- 8. Esteban Arcaute (Meta)
- 9. Vijay Raghavendra (Walmart Labs)
BibTeX Citation
@inproceedings{mudgal_sigmod18,
title = {{Deep Learning for Entity Matching: A Design Space Exploration}},
author = {Mudgal, Sidharth and Li, Han and Rekatsinas, Theodoros and Doan, AnHai and Park, Youngchoon and Krishnan, Ganesh and Deep, Rohit and Arcaute, Esteban and Raghavendra, Vijay},
series = {{SIGMOD} '18},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3183713.3196926},
url = {https://dl.acm.org/doi/10.1145/3183713.3196926},
year = {2018}
}
Incoming Citations (Sorted by Pagerank)
Showing 50 of 89 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 11 of 11 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 196 | CrowdER: Crowdsourcing Entity Resolution | 2012 | VLDB | 0.00025780596 |
| 205 | Snorkel: Rapid Training Data Creation with Weak Supervision | 2018 | VLDB | 0.00025235185 |
| 248 | Evaluation of entity resolution approaches on real-world match problems | 2010 | VLDB | 0.00023278354 |
| 439 | Corleone: Hands-Off Crowdsourcing for Entity Matching | 2014 | SIGMOD | 0.00018464913 |
| 516 | Data Curation at Scale: The Data Tamer System | 2013 | CIDR | 0.00017171198 |
| 529 | Magellan: Toward Building Entity Matching Management Systems | 2016 | VLDB | 0.00017096361 |
| 619 | Reasoning about Record Matching Rules | 2009 | VLDB | 0.00015707247 |
| 626 | Entity Resolution: Theory, Practice & Open Challenges | 2012 | VLDB | 0.00015656958 |
| 2,120 | Comparative Analysis of Approximate Blocking Techniques for Entity Resolution | 2016 | VLDB | 9.1406654e-05 |
| 3,192 | Fonduer: Knowledge Base Construction from Richly Formatted Data | 2018 | SIGMOD | 7.65035e-05 |
| 3,713 | Generating Concise Entity Matching Rules | 2017 | SIGMOD | 7.176496e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 4,056 | A Benchmarking Study of Embedding-based Entity Alignment for Knowledge Graphs | 2020 | VLDB |
| 2 | 248 | Evaluation of entity resolution approaches on real-world match problems | 2010 | VLDB |
| 3 | 6,192 | Pre-trained Embeddings for Entity Resolution: An Experimental Analysis | 2023 | VLDB |
| 4 | 5,606 | Analyzing How BERT Performs Entity Matching | 2022 | VLDB |
| 5 | 1,402 | Creating Embeddings of Heterogeneous Relational Datasets for Data Integration Tasks | 2020 | SIGMOD |
| 6 | 5,730 | Domain Adaptation for Deep Entity Resolution | 2022 | SIGMOD |
| 7 | 9,560 | Ground Truth Inference for Weakly Supervised Entity Matching | 2023 | SIGMOD |
| 8 | 2,463 | A Comprehensive Benchmark Framework for Active Learning Methods in Entity Matching | 2020 | SIGMOD |
| 9 | 9,610 | The Battleship Approach to the Low Resource Entity Matching Problem | 2023 | SIGMOD |
| 10 | 2,475 | Deep Learning for Blocking in Entity Matching: A Design Space Exploration | 2021 | VLDB |