Data Curation at Scale: The Data Tamer System
Summary: Introduces Data Tamer: an end-to-end, scalable data curation system that uses ML for attribute identification, schema/table grouping, transformations, and deduplication with human-in-the-loop visualization to assemble composites from sequences of sources. Evaluated on real enterprise workloads (up to tens of thousands of sources), showing ≈90% reduction in curation cost versus deployed production software. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Michael Stonebraker (Massachusetts Institute of Technology)
- 2. Daniel Bruckner (University of California Berkeley)
- 3. Ihab F. Ilyas (Qatar Computing Research Institute)
- 4. George Beskales (Qatar Computing Research Institute)
- 5. Mitch Cherniack (Brandeis University)
- 6. Stan Zdonik (Brown University)
- 7. Alexander Pagan (Massachusetts Institute of Technology)
- 8. Shan Xu (Verisk Analytics)
BibTeX Citation
@inproceedings{stonebraker_cidr13,
address = {Amsterdam, Netherlands},
series = {{CIDR} '13},
title = {{Data Curation at Scale: The Data Tamer System}},
booktitle = {Proceedings of the {Conference} on {Innovative} {Data} {Systems} {Research}},
author = {Stonebraker, Michael and Bruckner, Daniel and Ilyas, Ihab F. and Beskales, George and Cherniack, Mitch and Zdonik, Stan and Pagan, Alexander and Xu, Shan},
year = {2013}
}
Incoming Citations (Sorted by Pagerank)
Showing 38 of 38 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 3 of 3 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 91 | WebTables: Exploring the Power of Tables on the Web | 2008 | VLDB | 0.00034838835 |
| 94 | Potter's Wheel: An Interactive Data Cleaning System | 2001 | VLDB | 0.00034616103 |
| 4,239 | Semi-Automatic Schema Integration in Clio | 2007 | VLDB | 6.8107558e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 3,690 | Toward Large Scale Integration: Building a MetaQuerier over Databases on the Web | 2005 | CIDR |
| 2 | 5,644 | A Demo of the Data Civilizer System | 2017 | SIGMOD |
| 3 | 12,105 | Knowledge Curation and Knowledge Fusion: Challenges, Models, and Applications | 2015 | SIGMOD |
| 4 | 351 | CURE: An Efficient Clustering Algorithm for Large Databases | 1998 | SIGMOD |
| 5 | 13,375 | Reimagining Deep Learning Systems Through the Lens of Data Systems | 2024 | VLDB |
| 6 | 6,024 | Expand your Training Limits! Generating Training Data for ML-based Data Management | 2021 | SIGMOD |
| 7 | 5,397 | KATARA: Reliable Data Cleaning with Knowledge Bases and Crowdsourcing | 2015 | VLDB |
| 8 | 3,983 | Crowd-Based Deduplication: An Adaptive Approach | 2015 | SIGMOD |
| 9 | 2,398 | BigDansing: A System for Big Data Cleansing | 2015 | SIGMOD |
| 10 | 963 | The Data Civilizer System | 2017 | CIDR |