Chimera: Large-Scale Classification using Machine Learning, Rules, and Crowdsourcing
Summary: Chimera classifies tens of millions of products into 5,000+ types by integrating machine learning, analyst-authored rules, and crowdsourcing. Its WalmartLabs deployment shows that scale invalidates conventional assumptions, making hybrid workflows and large-scale rule management essential for accuracy, cost, and continual improvement. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Chong Sun (WalmartLabs)
- 2. Narasimhan Rampalli (WalmartLabs)
- 3. Frank Yang (WalmartLabs)
- 4. AnHai Doan (University of Wisconsin; WalmartLabs)
BibTeX Citation
@article{sun_vldb14,
title = {{Chimera: Large-Scale Classification using Machine Learning, Rules, and Crowdsourcing}},
author = {Sun, Chong and Rampalli, Narasimhan and Yang, Frank and Doan, AnHai},
journal = {PVLDB},
series = {{VLDB} '14},
volume = {7},
number = {13},
pages = {1529--1540},
doi = {10.14778/2733004.2733024},
url = {https://doi.org/10.14778/2733004.2733024},
year = {2014}
}
Incoming Citations (Sorted by Pagerank)
Showing 10 of 10 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 1 of 1 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 439 | Corleone: Hands-Off Crowdsourcing for Entity Matching | 2014 | SIGMOD | 0.00018464913 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 7,357 | Crowdsourced Data Management: Overview and Challenges | 2017 | SIGMOD |
| 2 | 1,791 | Cheetah: A High Performance, Custom Data Warehouse on Top of MapReduce | 2010 | VLDB |
| 3 | 6,038 | Efficient Construction of Approximate Ad-Hoc ML models Through Materialization and Reuse | 2018 | VLDB |
| 4 | 4,812 | Clustering by Pattern Similarity in Large Data Sets | 2002 | SIGMOD |
| 5 | 3,714 | Large-Scale Machine Learning at Twitter | 2012 | SIGMOD |
| 6 | 6,908 | Cost-Effective Data Annotation using Game-Based Crowdsourcing | 2019 | VLDB |
| 7 | 4,859 | Machine Learning for Big Data | 2013 | SIGMOD |
| 8 | 8,495 | CrowdGame: A Game-Based Crowdsourcing System for Cost-Effective Data Labeling | 2019 | SIGMOD |
| 9 | 2,626 | Scaling Up Crowd-Sourcing to Very Large Datasets: A Case for Active Learning | 2015 | VLDB |
| 10 | 6,068 | Why Big Data Industrial Systems Need Rules and What We Can Do About It | 2015 | SIGMOD |