Mixer: Efficiently Understanding and Retrieving Visual Content at Web-scale
Summary: Mixer: class-based features with separate production/execution layers for images and videos. Two retrieval layers enable aggregation; on Baidu, model production time halved and throughput 9.14x, with 95% precision and 97% recall for video retrieval. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. An Qin (Baidu)
- 2. Mengbai Xiao (Shandong University)
- 3. Yongwei Wu (Baidu)
- 4. Xinjie Huang (Baidu)
- 5. Xiaodong Zhang (Ohio State University)
BibTeX Citation
@article{qin_vldb21,
title = {{Mixer: Efficiently Understanding and Retrieving Visual Content at Web-scale}},
author = {Qin, An and Xiao, Mengbai and Wu, Yongwei and Huang, Xinjie and Zhang, Xiaodong},
journal = {PVLDB},
series = {{VLDB} '21},
volume = {14},
number = {12},
pages = {2906--2917},
doi = {10.14778/3476311.3476371},
url = {https://doi.org/10.14778/3476311.3476371},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,631 | High-Throughput Vector Similarity Search in Knowledge Graphs | 2023 | SIGMOD | 0.00010174628 |
| 4,256 | VIVA: An End-to-End System for Interactive Video Analytics | 2022 | CIDR | 6.8018439e-05 |
| 9,915 | MicroNN: An On-device Disk-resident Updatable Vector Database | 2025 | SIGMOD | 5.1955087e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 11 of 11 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 4,618 | Towards Effective Indexing for Very Large Video Sequence Database | 2005 | SIGMOD |
| 2 | 12,371 | Effective Data Co-Reduction for Multimedia Similarity Search | 2011 | SIGMOD |
| 3 | 9,174 | Qcluster: Relevance Feedback Using Adaptive Clustering for Content-Based Image Retrieval | 2003 | SIGMOD |
| 4 | 4,459 | Efficient and Cost-effective Techniques for Browsing and Indexing Large Video Databases | 2000 | SIGMOD |
| 5 | 8,589 | Optimizing Video Selection LIMIT Queries With Commonsense Knowledge | 2024 | VLDB |
| 6 | 13,779 | Studying Interaction Methodologies in Video Retrieval | 2008 | VLDB |
| 7 | 12,214 | Effective Multi-Modal Retrieval based on Stacked Auto-Encoders | 2014 | VLDB |
| 8 | 12,419 | Multiple Feature Fusion for Social Media Applications | 2010 | SIGMOD |
| 9 | 3,365 | Spatial and Temporal Constrained Ranked Retrieval over Videos | 2022 | VLDB |
| 10 | 4,060 | Inter-Media Hashing for Large-scale Retrieval from Heterogeneous Data Sources | 2013 | SIGMOD |