Mixer: Efficiently Understanding and Retrieving Visual Content at Web-scale
Summary: Mixer: class-based features with separate production/execution layers for images and videos. Two retrieval layers enable aggregation; on Baidu, model production time halved and throughput 9.14x, with 95% precision and 97% recall for video retrieval. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. An Qin (Baidu)
- 2. Mengbai Xiao (Shandong University)
- 3. Yongwei Wu (Baidu)
- 4. Xinjie Huang (Baidu)
- 5. Xiaodong Zhang (Ohio State University)
BibTeX Citation
@article{qin_vldb21,
title = {{Mixer: Efficiently Understanding and Retrieving Visual Content at Web-scale}},
author = {Qin, An and Xiao, Mengbai and Wu, Yongwei and Huang, Xinjie and Zhang, Xiaodong},
journal = {PVLDB},
series = {{VLDB} '21},
volume = {14},
number = {12},
pages = {2906--2917},
doi = {10.14778/3476311.3476371},
url = {https://doi.org/10.14778/3476311.3476371},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,373 | High-Throughput Vector Similarity Search in Knowledge Graphs | 2023 | SIGMOD | 0.0001088854 |
| 4,287 | VIVA: An End-to-End System for Interactive Video Analytics | 2022 | CIDR | 6.6869736e-05 |
| 8,370 | MicroNN: An On-device Disk-resident Updatable Vector Database | 2025 | SIGMOD | 5.34483e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 11 of 11 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 12,662 | Effective Data Co-Reduction for Multimedia Similarity Search | 2011 | SIGMOD |
| 2 | 9,337 | Qcluster: Relevance Feedback Using Adaptive Clustering for Content-Based Image Retrieval | 2003 | SIGMOD |
| 3 | 4,456 | Efficient and Cost-effective Techniques for Browsing and Indexing Large Video Databases | 2000 | SIGMOD |
| 4 | 8,749 | Optimizing Video Selection LIMIT Queries With Commonsense Knowledge | 2024 | VLDB |
| 5 | 10,839 | Craw: A Unified and Efficient Querying Framework for Large-Scale Video Datasets | 2026 | VLDB |
| 6 | 14,090 | Studying Interaction Methodologies in Video Retrieval | 2008 | VLDB |
| 7 | 12,505 | Effective Multi-Modal Retrieval based on Stacked Auto-Encoders | 2014 | VLDB |
| 8 | 12,710 | Multiple Feature Fusion for Social Media Applications | 2010 | SIGMOD |
| 9 | 3,411 | Spatial and Temporal Constrained Ranked Retrieval over Videos | 2022 | VLDB |
| 10 | 4,151 | Inter-Media Hashing for Large-scale Retrieval from Heterogeneous Data Sources | 2013 | SIGMOD |