Effective Multi-Modal Retrieval based on Stacked Auto-Encoders
Summary: Proposes an effective multi-modal retrieval framework using stacked auto-encoders to map heterogeneous media features into a shared low-dimensional space for cross-modal similarity search. Introduces a novel objective that models intra- and inter-modal semantics with minimal prior knowledge, and uses mini-batch, memory-efficient training, achieving state-of-the-art accuracy on real datasets. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Wei Wang (National University of Singapore)
- 2. Beng Chin Ooi (National University of Singapore)
- 3. Xiaoyan Yang (Advanced Digital Sciences Center)
- 4. Dongxiang Zhang (National University of Singapore)
- 5. Yueting Zhuang (Zhejiang University)
BibTeX Citation
@article{wang_vldb14,
title = {{Effective Multi-Modal Retrieval based on Stacked Auto-Encoders}},
author = {Wang, Wei and Ooi, Beng Chin and Yang, Xiaoyan and Zhang, Dongxiang and Zhuang, Yueting},
journal = {PVLDB},
series = {{VLDB} '14},
volume = {7},
number = {8},
pages = {649--660},
doi = {10.14778/2732296.2732300},
url = {https://doi.org/10.14778/2732296.2732300},
year = {2014}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 2 of 2 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 46 | A Quantitative Analysis and Performance Study for Similarity-Search Methods in High-Dimensional Spaces | 1998 | VLDB | 0.00044853085 |
| 4,060 | Inter-Media Hashing for Large-scale Retrieval from Heterogeneous Data Sources | 2013 | SIGMOD | 6.9322681e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 12,972 | Facilitating Multimedia Database Exploration through Visual Interfaces and Perpetual Query Reformulations | 1997 | VLDB |
| 2 | 11,378 | Effectiveness Perspectives and a Deep Relevance Model for Spatial Keyword Queries | 2023 | SIGMOD |
| 3 | 1,153 | Efficient Diversity-Aware Search | 2011 | SIGMOD |
| 4 | 13,779 | Studying Interaction Methodologies in Video Retrieval | 2008 | VLDB |
| 5 | 9,174 | Qcluster: Relevance Feedback Using Adaptive Clustering for Content-Based Image Retrieval | 2003 | SIGMOD |
| 6 | 3,353 | Efficient EMD-based Similarity Search in Multimedia Databases via Flexible Dimensionality Reduction | 2008 | SIGMOD |
| 7 | 4,060 | Inter-Media Hashing for Large-scale Retrieval from Heterogeneous Data Sources | 2013 | SIGMOD |
| 8 | 3,365 | Spatial and Temporal Constrained Ranked Retrieval over Videos | 2022 | VLDB |
| 9 | 12,371 | Effective Data Co-Reduction for Multimedia Similarity Search | 2011 | SIGMOD |
| 10 | 8,343 | An Interactive Multi-modal Query Answering System with Retrieval-Augmented Large Language Models | 2024 | VLDB |