Learning Deep Structure-Preserving Image-Text Embeddings

Wang, Liwei; Li, Yin; Lazebnik, Svetlana

doi:10.1109/cvpr.2016.541

preprintJun 1, 2016Closed access

Learning Deep Structure-Preserving Image-Text Embeddings

LWLiwei Wang YLYin Li SLSvetlana Lazebnik

University of Illinois Urbana-Champaign · Georgia Institute of Technology

Indexed incrossref

Abstract

This paper proposes a method for learning joint embeddings of images and text using a two-branch neural network with multiple layers of linear projections followed by nonlinearities. The network is trained using a largemargin objective that combines cross-view ranking constraints with within-view neighborhood structure preservation constraints inspired by metric learning literature. Extensive experiments show that our approach gains significant improvements in accuracy for image-to-text and textto-image retrieval. Our method achieves new state-of-theart results on the Flickr30K and MSCOCO image-sentence datasets and shows promise on the new task of phrase localization on the Flickr30K Entities dataset.

Citation impact

804

total citations

FWCI: 53.56
Percentile: 100%
References: 92

Citations per year

Authors

3

Topics & keywords

Topics

Keywords

Computer science
Artificial intelligence
Ranking (information retrieval)
Sentence
Metric (unit)
Task (project management)
Phrase
Image (mathematics)

UN Sustainable Development Goals

Sustainable cities and communities

No related works found for this paper.

Funding

NS
National Science Foundation