Vision-Language Pre-Training with Triple Contrastive Learning

Yang, Jinyu; Duan, Jiali; Tran, Son N.; Xu, Yi; Chanda, Sampath; Chen, Li‐Qun; Zeng, Belinda; Chilimbi, Trishul; Huang, Junzhou

doi:10.1109/cvpr52688.2022.01522

article2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)Jun 1, 2022Closed access

Vision-Language Pre-Training with Triple Contrastive Learning

JYJinyu Yang JDJiali Duan SNSon N. Tran YXYi Xu SCSampath Chanda

The University of Texas at Arlington · Amazon (Germany)

Indexed incrossref

Abstract

Vision-language representation learning largely benefits from image-text alignment through contrastive losses (e.g., InfoNCE loss). The success of this alignment strategy is attributed to its capability in maximizing the mutual information (MI) between an image and its matched text. However, simply performing cross-modal alignment (CMA) ignores data potential within each modality, which may result in degraded representations. For instance, although CMA-based models are able to map image-text pairs close together in the embedding space, they fail to ensure that similar inputs from the same modality stay close by. This problem can get even worse when the pre-training data is noisy. In this paper, we propose…

Citation impact

268

total citations

FWCI: 15.17
Percentile: 100%
References: 75

Citations per year

Authors

9

Topics & keywords

Topics

Keywords

Computer science
Modality (human–computer interaction)
Embedding
Representation (politics)
Artificial intelligence
Image (mathematics)
Feature learning
Modal

UN Sustainable Development Goals

Quality Education

No related works found for this paper.

Funding

CP
Cancer Prevention and Research Institute of Texas
Award: RP190107