LAVT: Language-Aware Vision Transformer for Referring Image Segmentation

Yang, Zhao; Wang, Jiaqi; Tang, Yansong; Chen, Kai; Zhao, Hengshuang; Torr, Philip

doi:10.1109/cvpr52688.2022.01762

article2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)Jun 1, 2022Closed access

LAVT: Language-Aware Vision Transformer for Referring Image Segmentation

ZYZhao Yang JWJiaqi Wang YTYansong Tang KCKai Chen HZHengshuang Zhao

University of Oxford · Shanghai Artificial Intelligence Laboratory · +4 more institutions

Indexed incrossref

Abstract

Referring image segmentation is a fundamental vision-language task that aims to segment out an object referred to by a natural language expression from an image. One of the key challenges behind this task is leveraging the referring expression for highlighting relevant positions in the image. A paradigm for tackling this problem is to leverage a powerful vision-language (“cross-madal”) decoder to fuse features independently extracted from a vision encoder and a language encoder. Recent methods have made remarkable advancements in this paradigm by exploiting Transformers as cross-modal decoders, concurrent to the Transformer's overwhelming success in many other vision-language tasks. Adopting a different…

Citation impact

326

total citations

FWCI: 17.23
Percentile: 100%
References: 81

Citations per year

Authors

6

Topics & keywords

Topics

Keywords

Computer science
Computer vision
Artificial intelligence
Image segmentation
Transformer
Segmentation
Engineering
Electrical engineering

UN Sustainable Development Goals

Quality Education

No related works found for this paper.