Multimodal Learning With Transformers: A Survey

Tsinghua University · University of Surrey · +1 more institution

PubMed
Indexed incrossrefpubmed

Abstract

Transformer is a promising neural network learner, and has achieved great success in various machine learning tasks. Thanks to the recent prevalence of multimodal applications and Big Data, Transformer-based multimodal learning has become a hot topic in AI research. This paper presents a comprehensive survey of Transformer techniques oriented at multimodal data. The main contents of this survey include: (1) a background of multimodal learning, Transformer ecosystem, and the multimodal Big Data era, (2) a systematic review of Vanilla Transformer, Vision Transformer, and multimodal Transformers, from a geometrically topological perspective, (3) a review of multimodal Transformer applications, via two important…

Citation impact

837
total citations
FWCI
94.27
Percentile
100%
References
443
Citations per year

Authors

3

Topics & keywords

Keywords
  • Transformer
  • Computer science
  • Multimodal learning
  • Multimodal therapy
  • Artificial intelligence
  • Machine learning
  • Engineering
  • Electrical engineering
No related works found for this paper.

Funding