CogAgent: A Visual Language Model for GUI Agents
Tsinghua University · Zhipu AI (China)
Abstract
People are spending an enormous amount of time on dig-ital devices through graphical user interfaces (GUIs), e.g., computer or smartphone screens. Large language models (LLMs) such as ChatGPT can assist people in tasks like writing emails, but struggle to understand and interact with GUIs, thus limiting their potential to increase automation levels. In this paper, we introduce CogAgent, an 18-billion-parameter visual language model (VLM) specializing in GUI understanding and navigation. By utilizing both low-resolution and high-resolution image encoders, CogA-gent supports input at a resolution of1120 × 1120, enabling it to recognize tiny page elements and text. As a general-ist visual language model, CogAgent…
Citation impact
- FWCI
- 30.02
- Percentile
- 100%
- References
- 60
Authors
11Topics & keywords
- Computer science
- Programming language
- Human–computer interaction
- Natural language processing
- Artificial intelligence