Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
University of Illinois Urbana-Champaign · Broad Institute · +1 more institution
Abstract
The Flickr30k dataset has become a standard benchmark for sentence-based image description. This paper presents Flickr30k Entities, which augments the 158k captions from Flickr30k with 244k coreference chains linking mentions of the same entities in images, as well as 276k manually annotated bounding boxes corresponding to each entity. Such annotation is essential for continued progress in automatic image description and grounded language understanding. We present experiments demonstrating the usefulness of our annotations for text-to-image reference resolution, or the task of localizing textual entity mentions in an image, and for bidirectional image-sentence retrieval. These experiments confirm that we can…
Citation impact
- FWCI
- 24.33
- Percentile
- 100%
- References
- 99
Authors
6- BABryan A. PlummerCorresponding
University of Illinois Urbana-Champaign
- LWLiwei Wang
University of Illinois Urbana-Champaign
- CMChris M. Cervantes
University of Illinois Urbana-Champaign
- JCJuan Carlos Caicedo
Broad Institute, Fundación Universitaria Konrad Lorenz
- JHJulia Hockenmaier
University of Illinois Urbana-Champaign
Topics & keywords
- Computer science
- Coreference
- Natural language processing
- Artificial intelligence
- Sentence
- Annotation
- Phrase
- Task (project management)
- Quality Education