The “Something Something” Video Database for Learning and Evaluating Visual Common Sense

Goyal, Raghav; Kahou, Samira Ebrahimi; Michalski, Vincent; Materzyńska, Joanna; Westphal, Susanne; Kim, Heuna; Haenel, Valentin; Fruend, Ingo; Yianilos, P.N.; Mueller-Freitag, Moritz; Hoppe, Florian; Thurau, Christian; Bax, Ingo; Memisevic, Roland

doi:10.1109/iccv.2017.622

articleOct 1, 2017Closed access

The “Something Something” Video Database for Learning and Evaluating Visual Common Sense

RGRaghav Goyal SESamira Ebrahimi Kahou VMVincent Michalski JMJoanna Materzyńska SWSusanne Westphal

University of Tabriz

Indexed incrossref

Abstract

Neural networks trained on datasets such as ImageNet have led to major advances in visual object classification. One obstacle that prevents networks from reasoning more deeply about complex scenes and situations, and from integrating visual knowledge with natural language, like humans do, is their lack of common sense knowledge about the physical world. Videos, unlike still images, contain a wealth of detailed information about the physical world. However, most labelled video datasets represent high-level concepts rather than detailed physical aspects about actions and scenes. In this work, we describe our ongoing collection of the “something-something” database of video prediction tasks whose solutions…

Citation impact

1,430

total citations

FWCI: 23.10
Percentile: 100%
References: 54

Citations per year

Authors

14

Topics & keywords

Topics

Keywords

Computer science
Obstacle
Artificial intelligence
Commonsense reasoning
Object (grammar)
Common sense
Scale (ratio)
Visualization

UN Sustainable Development Goals

Quality Education

No related works found for this paper.