login

Natural Language Description of Human Activities from Video Images Based on Concept Hierarchy of Actions

International Journal of Computer VisionPublished 1 November 2002
Atsuhiro Kojima, Takeshi Tamura, Kunio Fukunaga
Citations356
SJR quartileQ1
SJR score3.14
SNIP5.15

TL;DR

A method for describing human activities from video images based on concept hierarchies of actions based on semantic primitives, which demonstrates the performance of the proposed method by several experiments.

Abstract

We propose a method for describing human activities from video images based on concept hierarchies of actions. Major difficulty in transforming video images into textual descriptions is how to bridge a semantic gap between them, which is also known as inverse Hollywood problem. In general, the concepts of events or actions of human can be classified by semantic primitives. By associating these concepts with the semantic features extracted from video images, appropriate syntactic components such as verbs, objects, etc. are determined and then translated into natural language sentences. We also demonstrate the performance of the proposed method by several experiments.

Keywords

Computer ScienceSocial Sciences