login

Generating natural language description of human behavior from video images

Published 11 November 2002
Atsuhiro Kojima, Masayuki Izumi, T. Tamura, Koichi Fukunaga
Citations37

TL;DR

This work proposes an approach to generating a natural language description of human behavior appearing in real video images using a model based method and a technique of machine translation.

Abstract

In visual surveillance applications, it is becoming popular to perceive video images and to interpret them using natural language concepts. We propose an approach to generating a natural language description of human behavior appearing in real video images. First, a head region of a human, on behalf of the whole body, is extracted from each frame. Using a model based method, three dimensional pose and position of the head are estimated. Next, the trajectory of these parameters is divided into segments of monotonous motions. For each segment, we evaluate conceptual features such as degree of change of pose and position and that of relative distance to some objects in the surroundings, and so on. By calculating the product of these feature values, a most suitable verb is selected and other syntactic elements are supplied. Finally natural language text is generated using a technique of machine translation.

Keywords

Computer ScienceSocial Sciences