login

Detecting Visual Text

Published 3 June 2012
Jesse Dodge, Amit Goyal, Xufeng Han, Alyssa Mensch, Margaret Mitchell, Karl Stratos
Citations42

TL;DR

This work concretely defines what it means to be visual, annotate visual text and develops algorithms to automatically classify noun phrases as visual or non-visual, and finds that using text alone, it is able to achieve high accuracies at this task, and that incorporating features derived from computer vision algorithms improves performance.

Abstract

When people describe a scene, they often include information that is not visually apparent; sometimes based on background knowledge, sometimes to tell a story. We aim to separate visual text—descriptions of what is being seen—from non-visual text in natural images and their descriptions. To do so, we first concretely define what it means to be visual, annotate visual text and then develop algorithms to automatically classify noun phrases as visual or non-visual. We find that using text alone, we are able to achieve high accuracies at this task, and that incorporating features derived from computer vision algorithms improves performance. Finally, we show that we can reliably mine visual nouns and adjectives from large corpora and that we can use these effectively in the classification task. 1

Keywords

Computer Science