login

An experimental evaluation of OCR text representations for learning document classifiers

International Journal on Document Analysis and Recognition (IJDAR)Published 1 July 1998
Markus Junker, Rainer Hoch
Citations24
SJR quartileQ1
SJR score0.83
SNIP1.96

TL;DR

The results indicate that the use of n-grams is an attractive technique which can even compare to techniques relying on a morphological analysis, which holds for OCR texts as well as for correct ASCII texts.

Abstract

In the literature, many feature types are proposed for document classification. However, an extensive and systematic evaluation of the various approaches has not yet been done. In particular, evaluations on OCR documents are very rare. In this paper we investigate seven text representations based on n-grams and single words. We compare their effectiveness in classifying OCR texts and the corresponding correct ASCII texts in two domains: business letters and abstracts of technical reports. Our results indicate that the use of n-grams is an attractive technique which can even compare to techniques relying on a morphological analysis. This holds for OCR texts as well as for correct ASCII texts.

Keywords

Computer Science