login

Improvements in BBN's HMM-Based Offline Arabic Handwriting Recognition System

Published 1 January 2009
Shirin Saleem, Huaigu Cao, Krishna Subramanian, Matin Kamali, Rohit Prasad, Prem Natarajan
Citations33

TL;DR

A novel integration of structural features in the HMM framework which exclusively results in a 9% relative improvement in performance is proposed, and a relative reduction of 17% in word error rate over the baseline Arabic handwriting recognition system is demonstrated.

Abstract

Offline handwriting recognition of free-flowing Arabic text is a challenging task due to the plethora of factors that contribute to the variability in the data. In this paper, we address some of these sources of variability, and present experimental results on a large corpus of handwritten documents. Specific techniques such as the application of context-dependent Hidden Markov Models (HMMs) for the cursive Arabic script, unsupervised adaptation to account for the stylistic variations across scribes, and image pre-processing to remove ruled-lines are explored. In particular, we proposed a novel integration of structural features in the HMM framework which exclusively results in a 9% relative improvement in performance. Overall, we demonstrate a relative reduction of 17% in word error rate over our baseline Arabic handwriting recognition system.

Keywords

Computer Science