Arabic diacritization through full morphological tagging
Published 1 January 2007Open access
Nizar Habash, Owen Rambow
Citations139
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A diacritization system for written Arabic which is based on a lexical resource which combines a tagger and a lexeme language model which improves on the best results reported in the literature.
Abstract
We present a diacritization system for written Arabic which is based on a lexical resource. It combines a tagger and a lexeme language model. It improves on the best results reported in the literature.
Keywords
Computer Science
Arabic tokenization, part-of-speech tagging and morphological disambiguation in one fell swoop
443 Citations2005Nizar Habash, Owen Rambow
An approach to using a morphological analyzer for tokenizing and morphologically tagging Arabic words in one process using classifiers for individual morphological features, as well as ways of using these classifiers to choose among entries from the output of the analyzer.
SVMTool: A general POS Tagger Generator Based on Support Vector Machines
336 Citations2004Jesús Giménez, Lluı́s Màrquez
Fast methods for kernel-based text analysis
207 Citations2003Taku Kudo, Yūji Matsumoto
A Basket Mining algorithm is extended to convert a kernel-based classifier into a simple and fast linear classifier, showing results that show that these new classifiers are about 30 to 300 times faster than the standard kernel- based classifiers.
Maximum entropy based restoration of Arabic diacritics
140 Citations2006Imed Zitouni, Jeffrey Sorensen +1 more
A maximum entropy approach for restoring diacritics in a document that can easily integrate and make effective use of diverse types of information and integrates a wide array of lexical, segment-based and part-of-speech tag features.
Morphological tagging: data vs. dictionaries
108 Citations2000Jan Hajič
It is argued strongly that the use of an independent morphological dictionary is the preferred choice to more annotated data under such circumstances as data sparseness issue for inflectionally rich languages.
Automatic diacritization of Arabic for acoustic modeling in speech recognition
95 Citations2004Dimitra Vergyri, Katrin Kirchhoff
Various procedures that enable us to use acoustic information in combination with different levels of morphological and contextual constraints by automatically inserting the missing diacritics into the transcription are investigated.
Context-based morphological disambiguation with random fields
55 Citations2005Noah A. Smith, David A. Smith +1 more
A novel source-channel model is applied to the problem of morphological disambiguation (segmentation into morphemes, lemmatization, and POS tagging) for concatenative, templatic, and inflectional languages.
