Introduction to Arabic Natural Language Processing
Synthesis lectures on human language technologiesPublished 1 January 2010
Nizar Habash
Citations290
SJR quartileQ3
SJR score0.12
SNIP0.00
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
This book provides system developers and researchers in natural language processing and computational linguistics with the necessary background information for working with the Arabic language. The go
Keywords
Computer Science
Computational LinguisticsA Systematic Comparison of Various Statistical Alignment Models
3,928 Citations2003Franz Josef Och, Hermann Ney
An important result is that refined alignment models with a first-order dependence and a fertility model yield significantly better results than simple heuristic models.
Speech and language processing
3,589 Citations2010Dan Jurafsky, James Martin +1 more
It is now clear that HAL’s creator, Arthur C. Clarke, was a little optimistic in predicting when an artificial agent such as HAL would be avail-able.
Accurate unlexicalized parsing
3,055 Citations2003Dan Klein, Christopher D. Manning
It is demonstrated that an unlexicalized PCFG can parse much more accurately than previously shown, by making use of simple, linguistically motivated state splits, which break down false independence assumptions latent in a vanilla treebank grammar.
The Berkeley FrameNet Project
2,564 Citations1998Collin F. Baker, Charles J. Fillmore +1 more
This report will present the project's goals and workflow, and information about the computational tools that have been adapted or created in-house for this work.
Computational LinguisticsThe Proposition Bank: An Annotated Corpus of Semantic Roles
2,302 Citations2005Martha Palmer, Daniel Gildea +1 more
An automatic system for semantic role tagging trained on the corpus is described and the effect on its performance of various types of information is discussed, including a comparison of full syntactic parsing with a flat representation and the contribution of the empty trace categories of the treebank.
EuroWordNet: A multilingual database with lexical semantic networks
853 Citations1998Piek Vossen
Cross-Linguistic Alignment of Wordnets with an Inter-Lingual-Index W. Peters, et al.
Language<b>Meaning and grammar:</b> An introduction to semantics. By Gennaro Chierchia and Sally McConnell-Ginet. Cambridge, MA: MIT Press, 1990. Pp. xiii, 460. $29.95.
814 Citations1991Greg N. Carlson
This self-contained introduction to natural language semantics addresses the major theoretical questions in the field and introduces the systematic study of linguistic meaning through a sequence of formal tools and their linguistic applications.
Natural Language EngineeringMaltParser: A language-independent system for data-driven dependency parsing
805 Citations2007Joakim Nivre, Johan Hall +6 more
Experimental evaluation confirms that MaltParser can achieve robust, efficient and accurate parsing for a wide range of languages without language-specific enhancements and with rather limited amounts of training data.
The Phonology And Morphology Of Arabic
680 Citations2002Janet C. E. Watson
The Phoneme System of Arabic is described as a system of phonemes that combines syllable structure, word stress, and phonological features of Arabic with those of other Romance languages.
Lecture notes in computer sciencePharaoh: A Beam Search Decoder for Phrase-Based Statistical Machine Translation Models
634 Citations2004Philipp Koehn
Pharaoh, a freely available decoder for phrase-based statistical machine translation models is described, which is the implement at ion of an efficient dynamic programming search algorithm with lattice generation and XML markup for external components.
Modern Language JournalA Grammar of the Arabic Language
572 Citations1970Sami A. Hanna, William Wright
IEEE Transactions on Pattern Analysis and Machine IntelligenceOffline Arabic handwriting recognition: a survey
492 Citations2006Liana M. Lorigo, Venu Govindaraju
This paper is the first survey to focus on Arabic handwriting recognition and the first Arabic character recognition survey to provide recognition rates and descriptions of test data for the approaches discussed.
Arabic tokenization, part-of-speech tagging and morphological disambiguation in one fell swoop
443 Citations2005Nizar Habash, Owen Rambow
An approach to using a morphological analyzer for tokenizing and morphologically tagging Arabic words in one process using classifiers for individual morphological features, as well as ways of using these classifiers to choose among entries from the output of the analyzer.
Improving stemming for Arabic information retrieval
360 Citations2002Leah S. Larkey, Lisa Ballesteros +1 more
Several light stemmers based on heuristics and a statistical stemmer based on co-occurrence for Arabic retrieval and the retrieval effectiveness of these stemmers and of a morphological analyzer on the TREC-2001 data were compared.
Automatic tagging of Arabic text
311 Citations2004Mona Diab, Kadri Hacıoğlu +1 more
A Support Vector Machine (SVM) based approach to automatically tokenize, tag and annotate base phrases (BPs) in Arabic text and adapt highly accurate tools that have been developed for English text and apply them to Arabic text.
Arabic preprocessing schemes for statistical machine translation
254 Citations2006Nizar Habash, Fatiha Sadat
Journal of the American Society for Information Science and TechnologyArabic morphological analysis techniques: A comprehensive survey
239 Citations2003Imad A. Al‐Sughaiyer, Ibrahim A. Al‐Kharashi
This paper introduces, classifies, and surveys Arabic morphological analysis techniques, and summarizes and organize the information available in the literature in an attempt to motivate researchers to look into these techniques and try to develop more advanced ones.
Kluwer eBooksMultilingual text-to-speech synthesis : the Bell Labs approach
225 Citations1998Richard Sproat, L.C.W. Pols
This chapter discusses Multilingual Text Analysis with a focus on Character Set Encodings and Grammatical Labels, and concludes with a discussion of the role of tone in the construction of sentences.
Americanae (AECID Library)A Short Reference Grammar of Moroccan Arabic
220 Citations1962Richard S. Harrell
Digital Academic REpository of VU University Amsterdam (Vrije Universiteit Amsterdam)Introducing the Arabic WordNet project
215 Citations2006Christiane Fellbaum, Munya Alkhalifa +2 more
The approach towards building a lexical resource in Standard Arabic will be based on the design and contents of the universally accepted Princeton WordNet and will be mappable straightforwardly onto PWN 2.0 and EuroWordNet, enabling translation on the lexical level to English and dozens of other languages.
Text, speech and language technologyArabic Computational Morphology: Knowledge-based and Empirical Methods
198 Citations2007Abdelhadi Soudi, Günter Neumann +1 more
Americanae (AECID Library)A short reference grammar of Iraqi Arabic
179 Citations1963Wallace M. Erwin
Morphological analysis for statistical machine translation
177 Citations2004Young‐Suk Lee
A novel morphological analysis technique which induces a morphological and syntactic symmetry between two languages with highly asymmetrical morphological structures to improve statistical machine translation qualities.
Building a shallow Arabic Morphological Analyzer in one day
173 Citations2002Kareem Darwish
The paper presents a rapid method of developing a shallow Arabic morphological analyzer based on automatically derived rules and statistics that will only be concerned with generating the possible roots of any given Arabic word.
On arabic search
172 Citations2002Mohammed Aljlayl, Ophir Frieder
This work proposes a novel light-stemming algorithm for Arabic texts that significantly outperforms the root-based algorithm and shows that a significant improvement in retrieval precision can be achieved with light inflectional analysis of Arabic words.
Arabic finite-state morphological analysis and generation
162 Citations1996Kenneth R. Beesley
A large-scale system that performs morphological analysis and generation of on-line Arabic words represented in the standard orthography, whether fully voweled, partially voweled or unvoweled, using Xerox Finite-State Morphology tools.
Machine transliteration of names in Arabic text
160 Citations2002Yaser Al-Onaizan, Kevin Knight
A new spelling-based model is introduced that is much more accurate than state-of-the-art phonetic-based models and can be trained on easier-to-obtain training data.
Arabic morphological tagging, diacritization, and lemmatization using lexeme models and feature ranking
152 Citations2008Ryan M. Roth, Owen Rambow +3 more
This work investigates the tasks of general morphological tagging, diacritization, and lemmatization for Arabic and shows that for all tasks, both modeling the lexeme explicitly and retuning the weights of individual classifiers for the specific task, improve the performance.
Spoken Arabic dialect identification using phonotactic modeling
149 Citations2009Fadi Biadsy, Julia Hirschberg +1 more
A system that automatically identifies the Arabic dialect of a speaker given a sample of his/her speech is described, and the phonotactic approach proves to be effective in identifying these dialects with considerable overall accuracy.
Maximum entropy based restoration of Arabic diacritics
140 Citations2006Imed Zitouni, Jeffrey Sorensen +1 more
A maximum entropy approach for restoring diacritics in a document that can easily integrate and make effective use of diverse types of information and integrates a wide array of lexical, segment-based and part-of-speech tag features.
Arabic diacritization through full morphological tagging
139 Citations2007Nizar Habash, Owen Rambow
A diacritization system for written Arabic which is based on a lexical resource which combines a tagger and a lexeme language model which improves on the best results reported in the literature.
Arabic named entity recognition using optimized feature sets
126 Citations2008Yassine Benajiba, Mona Diab +1 more
This paper investigates the impact of using different sets of features in two discriminative machine learning frameworks, namely, Support Vector Machines and Conditional Random Fields using Arabic data.
Four techniques for online handling of out-of-vocabulary words in Arabic-English statistical machine translation
107 Citations2008Nizar Habash
Four techniques for online handling of Out-of-Vocabulary words in Phrase-based Statistical Machine Translation using spelling expansion, morphological expansion, dictionary term expansion and proper name transliteration to reuse or extend a phrase table are presented.
Computer Speech & LanguageMorphology-based language modeling for conversational Arabic speech recognition
105 Citations2005Katrin Kirchhoff, Dimitra Vergyri +3 more
Four different approaches to morphology-based language modeling are presented, including a novel technique called factored language models, and results are presented for both rescoring and first-pass recognition experiments.
Natural Language EngineeringAdding semantic roles to the Chinese Treebank
104 Citations2008Nianwen Xue, Martha Palmer
This discussion focuses on the syntactic variations in the realization of the arguments as well as the approach to annotating dislocated and discontinuous arguments, and the creation of a lexical database of frame files and its role in guiding predicate-argument annotation.
Language Resources and EvaluationMorphological Annotation of Quranic Arabic
103 Citations2010Kais Dukes, Nizar Habash
Those who reported purchasing by the box/pack had greater odds of every day e-cigarette use compared to some day use, and this finding held for both males, females, and all device types.
2003 IEEE International Conference on Acoustics, Speech, and Signal Processing, 2003. Proceedings. (ICASSP '03).Novel approaches to Arabic speech recognition: report from the 2002 Johns-Hopkins Summer Workshop
101 Citations2003Katrin Kirchhoff, Jeffrey A. Bilmes +11 more
Novel approaches to automatic vowel restoration, morphology-based language modeling and the integration of out-of-corpus language model data are presented, and significant word error rate improvements on the LDC Arabic CallHome task are reported.
Developing an Arabic treebank
94 Citations2004Mohamed Maamouri, Ann Bies
This paper addresses the following questions from the experience of developing a large-scale corpus of Arabic text annotated for morphological information, part-of-speech, English gloss, and syntactic structure.
Arabic diacritization using weighted finite-state transducers
91 Citations2005Rani Nelken, Stuart M. Shieber
A novel algorithm for restoring symbols of Arabic without short vowels and additional diacritics is presented, using a cascade of probabilistic finite-state transducers trained on the Arabic treebank, integrating a word-based language model, a letter-basedlanguage model, and an extremely simple morphological model.
Columbia Academic Commons (Columbia University)Online Arabic Handwriting Recognition Using Hidden Markov Models
90 Citations2006Fadi Biadsy, Jihad El‐Sana +1 more
This paper introduces a Hidden Markov Model (HMM) based system to provide solutions for most of the difficulties inherent in recognizing Arabic script including: letter connectivity, position-dependent letter shaping, and delayed strokes.
Empirical studies in strategies for Arabic retrieval
90 Citations2002Jinxi Xu, Alexander Fraser +1 more
Evaluated search strategies for Arabic monolingual and cross-lingual retrieval, using the TREC Arabic corpus as the test-bed, found that spelling normalization and stemming have little impact and a novel thesaurus-based technique is proposed.
Computer Speech & LanguagePhonetization of Arabic: rules and algorithms
85 Citations2003Yousif A. El-Imam
This paper presents a detailed investigation into all aspects of the phonetization of SA for the purpose of developing a comprehensive system for letter-to-sound conversion for the standard Arabic language and assessing the quality of the letter- to-sound transcription system.
Machine translation of very close languages
84 Citations2000Jan Hajič, Jan Hric +1 more
It is argued that for really close languages it is possible to obtain better translation quality by means of simpler methods by using transfer-based and word-for-word MT systems.
Combination of Arabic preprocessing schemes for statistical machine translation
84 Citations2006Fatiha Sadat, Nizar Habash
This paper studies the effect of different word-level preprocessing schemes for Arabic on the quality of phrase-based statistical machine translation and presents and evaluates different methods for combining pre processing schemes resulting in improved translation quality.
Computational LinguisticsMultitiered Nonlinear Morphology Using Multitape Finite Automata: A Case Study on Syriac and Arabic
74 Citations2000George Kiraz
This paper presents a computational model for nonlinear morphology with illustrations from Syriac and Arabic that allows for multiple lexical representations corresponding to the multiple tiers of autosegmental phonology.
Segmentation for English-to-Arabic statistical machine translation
68 Citations2008Ibrahim Badr, Rabih Zbib +1 more
It is shown that morphological decomposition of the Arabic source is beneficial, especially for smaller-size corpora, and recombination techniques are investigated, and the use of Factored Translation Models for English-to-Arabic translation is reported on.
Morphological analysis and generation for Arabic dialects
68 Citations2005Nizar Habash, Owen Rambow +1 more
Magead provides an analysis to a root+pattern representation, it has separate phonological and orthographic representations, and it allows for combining morphemes from different dialects.
CLIR Experiments at Maryland for TREC-2002: Evidence Combination for Arabic-English Retrieval
65 Citations2003Kareem Darwish, Douglas W. Oard
The focus of the experiments reported in this paper was techniques for combining evidence for crosslanguage retrieval, searching Arabic documents using English queries, and a new technique that exploits translation probability information was found to outperform a comparable technique in which that information was not used.
Text REtrieval ConferenceThe TREC-2001 Cross-Language Information Retrieval Track: Searching Arabic using English, French or Arabic Queries
64 Citations2001Fredric C. Gey, Douglas W. Oard
A variety of approaches were tried and a rich set of experiments performed using resources such as machine translation, parallel corpora, several approaches to stemming and/or morphology, and both pre-translation and post-translation blind relevance feedback.
On the use of morphological analysis for dialectal Arabic speech recognition
62 Citations2006Mohamed Afify, Ruhi Sarikaya +3 more
A simple word decomposition algorithm is introduced which only requires a text corpus and a predefined list of affixes to create the lexicon for Iraqi Arabic ASR and results in about 10% relative improvement in word error rate (WER).
Synthesis lectures on human language technologiesIntroduction to Chinese Natural Language Processing
62 Citations2010Kam‐Fai Wong, Wenjie Li +2 more
MaTrEx: the DCU Machine Translation System for IWSLT 2007
60 Citations2006Hany Hassan, Yanjun Ma +1 more
The machine translation system developed at DCU that was used for the second participation in the evaluation campaign of the International Workshop on Spoken Language Translation (IWSLT 2007) is described and some new methods to improve system quality are focused on.
A Dependency Treebank of the Quran using traditional Arabic grammar
59 Citations2010Kais Dukes, Tim Buckwalter
The Quranic Arabic Dependency Treebank (QADT) is presented and it is reported on how online collaborative annotation was used to bring together Quranic scholars and Arabic language experts to ensure a high level of accuracy for grammatical analysis of the entire Quran.
Synchronous tree adjoining machine translation
58 Citations2009Steve DeNeefe, Kevin Knight
A novel method for learning a type of Synchronous Tree Adjoining Grammar and associated probabilities from aligned tree/string training data and a method of converting these grammars to a weakly equivalent tree transducer for decoding is introduced.
Term selection for searching printed Arabic
58 Citations2002Kareem Darwish, Douglas W. Oard
Alternative choices of indexing terms are explored using both an existing electronic text collection and a newly developed collection built from images of actual printed Arabic documents, and character n-grams or lightly stemmed words were found to typically yield near-optimal retrieval effectiveness.
A cross-cultural comparison of american, Palestinian, and Swedish perception of charismatic speech
57 Citations2008Fadi Biadsy, Andrew Rosenberg +3 more
Medical Entomology and ZoologyA reference grammar of Egyptian Arabic
56 Citations2009Ernest T. Abdel-Massih, Zaki N. Abdel-Malek +1 more
A chart parser for analyzing modern standard Arabic sentence
56 Citations2003Eman Othman, Khaled Shaalan +1 more
Arabic morphology generation using a concatenative strategy
56 Citations2000Violetta Cavalli‐Sforza, Abdelhadi Soudi +1 more
This paper describes an approach to reducing the complexity of Arabic morphology generation using discrimination trees and transformational rules, and gains a significant reduction in the number of rules required, as much as a factor of three for certain verb types.
Context-based morphological disambiguation with random fields
55 Citations2005Noah A. Smith, David A. Smith +1 more
A novel source-channel model is applied to the problem of morphological disambiguation (segmentation into morphemes, lemmatization, and POS tagging) for concatenative, templatic, and inflectional languages.
Arabic diacritization in the context of statistical machine translation
55 Citations2007Mona Diab, Mahmoud Ghoneim +1 more
Improving the Arabic pronunciation dictionary for phone and word recognition with linguistically-based pronunciation rules
54 Citations2009Fadi Biadsy, Nizar Habash +1 more
It is demonstrated that linguistically motivated pronunciation rules can significantly improve both MSA phone recognition and MSA word recognition accuracies over a baseline system using pronunciation rules typically employed in previous work on MSA Automatic Speech Recognition (ASR).
Journal of Colloid and Interface ScienceParsing the Arabic Treebank: Analysis and Improvements
53 Citations2006Seth Kulick, Ryan Gabbard +1 more
Some issues around evaluation are discussed and it is shown that current Arabic parsing performance is not quite as bad as previously thought and some modifications to the parser are presented which provide modest increases in performance.
Improving Arabic-Chinese statistical machine translation using English as pivot language
51 Citations2009Nizar Habash, Jun Hu
A comparison of two approaches for Arabic-Chinese machine translation using English as a pivot language: sentence pivoting and phrase-table pivoting shows that using Englishas a pivot in either approach outperforms direct translation from Arabic to Chinese.
Bridging the inflection morphology gap for Arabic statistical machine translation
46 Citations2006Andreas Zollmann, Ashish Venugopal +1 more
This work presents techniques that select appropriate word segmentations in the morphologically rich source language based on contextual relationships in the target language to improve translation quality above state-of-the-art on a limited-data Arabic to English speech translation task.
Automatic treebank-based acquisition of Arabic LFG dependency structures
45 Citations2009Lamia Tounsi, Mohammed Attia +1 more
The LFG grammar acquisition approach to Arabic and the Penn Arabic Treebank (ATB) is extended, adapting and extending the methodology of (Cahill and al., 2004) originally developed for English.
The impact of morphological stemming on Arabic mention detection and coreference resolution
44 Citations2005Imed Zitouni, Jeff Sorensen +2 more
This paper presents an in-depth investigation of the entity detection and recognition task for Arabic, highlighting why segmentation is a necessary prerequisite for EDR, presenting a finite-state statistical segmenter, and examining how the resulting segments can be better included into a mention detection system and an entity recognition system.
Improving Arabic Dependency Parsing with Lexical and Inflectional Morphological Features
44 Citations2010Yuval Marton, Nizar Habash +1 more
It is shown that training the parser using a simple regular expressive extension of an impoverished POS tagset with high prediction accuracy does better than using a highly informative POS tag set with only medium prediction accuracy, although the latter performs best on gold input.
International Journal of Computer Processing Of LanguagesDetection and Correction of Non-Words in Arabic: A Hybrid Approach
43 Citations2007Bassam Haddad, Mustafa Yaseen
Novel probabilistic measures for completing the task of the correction by locating, reducing and ranking of the most probable correction candidates in Arabic derivative words are proposed.
Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE<title>Robust language-independent OCR system</title>
43 Citations1999Zhidong Lu, Issam Bazzi +4 more
A language-independent optical character recognition system that is capable, in principle, of recognizing printed text from most of the world's languages, using hidden Markov modeling technology to model each character.
Improving NER in Arabic Using a Morphological Tagger.
43 Citations2008B. Farber, Dayne Freitag +2 more
A named entity recognition system for Arabic is discussed, and how MADA, a full morphological tagger which uses a morphological analyzer is incorporated, yielding a 14\% reduction in error over the baseline.
Using prosody and phonotactics in Arabic dialect identification
42 Citations2009Fadi Biadsy, Julia Hirschberg
It is shown that prosodic features can significantly improve identification, over a purely phonotactic-based approach, with an identification accuracy of 86.33% for 2m utterances.
Multi-tape two-level morphology
39 Citations1994George Kiraz
It is illustrated that if finite-state transducers in a standard two-level morphology model are replaced with multi-tape auxiliary versions (AFSTs), one can account for Semitic root-and-pattern morphology using high level notation.
Arabic OCR error correction using character segment correction, language modeling, and shallow morphology
39 Citations2006Walid Magdy, Kareem Darwish
Experimentation shows that character segment based correction is superior to single character correction and that language modeling boosts correction, by improving the ranking of candidate corrections, while shallow morphology had a small adverse effect.
Determining Case in Arabic: Learning Complex Linguistic Behavior Requires Complex Linguistic Features
39 Citations2007Nizar Habash, Ryan Gabbard +3 more
A careful error analysis suggests that when one accounts for annotation errors in the gold standard, the error rate in the automatic determination of case in Arabic drops to 0.8%, with the hand-written rules outperforming the machine learning-based system.
Development of the SRI/nightingale Arabic ASR system
37 Citations2008Dimitra Vergyri, Arindam Mandal +10 more
The large vocabulary automatic speech recognition system developed for Modern Standard Arabic by the SRI/Nightingale team is described and how system performance is affected by different development choices, ranging from text processing and lexicon to decoding system architecture design.
ElixirFM
36 Citations2007Otakar Smrž
ElixirFM is its high-level implementation that reuses and extends the Functional Morphology library for Haskell, and the lexicon is derived from the open-source Buckwalter lexicon and is enhanced with information sourcing from the syntactic annotations of the Prague Arabic Dependency Treebank.
Syntactic Annotation in Columbia Arabic Treebank
35 Citations2009Nizar Habash, Reem Faraj +1 more
CATiB uses linguistic representation and terminology inspired by the long tradition of Arabic syntactic studies to make it easier to train annotators and not be restricted to hire annotators who have degrees in linguistics.
Morpho-syntactic Arabic preprocessing for Arabic-to-English statistical machine translation
35 Citations2006Anas El Isbihani, Shahram Khadivi +2 more
Some statistically and linguistically motivated methods for Arabic word segmentation are described and the efficiency of proposed methods on the Arabic-English BTEC and NIST tasks is shown.
Improved Arabic base phrase chunking with a new enriched POS tag set
35 Citations2007Mona Diab
A BPC system is introduced that improves over state of the art performance in BPC using a new part of speech tag (POS) set, ERTS, that reflects some of the morphological features specific to Modern Standard Arabic.
Language Resources and EvaluationMorphological Analysis and Generation of Arabic Nouns: A Morphemic Functional Approach
35 Citations2010Mohamed Altantawy, Nizar Habash +2 more
The work is directed toward the design of linker molecules which could form part of new metal-organic framework materials with enhanced affinity for CO(2) adsorption at low pressure.
Syntactic reordering for English-Arabic phrase-based machine translation
34 Citations2009Jakob Elming, Nizar Habash
A pre-translation syntactic reordering approach developed on a close language pair (English-Danish) to the distant language pair, English-Arabic, proves the viability of this approach for distant languages.
Towards a Multi-Representational Treebank
34 Citations2008Fei Xia, Owen Rambow +3 more
This paper shows that high-quality DS-to-PS conversion is possible if the conversion process is performed at the designing stage of treebank construction to ensure that all information the authors wish to represent in PS is provided in DS.
Improvements in BBN's HMM-Based Offline Arabic Handwriting Recognition System
33 Citations2009Shirin Saleem, Huaigu Cao +4 more
A novel integration of structural features in the HMM framework which exclusively results in a 9% relative improvement in performance is proposed, and a relative reduction of 17% in word error rate over the baseline Arabic handwriting recognition system is demonstrated.
Cambridge University Press eBooksA Student Grammar of Modern Standard Arabic
32 Citations2001Eckehard Schulz
Syntactic phrase reordering for English-to-Arabic statistical machine translation
31 Citations2009Ibrahim Badr, Rabih Zbib +1 more
The effect of combining reordering with Arabic morphological segmentation, a preprocessing technique that has been shown to improve Arabic-English and English-Arabic translation, is studied.
Machine TranslationThe impact of Arabic morphological segmentation on broad-coverage English-to-Arabic statistical machine translation
29 Citations2011Hassan Al-Haj, Alon Lavie
An in-depth analysis on the effect of segmentation choices on the components of a PBSMT system reveals that text fragmentation has a negative effect on the perplexity of the language models and that aggressive segmentation can significantly increase the size of the phrase table and the uncertainty in choosing the candidate translation phrases during decoding.
Towards Automatic Spell Checking for Arabic
29 Citations2003Khaled Shaalan, Amin Allam +1 more
The approach is heuristic and involves developing an Arabic morphological analyzer, techniques of spelling checking and spelling correction, and efficient methods of lexicon operations, which is able to recognize common spelling errors for standard Arabic and Egyptian dialects.
Using shallow syntax information to improve word alignment and reordering for SMT
28 Citations2008Josep Crego, Nizar Habash
Two methods to improve SMT accuracy using shallow syntax information are described, which use chunks to refine the set of word alignments typically used as a starting point in SMT systems and extend an N-gram-based SMT system with chunk tags to better account for long-distance reorderings.
The architecture of a standard Arabic lexical database
26 Citations2004Ramzi Abbès, Joseph Dichy +1 more
This work aims at giving a first answer to the question of the ratios between the number of lemma-entries and inflected word-forms that can be expected to be included in, or generated by, a Standard Arabic lexical dB.
An integrated approach for Arabic-English named entity translation
26 Citations2005Hany Hassan, Jeffrey Sorensen
An integrated approach for named entity translation deploying phrase- based translation, word-based translation, and transliteration modules into a single framework that can be applied to NE translation for any language pair is introduced.
Lecture notes in computer scienceEfficient Automatic Correction of Misspelled Arabic Words Based on Contextual Information
25 Citations2003Chiraz Ben Othmane Zribi, Mohammed Ben Ahmed
A new method aiming to reduce the number of proposals given by automatic Arabic spelling correction tools by using the use of error’s context in order to eliminate some correction candidates is addressed.
Functional morphology
23 Citations2004Markus Forsberg, Aarne Ranta
Finite functions over hereditarily finite algebraic datatypes are used to implement natural language morphology in the functional language Haskell to make it easy for linguists, who are not trained as functional programmers, to apply the ideas to new languages.
Localization of difficult-to-translate phrases
23 Citations2007Behrang Mohit, Rebecca Hwa
It is verified that by isolating difficult-to-translate phrases and processing them as special cases, their negative impact on the translation of the rest of the sentences can be reduced.
…
