Language Independent Named Entity Recognition Combining Morphological and Contextual Evidence.
Published 1 January 1999
Silviu Cucerzan, David Yarowsky
Citations223
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
Identifying and classifying personal, geographic, institutional or other names in a text is an important task for numerous applications. This paper describes and evaluates a language-independent bootstrapping algorithm based on iterative learning and re-estimation of contextual and morphological patterns captured in hierarchicaily smoothed trie models. The algorithm learns from unannotated text and achieves competitive performance when trained on a very short labelled name list with no other required language-specific information, tokenizers or tools.
Keywords
Computer Science
Computers and the HumanitiesA method for disambiguating word senses in a large corpus
534 Citations1992William A. Gale, Kenneth Church +1 more
The proposed method was designed to disambiguate senses that are usually associated with different topics using a Bayesian argument that has been applied successfully in related tasks such as author identification and information retrieval.
One sense per discourse
524 Citations1992William A. Gale, Kenneth Church +1 more
An experiment confirmed the hypothesis that if a polysemous word such as sentence appears two or more times in a well-written discourse, it is extremely likely that they will all share the same sense and found that the tendency to share sense in the same discourse is extremely strong.
Natural Language EngineeringDistribution of content words and phrases in text and language modelling
257 Citations1996Slava M. Katz
The derivation of models describing word distribution in text is based on a linguistic interpretation of the process of text formation, with the probabilities of word occurrence being functions of observable and linguistically meaningful text characteristics.
Learning to Tag Multilingual Texts Through Observation
44 Citations1997Scott Bennett, Chinatsu Aone
This paper describes RoboTag, an advanced prototype for a machine learningbased multilingual information extraction system, and describes a general client/server architecture used in learning from observation and presents experimental results which compare RoboTag to both human-tagged keys and to the best hand-coded rule systems.
Discrimination Decisions for 100,000-Dimensional Spaces
35 Citations1994William A. Gale, Kenneth Church +1 more
