Natural language processing
Annual Review of Information Science and TechnologyPublished 1 January 2003Open access
Gobinda Chowdhury
Citations777
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This chapter presents the challenges of NLP, progress so far made in this field, NLP applications, components of N LP, and grammar of English language—the way machine requires it.
Abstract
conducted domain-specific NLP studies
Keywords
Computer Science
Foundations of statistical natural language processing
9,996 Citations1999Christopher D. Manning, Hinrich Schütze
Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition
4,165 Citations2000Daniel Jurafsky, James Martin
This book takes an empirical approach to language processing, based on applying statistical and other machine-learning algorithms to large corpora, to demonstrate how the same algorithm can be used for speech recognition and word-sense disambiguation.
A re-examination of text categorization methods
2,660 Citations1999Yiming Yang, Xin Liu
The results show that SVM, kNN and LLSF signi cantly outperform NNet and NB when the number of positive training instances per category are small, and that all the methods perform comparably when the categories are over 300 instances.
Journal of the American Statistical AssociationStatistical Methods for Speech Recognition
1,988 Citations1999Don X. Sun, Frederick Jelinek
The speech recognition problem hidden Markov models the acoustic model basic language modelling the Viterbi search hypothesis search on a tree and the fast match elements of information theory.
Introduction to Modern Information Retrieval
921 Citations1999Gobinda Chowdhury
It is this book's contention that it also benefits information professionals to learn the theory, techniques and tools that constitute the traditional approaches to the organization and processing of information.
The TREC-8 Question Answering Track Report
857 Citations1999Ellen M. Voorhees
Proceedings of the IEEETwo decades of statistical language modeling: where do we go from here?
726 Citations2000Roni Rosenfeld
A Bayesian approach to integration of linguistic theories with data is argued for inStatistical language models estimate the distribution of various natural language phenomena for the purpose of speech recognition and other language technologies.
Communications of the ACMInformation extraction
726 Citations1996Jim Cowie, Wendy G. Lehnert
A relatively new development—information extraction (IE)—is the subject of this article and can transform the raw material, refining and reducing it to a germ of the original text.
Data Mining and Knowledge Discovery
690 Citations1998Krzysztof J. Cios, Witold Pedrycz +1 more
Summarizing text documents
464 Citations1999Jade Goldstein, Mark Kantrowitz +2 more
An analysis of news-article summaries generated by sentence selection, using a normalized version of precision-recall curves with a baseline of random sentence selection to evaluate features and empirical results show the importance of corpus-dependent baseline summarization standards, compression ratios and carefully crafted long queries.
Supertagging: an approach to almost parsing
393 Citations1999Srinivas Bangalore, Aravind K. Joshi
Novel methods for robust parsing that integrate the flexibility of linguistically motivated lexical descriptions with the robustness of statistical techniques are proposed.
Text, speech and language technologyNatural Language Information Retrieval
324 Citations1999Tomek Strzalkowski
arXiv (Cornell University)A Corpus-Based Investigation of Definite Description Use
322 Citations1997Massimo Poesio, Renata Vieira
Communications of the ACMNatural language processing for information retrieval
281 Citations1996David Lewis, Karen Spärck Jones
The paper considers the new opportunities and challenges presented by the user’s ability to search full text directly (rather than e.g. titles and abstracts), and suggests appropriate approaches to doing this, with a focus on the potential role of natural language processing.
Encyclopedia of Machine Learning and Data MiningCross-Language Information Retrieval
246 Citations2017Claude Sammut, Geoffrey I. Webb
This chapter reviews research and practice in CLIR that allows users to state queries in their native language and retrieve documents in any other language supported by the system.
Lecture notes in computer scienceUsing Noun Phrase Heads to Extract Document Keyphrases
236 Citations2000Ken Barker, Nadia Cornacchia
The simple noun phrase-based system performs roughly as well as a state-of-the-art, corpus-trained keyphrase extractor; ratings for individual keyphrases do not necessarily correlate with ratings for sets of keyphRases for a document.
Cross-Language Information Retrieval
226 Citations1998Gregory Grefenstette
This work focuses on the development of a model for automatic Cross-Language Information Retrieval using Latent Semantic Indexing and its application to Machine Translation Technology.
Journal of DocumentationInformation extraction: beyond document retrieval
213 Citations1998Robert Gaizauskas, Yorick Wilks
A synoptic view of the growth of the text processing technology of information extraction whose function is to extract information about a pre‐specified set of entities, relations or events from natural language texts and to record this information in structured representations called templates is given.
Innovative Applications of Artificial IntelligenceCONSTRUE/TIS: A System for Content-Based Indexing of a Database of News Stories
162 Citations1990Philip J. Hayes, Steven P. Weinstein
The Construe news story categorization system assigns indexing terms to news stories according to their content using knowledge-based techniques and Reuters expects the speed and consistency of TIS to provide significant competitive advantage and, hence, an increased market share for Country Reports and other products from Reuters Historical Information Products Division.
Extracting sentence segments for text summarization
144 Citations2000Wesley T. Chuang, Jihoon Yang
An approach to the design of an automatic text summarizer that generates a summary by extracting sentence segments that compares very favorably with other approaches in terms of precision, recall, and classification accuracy.
Journal of the American Society for Information ScienceComparing noun phrasing techniques for use with medical digital library tools
132 Citations2000Kristin M. Tolle, Hsinchun Chen
Noun phrasing tools were evaluated as to their ability to isolate noun phrases from medical journal abstracts and it was shown that augmenting the AZ Noun Phraser by including the SPECIALIST Lexicon from the National Library of Medicine resulted in improved recall and precision.
Current theories of centering for pronoun interpretation: a critical evaluation
99 Citations1997Andrew Kehler
The fundamental concepts of centering theory are reviewed and some facets of the pronoun interpretation problem that motivate a centering-style analysis are discussed, as well as some problems with a popular Centering-based approach.
Text, speech and language technologyWhat is the Role of NLP in Text Retrieval?
91 Citations1999Karen Spärck Jones
It is concluded that LMI is not needed for effective retrieval, but has other important roles within information-selection systems.
Bulletin of the American Society for Information Science and TechnologyEnhanced Text Retrieval Using Natural Language Processing
87 Citations1998Elizabeth D. Liddy
Quelques ameliorations basees sur l'emploi du traitement du langage naturel ont atteint finalement les moteurs de recherche du commmerce, mais il n'y a pas d'usage normalise en terminologie pour decrire leur processus.
Text, speech and language technologyUsing NLP or NLP Resources for Information Retrieval Tasks
83 Citations1999Alan F. Smeaton
This chapter presents a precis of the experiments in information retrieval using NLP which have had mixed success over the last few years and introduces the respective roles of NLP and IR and summarises their early experiments on using syntactic analysis to derive term dependencies and structured representations of term-term relationships.
Applying data mining techniques for descriptive phrase extraction in digital document collections
81 Citations2002Helena Ahonen-Myka, O. P. Heinonen +2 more
This paper shows that general data mining methods are applicable to text analysis tasks such as descriptive phrase extraction and presents a general framework for text mining, based on generalized episodes and episode rules.
A memory-based approach to learning shallow natural language patterns
81 Citations1998Shlomo Argamon, Ido Dagan +1 more
A novel memory-based learning method that recognizes shallow patterns in new text based on a bracketed training corpus that enables easy porting to new domains and to sub-language patterns for information extraction.
Journal of DocumentationMorphological typology of languages for IR
80 Citations2001Ari Pirkola
The paper elaborates the linguistic morphological typology for the purposes of IR research and studies how the indexes of synthesis and fusion could be used as practical tools in mono‐ and cross‐lingual IR research.
Efficient text summarization using lexical chains
78 Citations2000H. Gregory Silber, Kathleen F. McCoy
This research presents a linear time algorithm for calculating lexical chains which is a method of capturing the “aboutness” of a document which is compared to previous, less efficient methods of lexical chain extraction.
Journal of the American Society for Information ScienceComparing noun phrasing techniques for use with medical digital library tools
77 Citations2000Kristin M. Tolle, Hsinchun Chen
Communications of the ACMEvaluating natural language processing systems
76 Citations1996Margaret King
Evaluating Natural Language Processing Systems Designing customized methods for testing various NLP systems may be costly and expensive, so post hoc justification is needed.
Question answering in TREC
66 Citations2001Ellen M. Voorhees
A brief summary of the findings of the TREC question answering track to date is provided and the future directions of the track are discussed.
Information Processing & ManagementAutomatic Identification and Back-Transliteration of Foreign Words for Information Retrieval
65 Citations1999Kil Soon Jeong, Sung Hyon Myaeng +2 more
An algorithm is developed that first identifies a phrase containing a foreign word and then extracts the foreign word part from the phrase based on statistical information and the method for back-transliteration of a foreignword to its English origin is presented.
Digital Commons - USU (Utah State University)Cross-Language Information Retrieval
64 Citations1998Oard, Douglas W., Diekama, Anne R.
Knowledge lean word sense disambiguation
60 Citations1997Ted Pedersen
A corpus-based approach to word-sense disambiguation that only requires information that can be automatically extracted from untagged text and Gibbs Sampling results in small but consistent improvement in disambigsuation accuracy over the EM algorithm.
Information Processing & ManagementNatural language information retrieval: progress report
59 Citations2000José Pérez-Carballo, Tomek Strzalkowski
The ‘stream architecture’ is described, a method designed to combine evidence obtained from several different document representations that involved the use of phrases and proper names computed using Natural Language Processing techniques.
Engineering Applications of Artificial IntelligenceInnovative applications of artificial intelligence
56 Citations1990Chao Yan
Extended finite state models of language
53 Citations1999András Kornai
D-Lib MagazineMultilingual Federated Searching Across Heterogeneous Collections
52 Citations1998James Powell, Edward A. Fox
A scalable system for searching heterogeneous multilingual collections on the World Wide Web is described, including a markup language for describing the characteristics of a search engine and its interface, and a protocol for requesting word translations between languages.
Lecture notes in computer scienceGETESS—Searching the Web Exploiting German Texts
52 Citations1999Steffen Staab, Christian Braun +11 more
An intelligent information agent is designed such that as background knowledge and linguistic coverage increase, its benefits improve, while it guarantees state-of-the-art information and database retrieval capabilities as its bottom line.
Communications of the ACMNatural language dialogue for personalized interaction
51 Citations2000Wlodek Zadrozny, Margo Budzikowska +4 more
NL research attempts to define extensive discourse models that in turn provide improved models of context-enabling HCI and personalization, which are key to true personalization.
Artificial Intelligence in MedicineMedical dictionaries for patient encoding systems: a methodology
47 Citations1998Christian Lovis, Robert Baud +3 more
The main aim of the proposed approach is that of coping with 'the lack of coverage of the medical lexical knowledge', in order to help physicians find the correct international classification for diseases (ICD) codes for a written diagnosis.
Information Processing & ManagementAspects of Swedish morphology and semantics from the perspective of mono- and cross-language information retrieval
47 Citations2001Turid Hedlund, Ari Pirkola +1 more
The results suggest that part-of-speech tagging might be useful in Swedish IR due to the high frequency of homographic words, and publicly available morphological analysis tools used for normalization and compound splitting have pitfalls that might decrease the effectiveness of IR and CLIR.
Information Processing & ManagementUsing cause-effect relations in text to improve information retrieval precision
47 Citations2001Christopher S. G. Khoo, Sung Hyon Myaeng +1 more
The best kind of causal relation matching was found to be one in which one member of the causal relation was represented as a wildcard that could match with any word.
Computer Assisted Language LearningMorphological Processing and Computer-Assisted Language Learning
43 Citations1998John Nerbonne, Duco Dokter +1 more
The position of NLP within CALL is discussed using GLOSSER, an intelligent assistant for Dutch students learning to read in French, which relies essentially on lemmatization, part-of-speech (POS) disambiguation, lexeme indexing, and...
Machine LearningA Machine Learning Approach to POS Tagging
43 Citations2000Lluı́s Màrquez, Lluís Padró +1 more
This work describes a tagger which is able to use information of any kind, and in particular to incorporate the machine-learned decision trees, and addresses the problem of tagging when only limited training material is available, which is crucial in any process of constructing, from scratch, an annotated corpus.
Information Processing & ManagementInformation filtering via hill climbing, WordNet, and index patterns
40 Citations1997Kenrick Mock, V. Rao Vemuri
This paper describes work implemented in the INFOS (Intelligent News Filtering Organizational System) project that is designed to reduce the user's search burden by automatically categorizing data as relevant or irrelevant based upon user interests.
Journal of the American Society for Information ScienceCross-language information access to multilingual collections on the internet
39 Citations2000Guo-Wei Bian, Hsin‐Hsi Chen
This paper deals with query translation and document translation in a Chinese‐English information retrieval system called MTIR, where bilingual dictionary and monolingual corpus‐based approaches are adopted to select suitable translated query terms and a machine transliteration algorithm is introduced to resolve proper name searching.
Lecture notes in computer scienceInformation retrieval: Still butting heads with natural language processing?
39 Citations1997Alan F. Smeaton
It is of interest to the IE community to see how a related task, perhaps the most-related task, IR, has managed to use the same NLP base technology in its development so far.
TREC-9 Cross Language, Web and Question-Answering Track Experiments Using PIRCS
39 Citations2006K. L. Kwok, Laszlo Grunfeld +2 more
Evaluation shows that English-Chinese cross-lingual retrieval using only wordlist query translation can achieve about 70-75% of monolingual average precision, and combination with MT query translation further brings this effectiveness to 80-85% ofmonolingual.
Topical clustering of MRD senses based on information retrieval techniques
39 Citations1998Jen Nan Chen, Jason S. Chang
A heuristic approach capable of automatically clustering senses in a machine-readable dictionary (MRD) and an implementation of the method for clustering definition sentences in the Longman Dictionary of Contemporary English (LDOCE) is described.
Computers and the HumanitiesCross-Lingual Sense Determination: Can It Work?
37 Citations2000Nancy Ide
Communications of the ACMNatural language processing
30 Citations1996Yorick Wilks
T h i s has changed and it seems c l e a r t h a t l anguage p r o c e s s i n g i s a c e n t r a l t o p i c w i t h i n A I .
Journal of the American Society for Information ScienceCross-language information access to multilingual collections on the internet
25 Citations2000Guo-Wei Bian, Hsin‐Hsi Chen
L. Erlbaum Associates Inc. eBooksIntelligent high-volume text processing using shallow, domain-specific techniques
25 Citations1992Philip J. Hayes
Information Processing & ManagementAutomatic text structuring and categorization as a first step in summarizing legal cases
24 Citations1997Marie‐Francine Moens, Caroline Uyttendaele
It is argued that prior knowledge of the text structure and its indicative cues may support automatic abstracting, and a text grammar is a promising form for representing the knowledge involved.
Illinois Digital Environment for Access to Learning and Scholarship (University of Illinois at Urbana-Champaign)Template Mining for Information Extraction from Digital Documents.
23 Citations1999Gobinda Chowdhury
This article briefly reviews template mining research and shows how templates are used in Web search engines- such as Alta Vista-and in meta-search engines-such as Ask Jeeves-for helping end-users generate natural language search expressions.
Information Processing & ManagementAutomatic association of news items
22 Citations1997Christina Carrick, Carolyn Watters
The specific association of text and photo news items is discussed although the approach to the problem of defining relationships between photos and stories applies to a larger domain of news items including scripted news video clips and scripted radio broadcasts.
Question Answering Using a Large NLP System.
21 Citations2000David Elworthy
The Microsoft Research question-answering system for TREC-9 was based on a combination of the Okapi retrieval engine, Microsoft’s natural language processing system (NLPWin), and a module for matching logical forms.
European Journal of Applied PhysiologyQuestion Answering from Large Document Collections
21 Citations1999Eric Breck, John D. Burger +1 more
A rehabilitation exercise program to remediate acute atrophy in females appears more effective if E2 is present.
Natural language processing
19 Citations1987Roger C. Schank, Alex Kass
It is only possible for research on understanding natural language to make progress when the researchers realize that the heart of NLU is the understanding process, not language per se.
Proceedings of the 17th international conference on Computational linguistics -A memory-based approach to learning shallow natural language patterns
19 Citations1998Shlomo Argamon, Ido Dagan +1 more
Expert Systems with ApplicationsThe role of knowledge-based technology in language applications development
19 Citations2000Paloma Martínez Fernández
A cognitive approach is presented that allows the design of linguistic applications that integrates different formalisms, reuses existing language resources and supports the implementation of the required control in a flexible way and shows the suitability of knowledge-based technology in linguistic engineering.
MITA: An Information-Extraction Approach to the Analysis of Free-Form Text in Life Insurance Applications
19 Citations1997Barry Glasgow, Alan Mandell +3 more
MITA, MetLife's Intelligent Text Analyzer, uses the Infonnation Extraction technique of Natural Language Processing to structure the extensive text fields on a life insurance application to increase underwriting productivity by 20 to 30%.
Information Processing & ManagementFreestyle vs. Boolean: A comparison of partial and exact match retrieval systems
17 Citations1998Lee Anne H. Paris, Helen R. Tibbo
Although Boolean searching has been the standard model for commercial information retrieval systems for the past three decades, natural language input and partial-match weighted retrieval have recently emerged from the laboratories to become a searching option in several well-known online systems.
Information Processing & ManagementText structuration leading to an automatic summary system: RAFI
17 Citations1999Abderrafih Lehmam
A system automatically and directly transforming a source text into a reduced target text, based on the identification of specific expressions allowing an evaluation of the relevance of the sentence concerned, which can then be selected for the elaboration of the summary.
The use of phrases from query texts in information retrieval (poster session)
15 Citations2000Masumi Narita, Yasushi Ogawa
This paper focuses on linguistically motivated phrases as extracted from query texts by natural language processing, and uses shallow syntactic processing instead of statistical processing to automatically identify candidate phrasal terms from querytexts.
ACM Computing SurveysNatural language learning
14 Citations1995Eugene Charniak
The branch of artificial intelligence dealing with natural language processing (NLP) has undergone a quiet revolution and AI NLP work could be characterized as knowledge-based in its orientation and hand-tooled in its methodology.
International Journal of Human-Computer StudiesNatural language querying of databases: an information extraction approach in the conceptual query language
14 Citations2000Vesper Owei
The conceptual query language-with-natural language (CQL/NL) is proposed, which uses information extraction methods to filter NL query statements for search predicates that are derived from constructs on conceptual schemas.
Journal of the American Society for Information ScienceAbstracts produced using computer assistance
12 Citations2000Timothy C. Craven
Experimental subjects wrote abstracts of articles using a simplified version of the TEXNET abstracting assistance software, and showed considerable variation among subjects, but 37% found the keywords or phrases “quite” or “very” useful in writing their abstracts.
Artificial Intelligence and LawNatural language processing for transparent communication between public administration and citizens
11 Citations2000Bernardo Magnini, Elena Not +2 more
The first project, GIST, is concerned with automatic multilingual generation of instructional texts for form-filling and TAMIC aims at providing an interface for interactive access to information, centered on natural language processing and supposed to be used by the clerk but with the active participation of the citizen.
ETRI JournalAn Algorithm for Predicting the Relation between Lemmas and Corpus Size
11 Citations2000Dan-Hee Yang, Pascual Cantos Gómez +1 more
This study shall reveal the flaws of several previous researches aiming to predict corpus size, especially those using pure regression or curve‐fitting methods, and contrive a new mathematical tool: a piecewise curve‐ fitting algorithm.
Information Processing & ManagementFromTo-CLIR: Web-Based Natural Language Interface for Cross-Language Information Retrieval
11 Citations1999Taewan Kim, Chul-Min Sim +6 more
This paper proposes a method that uses a semantic category tree and collocation to resolve the ambiguity of query translation and uses a hybrid translation engine that combines a pattern-based module with a rule-based translator and includes pre- and post-fail softeners.
Journal of the American Society for Information ScienceAbstracts produced using computer assistance
10 Citations2000Timothy C. Craven
Journal of DocumentationKnowledge discovery from databases: an introductory review
10 Citations1997Brian P. Vickery
The paper aims to provide a non‐technical introduction to the new procedures being used to extract knowledge from databases by describing such techniques as classification, clustering, and the detection of deviations from pre‐established norms.
Experiences with a multilingual ontology-based lexicon for news filtering
9 Citations2002H. Weigard, Stijn Hoppenbrouwers
The way the lexicon, including an ontology, is constructed is sketched and the methodology used for adding domain ontologies is presented, which is lexicon-driven in the sense that ontology and lexicon are developed in tandem.
Journal of the American Society for Information ScienceText segmentation for Chinese spell checking
8 Citations1999Kin Hong Lee, Mau Kit Michael Ng +1 more
A Block-of-Combinations (BOC) segmentation method based on frequency of word usage is proposed to reduce the word combinations from exponential growth to linear growth and to make the segmentation more suitable for spell checking, user interaction is suggested.
The extraction method of the word meaning class
7 Citations2003Kenichi Tsuda, Manabu Nakamura
This paper proposes a method of extracting the semantic class information on a word from a set of documents by using the characteristic that the frequency of abstract words is high while thefrequency of concrete words is small.
Computer Assisted Language LearningLinguistic Games for Language Learning: A Special Use of the ILLICO Library
7 Citations1998Robert Pasero, Paul Sabatier
The principles underlying ILLICO, a generic natural language software tool for building larger applications for performing specific linguistic tasks such as analysis, synthesis and guided composition, are described.
Speech Synthesis and Speech Recognition: Tomorrow's Human-Computer Interfaces?.
7 Citations1993Holley R. Lange
This chapter provides an introduction to the subject and discusses speech synthesis and speech recognition: history, current work, human factors, and applications; and looks to future use and development for these technologies.
Information Processing & ManagementGlean: using syntactic information in document filtering
7 Citations1998Raman Chandrasekar, Srinivas Bachu
It is shown that syntactic information does improve the effectiveness of filtering irrelevant documents, and that supertagging is more effective than part of speech tagging in filtering documents.
Computer Assisted Language LearningA Discourse Structure Analysis of Technical Japanese Texts and Its Implementation on the WWW
7 Citations2000Jie Chi Yang, Kanji Akahori
A CALL system that can be used for automatically detecting headlines and cohesive expressions of technical Japanese texts on any World Wide Web (WWW) browser is described, which can be considered as a new means of language learning for the future.
Lecture notes in computer scienceAutomatic Acquisition of Morphological Knowledge for Medical Language Processing
7 Citations1999Pierre Zweigenbaum, Natalia Grabar
A simple and powerful method to acquire automatically morphologically related words for medical language processing application that takes advantage of commonly available lists of synonym terms to bootstrap the acquisition process.
IFIP advances in information and communication technologyA knowledge-based methodology applied to linguistic engineering
7 Citations1998Paloma Martı́nez, Ana Garcı́a-Serrano
A methodological approach to the design of structured knowledge models for natural language processing (NLP) applications that takes the inherent interdisciplinarity of this area into account and deals with the linguistic knowledge involved in these systems.
Natural Language Processing: Toward Large-Scale, Robust Systems.
7 Citations1996Stephanie W. Haas
Self-organizing semantic maps of Japanese nouns in terms of adnominal constituents
6 Citations2000Qing Ma, K. Kanzaki +4 more
The construction of a semantic map of Japanese nouns mapped according to their adnominal constituents is described, which can be a powerful tool for supporting the analysis of the relation between head nouns and their ad Nominal constituents, an important issue in studies of Japanese pragmatics.
Lecture notes in computer scienceThe Power of the TSNLP: Lessons from a Diagnostic Evaluation of a Broad-Coverage Parser
6 Citations2000Elizabeth Scarlett, Stan Śzpakowicz
The TSNLP suite is considered as a diagnostic tool, and an alternative broader-coverage test suite of test sentences extracted from Quirk et al. is proposed.
Journal of the American Society for Information ScienceText segmentation for Chinese spell checking
6 Citations1999Kin Hong Lee, Mau Kit Michael Ng +1 more
An event-driven and ontology-based approach for the delivery and information extraction of e-mails
5 Citations2002Heng-Hsou Chang, Yau-Hwang Ko +1 more
The backpropagation (BP) learning algorithm is adopted to train the event detector and, by means of the well-trained event detector, the information extraction task can be certainly applied in wider domains.
Literary and Linguistic ComputingORCHID: building linguistic resources in Thai
5 Citations2000Hitoshi Isahara
A POS neuro tagger is described, which consists of a three-layer perceptron with elastic input, which is superior to the statistical models including the frequency model, local n-gram model, and HMM (Hidden Markov Model).
Communications of the ACMComputerizing computer science
5 Citations1998D. J. L. Smith
It is argued that computer scientists should aim to set an example, by organizing computer science resources into a standardized, online “archive of computer science knowledge,” using computer-friendly encodings to enable easy processing by computer.
Journal of DocumentationConsistency of textual expression in newspaper articles: an argument for semantically based query expansion
5 Citations2001Raija Lehtokangas, Kalervo Järvelin
It is argued that the expression inconsistency is a clear sign of a retrieval problem and that query expansion based on semantic relationships can significantly improve retrieval performance on free‐text sources.
Information RetrievalDR-LINK in TIPSTER III
4 Citations2000Elizabeth D. Liddy, Ted Diamond +1 more
Experimental results show that there is potential for improving retrieval through query-specific fusion and that analysts found the Detailed Multiple Document Summary to be extremely useful for almost every query, while the Thumbnail sketch was useful in approximately 50% of the queries.
Lecture notes in computer scienceLearning Word Segmentation Rules for Tag Prediction
4 Citations1999Dimitar Kazakov, Suresh Manandhar +1 more
The results show high correlation between the constituents generated by the segmentation rules, and the tags of the words in which they appear, thereby demonstrating the linguistic relevance of the segmentations produced by the hybrid approach.
Automatic Acquisition of Sense Tagged Corpora
4 Citations1999Rada Mihalcea, Dan Moldovan
This paper presents a method which enables the automatic acquisition of sense tagged corpora based on the information provided in WordNet, particularly the word definitions found within the glosses and the information gathered from Internet using existing search engines.
Artificial Intelligence in MedicineUnderstanding of medico-technical reports
3 Citations2000Michel Roux, V. Ledoray
The Aristotle project is to build an automatic data system that is capable of producing a semantic representation of the text in a canonical form, and the syntactic-semantic Interpreter processes one sentence at a time.
…
