Automatic Language-Specific Stemming in Information Retrieval
Lecture notes in computer sciencePublished 1 January 2001
John Goldsmith, Derrick Higgins, Svetlana Soglasnova
Citations33
SJR quartileQ2
SJR score0.35
SNIP0.55
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
Automorphology, an MDL-based algorithm that determines the suffixes present in a language-sample with no prior knowledge of the language in question, is employed, employing this stemmer in a SMARTbased IR engine.
Abstract
We employ Automorphology, an MDL-based algorithm that determines the suffixes present in a language-sample with no prior knowledge of the language in question, and describe our experiments on the usefulness of this approach for Information Retrieval, employing this stemmer in a SMARTbased IR engine.
Keywords
Computer Science
Foundations of statistical natural language processing
9,996 Citations1999Christopher D. Manning, Hinrich Schütze
Program electronic library and information systemsAn algorithm for suffix stripping
8,136 Citations1980Martin Porter
An algorithm for suffix stripping is described, which has been implemented as a short, fast program in BCPL, and performs slightly better than a much more elaborate system with which it has been compared.
Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition
4,165 Citations2000Daniel Jurafsky, James Martin
This book takes an empirical approach to language processing, based on applying statistical and other machine-learning algorithms to large corpora, to demonstrate how the same algorithm can be used for speech recognition and word-sense disambiguation.
Speech and language processing
3,589 Citations2010Dan Jurafsky, James Martin +1 more
It is now clear that HAL’s creator, Arthur C. Clarke, was a little optimistic in predicting when an artificial agent such as HAL would be avail-able.
Information Retrieval: Data Structures and Algorithms
2,428 Citations1992William B. Frakes, Ricardo Baeza‐Yates
For programmers and students interested in parsing text, automated indexing, its the first collection in book form of the basic data structures and algorithms that are critical to the storage and retrieval of documents.
Viewing morphology as an inference process
661 Citations1993Robert Krovetz
The role of morphological analysis in word sense disambiguation, and in identifying lexical semantic relationships in a machine-readable dictionary, is described.
Bulletin of the London Mathematical SocietySTOCHASTIC COMPLEXITY IN STATISTICAL INQUIRY
651 Citations1991A. P. Dawid
Journal of the American Society for Information ScienceHow effective is suffixing?
388 Citations1991Donna Harman
Three measures were selected which evaluate performance at given rank cutoff points, such as those cor- responding to a screenful of document titles,such as those responding to the lists of the top ranked documents.
Journal of the American Society for Information ScienceStemming algorithms: A case study for detailed evaluation
371 Citations1996David A. Hull
A case study of stemming algorithms is described which describes a number of novel approaches to evaluation and demonstrates their value.
Text, speech and language technologyNatural Language Information Retrieval
324 Citations1999Tomek Strzalkowski
ACM SIGIR ForumAnother stemmer
318 Citations1990Chris D. Paice
In natural language processing, conflation is the process of merging or lumping together nonidentical words which refer to the same principal concept.
ACM Transactions on Information SystemsCorpus-based stemming using cooccurrence of word variants
303 Citations1998Jinxi Xu, W. Bruce Croft
This work proposes a technique for using corpus-based word variant cooccurrence statistics to modify or create a stemmer, and demonstrates the viability of this technique and its advantages relative to conventional approaches that only employ morphological rules.
Information Retrieval Systems: Theory and Implementation
289 Citations1997Gerald Kowalski
This chapter discusses information processing systems in detail, including information system evaluation, cataloging and indexing, and the role of text search algorithms in this system.
Encyclopedia of Machine Learning and Data MiningCross-Language Information Retrieval
246 Citations2017Claude Sammut, Geoffrey I. Webb
This chapter reviews research and practice in CLIR that allows users to state queries in their native language and retrieve documents in any other language supported by the system.
Cross-Language Information Retrieval
226 Citations1998Gregory Grefenstette
This work focuses on the development of a model for automatic Cross-Language Information Retrieval using Latent Semantic Indexing and its application to Machine Translation Technology.
Information Storage and RetrievalWord segmentation by letter successor varieties
206 Citations1974Margaret A. Hafer, Stephen F. Weiss
Results show that this method for automatically segmenting words into their stems and affixes is capable of high quality word segmentation, and that its use in information retrieval produces results that are at least as good as those obtained using the more traditional stemming processes.
Journal of the American Society for Information ScienceThe effectiveness of stemming for natural-language access to Slovene textual data
167 Citations1992Mirko Popovič, Peter Willett
The use of stemming on Slovene-language documents and queries is reported, and it is demonstrated that the use of an appropriate stemming algorithm results in a large, and statistically significant, increase in retrieval effectiveness when compared with nonconflated processing.
Viewing stemming as recall enhancement
163 Citations1996Wessel Kraaij, Renée Pohlmann
Results show that linguistic stemming restricted to inflection yields a significant improvement over full linguistic and non-linquistic stemming, both in average Precision and R-Recall.
Information Storage and RetrievalThe use of an association measure based on character structure to identify semantically related pairs of words and document titles
134 Citations1974George W. Adamson, Jillian Boreham
Dice's Similarity Coefficient is computed from the number of matching digrams in pairs of character strings, and used to cluster sets of characterstrings, which successfully clustered into groups of semantically related words.
Journal of Information ScienceAn evaluation of some conflation algorithms for information retrieval
118 Citations1981Martin Lennon, David S. Peirce +2 more
Comparative experiments with a range of keyword dictionaries and with the Cranfield document test collection suggest that there is relatively little difference in the performance of conflation algorithms despite the widely disparate means by which they have been developed and byWhich they operate.
Text, speech and language technologyWhat is the Role of NLP in Text Retrieval?
91 Citations1999Karen Spärck Jones
It is concluded that LMI is not needed for effective retrieval, but has other important roles within information-selection systems.
Text, speech and language technologyNLP for Term Variant Extraction: Synergy Between Morphology, Lexicon, and Syntax
89 Citations1999Christian Jacquemin, Evelyne Tzoukermann
A natural language processing (NLP) approach to automatic indexing over controlled vocabulary which accounts for term variation is presented, applied to the French language.
Journal of the American Society for Information ScienceMethod for evaluation of stemming algorithms based on error counting
85 Citations1996Chris D. Paice
Performance is assessed by counting the number of identifiable errors during the stemming of words from various text samples and it appears that the Lovins stemmer is inferior to the other two in terms of general accuracy.
Natural Language EngineeringFinite state morphology and information retrieval
17 Citations1996Kimmo Koskenniemi
A source of potential systematic errors in information retrieval that occurs when base form reduction is applied with a (necessarily) finite dictionary is identified and discussed.
Multi-language text indexing for internet retrieval
17 Citations1997Martin Wechsler, Páraic Sheridan +1 more
The issues associated with indexing multilingual collections of information, as is found for example on the internet are addressed, in particular the task of language identification and the use of stemming algorithms for several European languages.
Lecture notes in computer scienceNatural Language in Information Retrieval
15 Citations2003Elżbieta Dura
The time is ripe for the two to meet: NLP has grown out of prototypes and IR is having hard time trying to improve precision, so two examples of possible approaches are considered.
