login

Aspects of Swedish morphology and semantics from the perspective of mono- and cross-language information retrieval

Information Processing & ManagementPublished 1 January 2001
Turid Hedlund, Ari Pirkola, Kalervo Järvelin
Citations47
SJR quartileQ1
SJR score2.06
SNIP2.91

TL;DR

The results suggest that part-of-speech tagging might be useful in Swedish IR due to the high frequency of homographic words, and publicly available morphological analysis tools used for normalization and compound splitting have pitfalls that might decrease the effectiveness of IR and CLIR.

Abstract

This paper analyzes the features of the Swedish language from the viewpoint of mono- and cross-language information retrieval (CLIR). The study was motivated by the fact that Swedish is known poorly from the IR perspective. This paper shows that Swedish has unique features, in particular gender features, the use of fogemorphemes in the formation of compound words, and a high frequency of homographic words. Especially in dictionary-based CLIR, correct word normalization and compound splitting are essential. It was shown in this study, however, that publicly available morphological analysis tools used for normalization and compound splitting have pitfalls that might decrease the effectiveness of IR and CLIR. A comparative study was performed to test the degree of lexical ambiguity in Swedish, Finnish and English. The results suggest that part-of-speech tagging might be useful in Swedish IR due to the high frequency of homographic words.

Keywords

Computer Science