login

Indexing with WordNet synsets can improve text retrieval.

Published 1 August 1998
Julio Gonzalo, Felisa Verdejo, Irina Chugur, Juan Cigarrán
Citations208

TL;DR

The classical, vector space model for text retrieval is shown to give better results if WordNet synsets are chosen as the indexing space, instead of word forms, if queries are not disambiguated.

Abstract

The classical, vector space model for text retrieval is shown to give better results (up to 29% better in our experiments) if WordNet synsets are chosen as the indexing space, instead of word forms. This result is obtained for a manually disambiguated test collection (of queries and documents) derived from the Semcor semantic concordance. The sensitivity of retrieval performance to (automatic) disambiguation errors when indexing documents is also measured. Finally, it is observed that if queries are not disambiguated, indexing by synsets performs (at best) only as good as standard word indexing. 1 Introduction Text retrieval deals with the problem of finding all the relevant documents in a text collection for a given user's query. A large-scale semantic database such as WordNet (Miller, 1990) seems to have a great potential for this task. There are, at least, two obvious reasons: ffl It offers the possibility to discriminate word senses in documents and queries. This would prevent m...

Keywords

Computer Science