login

Distributional measures of concept-distance

Published 1 January 2006Open access
Saif M. Mohammad, Graeme Hirst
Citations74
View PDF

TL;DR

This work proposes a framework to derive the distance between concepts from distributional measures of word co-occurrences, using the categories in a published thesaurus as coarse-grained concepts, and shows that the newly proposed concept-distance measures outperform traditional distributional word- distance measures in the tasks of ranking word pairs in order of semantic distance and correcting real-word spelling errors.

Abstract

We propose a framework to derive the distance between concepts from distributional measures of word co-occurrences. We use the categories in a published thesaurus as coarse-grained concepts, allowing all possible distance values to be stored in a concept--concept matrix roughly .01% the size of that created by existing measures. We show that the newly proposed concept-distance measures outperform traditional distributional word-distance measures in the tasks of (1) ranking word pairs in order of semantic distance, and (2) correcting real-word spelling errors. In the latter task, of all the WordNet-based measures, only that proposed by Jiang and Conrath outperforms the best distributional concept-distance measures.

Keywords

Computer Science