login

Text Mining - Knowledge extraction from unstructured textual data

Studies in classification, data analysis, and knowledge organizationPublished 1 January 1998
Martin Rajman, Romaric Besançon
Citations72
SJR quartileQ4
SJR score0.13
SNIP0.15

TL;DR

This paper presents two examples of information that can be automatically extracted from text collections: probabilistic associations of key-words and prototypical document instances and the Natural Language Processing tools necessary for such extractions.

Abstract

In the general context of Knowledge Discovery, specific techniques, called Text Mining techniques, are necessary to extract information from unstructured textual data. The extracted information can then be used for the classification of the content of large textual bases. In this paper, we present two examples of information that can be automatically extracted from text collections: probabilistic associations of key-words and prototypical document instances. The Natural Language Processing (NLP) tools necessary for such extractions are also presented.

Keywords

Computer Science