login

Term Extraction and Automatic Indexing

Oxford University Press eBooksPublished 18 September 2012
Christian Jacquemin, Didier Bourigault
Citations97

TL;DR

Fields of activity in term-oriented NLP, including monolingual term recognition and automatic indexing, and some milestones in monolingual term recognition and automatic indexing are outlined.

Abstract

Terms are pervasive in scientific and technical documents and their identification is a crucial issue for any application dealing with the analysis, understanding, generation, or translation of such documents. In particular, the ever-growing mass of specialized documentation available on-line, in industrial and governmental archives or in digital libraries, calls for advances in terminology processing for tasks such as information retrieval, cross-language querying, indexing of multimedia documents, translation aids, document routing and summarization, etc. This article presents a new domain of research and development in natural language processing (NLP) that is concerned with the representation, acquisition, and recognition of terms. It begins with presenting the basic notions about the concept of ‘terms’, ranging from the classical view, to the recent concepts. There are two main areas of research involving terminology in NLP, which are, term acquisition and term recognition. Finally, this article presents the recent advances and prospects in term acquisition and automatic indexing.

Keywords

Computer ScienceArts and Humanities