login

Tokenization as the initial phase in NLP

Published 1 January 1992Open access
Jonathan J. Webster, Chunyu Kit
Citations364
View PDF

TL;DR

The significance and complexity of tokenization, the beginning step of NLP, is addressed and practical approaches to identification of compound tokens in English, such as idioms, phrasal verbs and fixed expressions, are developed.

Abstract

In this paper, the authors address the significance and complexity of tokenization, the beginning step of NLP. Notions of word and token are discussed and defined from the viewpoints of lexicography and pragmatic implementation, respectively. Automatic segmentation of Chinese words is presented as an illustration of tokenization. Practical approaches to identification of compound tokens in English, such as idioms, phrasal verbs and fixed expressions, are developed.

Keywords

Computer ScienceArts and Humanities