Tokenization as the initial phase in NLP
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The significance and complexity of tokenization, the beginning step of NLP, is addressed and practical approaches to identification of compound tokens in English, such as idioms, phrasal verbs and fixed expressions, are developed.
Abstract
In this paper, the authors address the significance and complexity of tokenization, the beginning step of NLP. Notions of word and token are discussed and defined from the viewpoints of lexicography and pragmatic implementation, respectively. Automatic segmentation of Chinese words is presented as an illustration of tokenization. Practical approaches to identification of compound tokens in English, such as idioms, phrasal verbs and fixed expressions, are developed.
