Critical tokenization and its properties
Published 1 December 1997Open access
Jin Guo
Citations56
SJR quartileQ4
SJR score0.11
SNIP0.06
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
It is believed that critical tokenization provides a precise mathematical description of the principle of maximum tokenization, and forms the sound mathematical foundation for categorizing tokenization ambiguity into critical and hidden types.
Abstract
This paper sets out to study critical tokenization, a distinctive type of tokenization following the principle of maximum tokenization. The objective in this paper is to develop its mathematical description and understanding
Keywords
Computer ScienceArts and Humanities
Compilers: Principles, Techniques, and Tools
8,137 Citations1986Alfred V. Aho, Ravi Sethi +1 more
This book discusses the design of a Code Generator, the role of the Lexical Analyzer, and other topics related to code generation and optimization.
The Philosophical ReviewGeneralized Phrase Structure Grammar.
1,843 Citations1989Scott Soames, Gerald Gazdar +3 more
"Generalized Phrase Structure Grammar" provides the definitive exposition of the theory of grammar originally proposed by Gerald Gazdar and developed during half a dozen years' work with his colleagues Ewan Klein, Geoffrey Pullum, and Ivan Sag.
Tokenization as the initial phase in NLP
364 Citations1992Jonathan J. Webster, Chunyu Kit
The significance and complexity of tokenization, the beginning step of NLP, is addressed and practical approaches to identification of compound tokens in English, such as idioms, phrasal verbs and fixed expressions, are developed.
Journal of PragmaticsReadings in natural language processing
312 Citations1989
The book presents papers on natural language processing, focusing on the central issues of representation, reasoning, and recognition in syntactic models, semantic interpretation, discourse interpretation, language action and intentions, language generation, and systems.
A stochastic finite-state word-segmentation algorithm for Chinese
290 Citations1994Richard Sproat, William A. Gale +2 more
A stochastic finite-state model is presented for segmenting Chinese text into dictionary entries and productively derived words, and providing pronunciations for these words; the method incorporates a class-based model in its treatment of personal names.
Word identification for Mandarin Chinese sentences
183 Citations1992Keh-Jiann Chen, Shing-Huan Liu
This paper proposes a matching algorithm with 6 different heuristic rules to resolve the ambiguities of Chinese sentences and supports that the maximal matching algorithm is the most effective heuristics.
Syntactic graphs: a representation for the union of all ambiguous parse trees
27 Citations1989Jungyun Seo, Robert F. Simmons
It is claimed that a syntactic graph carries complete syntactic information provided by a parse forest---the set of all possible parse trees.
Analysis of Japanese compound nouns using collocational information
24 Citations1994Yosiyuki Kobayasi, Takenobu Tokunaga +1 more
This paper proposes a method to analyze structures of Japanese compound nouns by using both word collocations statistics and a thesaurus, which is about 80% accurate.
The COCOON platform (University of Paris)A statistically emergent approach for language processing: application to modeling context effects in ambiguous Chinese word boundary perception
24 Citations1996Kok-Wee Gan, Kim-Teng Lua +1 more
It is proposed that the process of language understanding can be modeled as a collective phenomenon that emerges from a myriad of microscopic and diverse activities, analogous to the crystallization process in chemistry.
Waseda University Repository (Waseda University)Ambiguity Resolution in Chinese Word Segmentation
16 Citations1995Maosong Sun, Benjamin K. Tsou
The problems in this approach are (i) the constructions of 0As in texts to be processed are nearly unpredictable, resulting in unwieldy complexity in rule base establislunent and maintenance, and (ii) it always fails in finding CAs.
Broad coverage automatic morphological segmentation of German words
6 Citations1992Thomas Pachunke, Oliver Mertineit +2 more
A system for the automatic segmentation of German words into morphs was developed and statistical evaluations showed that more than 99% of the segmented words got a correct segmentation.
Corpus-based speech and language research in the Institute of Systems Science
1 Citations2002Horng Jyh Paul Wu, Jin Guo +2 more
This paper describes the ongoing and planned research projects on speech and language modeling in the Institute of Systems Science and reveals two related characteristics of a practical natural language processing (NLP) system emerge as rather crucial: to prepare a high quality and large amount of tagged corpora as training examples.
