Building a large-scale annotated Chinese corpus
Published 1 January 2002Open access
Nianwen Xue, Fu-Dong Chiou, Martha Palmer
Citations135
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This paper addresses issues related to building a large-scale Chinese corpus and tries to answer four questions: how to speed up annotation, how to maintain high annotation quality, for what purposes is the corpus applicable, and finally what future work the authors anticipate.
Abstract
In this paper we address issues related to building a large-scale Chinese corpus. We try to answer four questions: (i) how to speed up annotation, (ii) how to maintain high annotation quality, (iii) for what purposes is the corpus applicable, and finally (iv) what future work we anticipate.
Keywords
Computer ScienceBiochemistry, Genetics and Molecular Biology
Maximum Entropy Model for Part-Of-Speech Tagging
1,276 Citations1996Adwait Ratnaparkhi
A statistical model which trains from a corpus annotated with Part Of Speech tags and assigns them to previously unseen text with state of the art accuracy and discusses the corpus consistency problems discovered during the implementation of these features.
Statistical parsing with an automatically-extracted tree adjoining grammar
163 Citations2000David Chiang
This work describes the induction of a probabilistic LTAG model from the Penn Treebank and finds that this induction method is an improvement over the EM-based method of (Hwa, 1998), and that the induced model yields results comparable to lexicalized PCFG.
Developing Guidelines and Ensuring Consistency for Chinese Text Annotation
119 Citations2000Fei Xia, Martha Palmer +7 more
This paper will address several challenges in building the corpus, namely, creating annotation guidelines, ensuring annotation accuracy and maintaining a high level of community involvement.
ScholarlyCommons (University of Pennsylvania)Automatic grammar generation from two different perspectives
73 Citations2001Fei Xia, Martha Palmer +1 more
Two systems that automatically generate grammars are built that solve two major problems in grammar development: namely, the redundancy caused by the reuse of structures in a grammar and the lack of explicit generalizations over the structures inA grammar.
Statistically-enhanced new word identification in a rule-based Chinese system
64 Citations2000Andi Wu, Zixin Jiang
This paper presents a mechanism of new word identification in Chinese text where probabilities are used to filter candidate character strings and to assign POS to the selected strings in a ruled-based system that improves parser coverage and provides a tool for the lexical acquisition of new words.
Facilitating treebank annotation using a statistical parser
29 Citations2001Fu-Dong Chiou, David Chiang +1 more
Corpora of phrase-structure-annotated text, or treebanks, are useful for supervised training of statistical models for natural language processing, as well as for corpus linguistics.
