Syntactic Annotation in Columbia Arabic Treebank
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
CATiB uses linguistic representation and terminology inspired by the long tradition of Arabic syntactic studies to make it easier to train annotators and not be restricted to hire annotators who have degrees in linguistics.
Abstract
The Columbia Arabic Treebank (CATiB) is a database of syntactic analyses of Arabic sentences. CATiB contrasts with previous ap-proaches to Arabic treebanking in its emphasis on faster production with some constraints on linguistic richness. Two basic ideas inspire the CATiB approach. First, CATiB avoids the annotation of redundant linguistic information that is determinable automatically from syntax and morphological analysis, e.g., nominal case. And secondly, CATiB uses linguistic representation and terminology inspired by the long tradition of Arabic syntactic studies. This makes it easier to train annotators and not be restricted to hire annotators who have degrees in linguistics. This paper describes CATiB’s representation and compares it to other Arabic treebanking efforts. 1.
