Development of a Dependency Treebank for Russian and its Possible Applications in NLP
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A tagging scheme designed for the Russian Treebank is described and tools used for corpus creation are presented, which gives a more fine-grained representation of syntactic phenomena.
Abstract
The paper describes a tagging scheme designed for the Russian Treebank and presents tools used for corpus creation. 1. Introductory Remarks The present paper describes a project aimed at developing the first annotated corpus of Russian texts. Large text corpora have been used in the computational linguistics community for quite a long time now; at present, over 20 large corpora for the main European languages are available, the largest of them containing hundreds of millions of words (Language Resources (1997); Marcus, Santorini and Marcinkiewicz (1993); Kurohashi, Nagao (1998)). For Russian, annotated corpora had been nonexistent until 2000 when the first part of the corpus reported here was compiled (Boguslavsky et al., 2000). Since then, Russian corpus linguistics has been evolving
