login

Development of a Dependency Treebank for Russian and its Possible Applications in NLP

Published 1 May 2002
Igor Boguslavsky, Ivan Chardin, Svetlana Grigorieva, Nikolai Grigoriev, Leonid Iomdin, Leonid Kreidlin
Citations47
SJR quartileQ1
SJR score0.48
SNIP1.79

TL;DR

A tagging scheme designed for the Russian Treebank is described and tools used for corpus creation are presented, which gives a more fine-grained representation of syntactic phenomena.

Abstract

The paper describes a tagging scheme designed for the Russian Treebank and presents tools used for corpus creation. 1. Introductory Remarks The present paper describes a project aimed at developing the first annotated corpus of Russian texts. Large text corpora have been used in the computational linguistics community for quite a long time now; at present, over 20 large corpora for the main European languages are available, the largest of them containing hundreds of millions of words (Language Resources (1997); Marcus, Santorini and Marcinkiewicz (1993); Kurohashi, Nagao (1998)). For Russian, annotated corpora had been nonexistent until 2000 when the first part of the corpus reported here was compiled (Boguslavsky et al., 2000). Since then, Russian corpus linguistics has been evolving

Keywords

Computer Science