login

CUCWeb

Published 1 January 2006Open access
Gemma Boleda, Stefan Bott, Rafael Meza, Carlos Castillo, Toni Badía, V. López
Citations20
View PDF

TL;DR

CUCWeb, a 166 million word corpus for Catalan built by crawling the Web, has been annotated with NLP tools and made available to language users through a flexible web interface.

Abstract

This paper presents CUCWeb, a 166 million word corpus for Catalan built by crawling the Web. The corpus has been annotated with NLP tools and made available to language users through a flexible web interface. The developed architecture is quite general, so that it can be used to create corpora for other languages.

Keywords

Computer Science