CUCWeb
Published 1 January 2006Open access
Gemma Boleda, Stefan Bott, Rafael Meza, Carlos Castillo, Toni Badía, V. López
Citations20
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
CUCWeb, a 166 million word corpus for Catalan built by crawling the Web, has been annotated with NLP tools and made available to language users through a flexible web interface.
Abstract
This paper presents CUCWeb, a 166 million word corpus for Catalan built by crawling the Web. The corpus has been annotated with NLP tools and made available to language users through a flexible web interface. The developed architecture is quite general, so that it can be used to create corpora for other languages.
Keywords
Computer Science
Computational LinguisticsIntroduction to the Special Issue on the Web as Corpus
926 Citations2003Adam Kilgarriff, Gregory Grefenstette
This special issue of Computational Linguistics explores ways in which this dream of freely available language data in vast quantity and freely available is being explored.
Scaling to very very large corpora for natural language disambiguation
685 Citations2001Michele Banko, Eric Brill
This paper examines methods for effectively exploiting very large corpora when labeled data comes at a cost, and evaluates the performance of different learning methods on a prototypical natural language disambiguation task, confusion set disambigsuation.
Lecture notes in computer scienceIdentifying and Filtering Near-Duplicate Documents
376 Citations2000Andrei Broder
The algorithm for filtering near-duplicate documents discussed here has been successfully implemented and has been used for the last three years in the context of the AltaVista search engine.
Computational LinguisticsUsing the Web to Obtain Frequencies for Unseen Bigrams
357 Citations2003Frank Keller, Mirella Lapata
It is shown that the Web can be employed to obtain frequencies for bigrams that are unseen in a given corpus by querying a search engine.
Spam, damn spam, and statistics
293 Citations2004Dennis Fetterly, Mark S. Manasse +1 more
This paper proposes that some spam web pages can be identified through statistical analysis, and examines a variety of properties, including linkage structure, page content, and page evolution, and finds that outliers in the statistical distribution of these properties are highly likely to be caused by web spam.
DIGITAL.CSIC (Spanish National Research Council (CSIC))Characteristics of the Web of Spain
68 Citations2005Ricardo Baeza‐Yates, Carlos Castillo +1 more
The results of an in-depth study over a large collection of Web pages found that some of the characteristics of this collection resemble the ones of the Web at large, while others are specific to the Web of Spain, or have not been studied in the past.
CoBWeb-a crawler for the Brazilian Web
31 Citations2003A.S. da Silva, Eveline Veloso +4 more
The paper describes CoBWeb, an automatic document collector whose architecture is distributed and highly scalable, which is part of the SIAM (Information Systems in Mobile Computing Environments) search engine which is being implemented to support the Brazilian Web.
International Journal of Corpus LinguisticsCreating and using Web corpora
16 Citations2005Mike Thelwall
It was evident that the university Web sites of three nations contained a significant amount of non-English text, and academic Web English seems to be more future-oriented than British National Corpus written English.
Dipòsit Digital de la Universitat de Barcelona (Universitat de Barcelona)Un corpus general de referència de la llengua catalana
16 Citations1994Joaquim Rafel
CATCG: a general purpose parsing tool applied.
15 Citations2002Àlex Alsina, Toni Badía +5 more
This paper focuses on the language processing tool being developed at the centre and briefly describes two of its applications, CATCG, a morphosyntactic analyser designed to deal with general written Catalan text and a modular system that allows the possibility to choose the best strategy for each specific task.
Computers and the HumanitiesAutomatic acquisition of syntactic verb classes with basic resources
7 Citations2005Laia Mayol, Gemma Boleda +1 more
This paper describes a methodology aimed at grouping Catalan verbs according to their syntactic behavior, and shows that it is possible to acquire this kind of information using only a POS-tagged corpus.
