Large linguistically-processed web corpora for multiple languages
Published 1 January 2006Open access
Marco Baroni, Adam Kilgarriff
Citations123
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This work provides Web access to the corpora in the query tool, the Sketch Engine, for German and Italian, with corpus sizes of over 1 billion words in each case.
Abstract
Comunicació presentada a: EACL '06: Eleventh Conference of the European Chapter of the Association for Computational Linguistics: Posters & Demonstrations celebrat del 5 al 6 d'abril de 2006 a Trento, Itàlia.
Keywords
Computer Science
Computer Networks and ISDN SystemsSyntactic clustering of the Web
1,347 Citations1997Andrei Broder, S. Glassman +2 more
An efficient way to determine the syntactic similarity of files is developed and applied to every document on the World Wide Web, and a clustering of all the documents that are syntactically similar is built.
Computational LinguisticsIntroduction to the Special Issue on the Web as Corpus
926 Citations2003Adam Kilgarriff, Gregory Grefenstette
This special issue of Computational Linguistics explores ways in which this dream of freely available language data in vast quantity and freely available is being explored.
Mining the Web: Discovering Knowledge from Hypertext Data
695 Citations2002Soumen Chakrabarti
Literary and Linguistic ComputingCorpus Design Criteria
363 Citations1992S. D. Atkins
Two stages in Corpus Building: Tasks, Expertise, Personnel and Personnel 3 2.4.2 Stages in Corpus building: T tasks, expertise, personnel.
Applied Corpus LinguisticsMaking the Web More Useful as a Source for Linguistic Corpora
81 Citations2004William H. Fletcher
With judicious selection Web pages provide representative language samples, often prove more useful than off-the-shelf corpora for special information needs, and complement and verify data from traditional corpora.
Introduction to the Special Issue on the Web as Corpus
37 Citations2008Adam Kilgarriff, Gregory Grefenstette
