The Good, the Bad, and the Hazy: Design Decisions in Web Corpus Construction
HAL (Le Centre pour la Communication Scientifique Directe)Published 22 July 2013
Roland Schäfer, Adrien Barbaresi, Felix Bildhauer
Citations24
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This paper examines notions of text quality in the context of web corpus construction, and describes the general approach to the construction of carefully cleansed and non-destructively normalized web corpora.
Abstract
International audience
Keywords
Computer Science
Journal of the American Statistical AssociationContent Analysis: An Introduction to its Methodology.
24,566 Citations1984Mack Shelley, Klaus Krippendorff
arXiv (Cornell University)Assessing agreement on classification tasks: the kappa statistic
2,092 Citations1996Jean Carletta
N-gram-based text categorization
1,503 Citations1994William B. Cavnar, John M. Trenkle
An N-gram-based approach to text categorization that is tolerant of textual errors is described, which worked very well for language classification and worked reasonably well for classifying articles from a number of different computer-oriented newsgroups according to subject.
Language Resources and EvaluationThe WaCky wide web: a collection of very large linguistically processed web-crawled corpora
1,168 Citations2009Marco Baroni, Silvia Bernardini +2 more
UkWaC, deWaC and itWaC are introduced, three very large corpora of English, German, and Italian built by web crawling, and the methodology and tools used in their construction are described.
Computer NetworksOn near-uniform URL sampling
276 Citations2000Monika Henzinger, Allan Heydon +2 more
This paper suggests ways of improving sampling based on random walks of the Web graph to make the samples closer to uniform and suggests a natural test bed based onrandom graphs for testing the effectiveness of the procedures.
Computational LinguisticsWhat Determines Inter-Coder Agreement in Manual Annotations? A Meta-Analytic Investigation
126 Citations2011Petra Saskia Bayerl, Karsten I. Paul
A meta-analytic investigation of annotation studies reporting agreement percentages found seven factors that influence reported agreement values: annotation domain, number of categories in a coding scheme,Number of annotators in a project, whether annotators received training, the intensity of annotator training, and the annotation purpose.
Methods for Sampling Pages Uniformly from the World Wide Web
86 Citations2001Paat Rusmevichientong, David M. Pennock +2 more
Two new algorithms for generating uniformly random samples of pages from the World Wide Web are presented, building upon recent work by Henzinger et al. (2000) and Bar-Yossef et al (2000), based on a weighted random-walk methodology.
Synthesis lectures on human language technologiesWeb Corpus Construction
38 Citations2013Roland Schäfer, Felix Bildhauer
LDV-Forum/Journal for language technology and computational linguisticsScalable Construction of High-Quality Web Corpora
37 Citations2013Chris Biemann, Felix Bildhauer +7 more
This article first focuses on web crawling and the pros and cons of the existing crawling strategies, and describes how the crawled data can be linguistically pre-processed in a parallelized way that allows the processing of web-scale input data.
Building a 70 billion word corpus of English from ClueWeb
27 Citations2012Jan Pomikálek, Miloš Jakubíček +1 more
The tools used for boilerplate cleaning and for de-duplication that was performed not only on full (document-level) duplicates but also on the level of near-duplicate texts are described and the impact of each of the performed pre-processing steps on the final corpus size is shown.
