The Construction of a 500-Million-Word Reference Corpus of Contemporary Written Dutch
Theory and applications of natural language processingPublished 11 November 2012Open access
Nelleke Oostdijk, Martin Reynaert, Véronique Hoste, Ineke Schuurman
Citations168
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The present chapter describes how in two consecutive STEVIN-funded projects, viz.
Abstract
\n Contains fulltext :\n 116265.pdf (Publisher’s version ) (Open Access)\n
Keywords
Computer ScienceArts and Humanities
Computational LinguisticsThe Proposition Bank: An Annotated Corpus of Semantic Roles
2,302 Citations2005Martha Palmer, Daniel Gildea +1 more
An automatic system for semantic role tagging trained on the corpus is described and the effect on its performance of various types of information is discussed, including a comparison of full syntactic parsing with a flat representation and the contribution of the empty trace categories of the treebank.
arXiv (Cornell University)Assessing agreement on classification tasks: the kappa statistic
2,092 Citations1996Jean Carletta
A model-theoretic coreference scoring scheme
682 Citations1995Marc Vilain, John D. Burger +3 more
This note describes a scoring scheme for the coreference task in MUC6 that improves on the original approach by grounding the scoring scheme in terms of a model; producing more intuitive recall and precision scores; and not requiring explicit computation of the transitive closure of coreference.
The BNC Handbook: Exploring the British National Corpus with SARA
481 Citations1998Guy Aston, Lou Burnard
This textbook is designed to provide a detailed understanding of the principles and practices underlying the use of large language corpora in exploratory learning and English language teaching and research and on the search tool SARA (SGML Aware Retrieval Application).
arXiv (Cornell University)The JRC-Acquis: A multilingual aligned parallel corpus with 20+ languages
429 Citations2006Ralf Steinberger, Bruno Pouliquen +5 more
Named entity recognition through classifier combination
403 Citations2003Radu Florian, Abe Ittycheriah +2 more
A classifier-combination experimental framework for named entity recognition in which four diverse classifiers (robust linear classifier, maximum entropy, transformation-based learning, and hidden Markov model) are combined under different conditions is presented.
At Last Parsing Is Now Operational
181 Citations2006Gertjan van Noord
A variety of cases where large amounts of parser output are used to improve the parser, and how corpus-based methods are essential to obtain accurate knowledge-based parsers are shown.
Research portal (Tilburg University)An efficient memory-based morphosyntactic tagger and parser for Dutch
158 Citations2007Antal van den Bosch, Bertjan Busser +2 more
A global analysis of the TADPOLE system shows that it is able to process text in linear time close to an estimated 2,500 words per second, while maintaining sufficient accuracy.
Anaphoric Annotation in the ARRAU Corpus
124 Citations2008Massimo Poesio, Ron Artstein
Arrau is a new corpus annotated for anaphoric relations, with information about agreement and explicit representation of multiple antecedents for ambiguous anaphic expressions and discourse antecedent for expressions which refer to abstract entities such as events, actions and plans.
Cambridge University Press eBooksComputational linguistics
101 Citations2006Masayuki Asahara, Yasuharu Den +1 more
Language Resources and EvaluationAnCora-CO: Coreferentially annotated corpora for Spanish and Catalan
85 Citations2009Marta Vilar Recasens, M. Antònia Martí
The enrichment of the AnCora corpora of Spanish and Catalan with coreference links between pronouns, full noun phrases, and discourse segments makes it possible to train and test learning-based algorithms for automatic coreference resolution, as well as to carry out bottom-up linguistic descriptions of coreference relations as they occur in real data.
University of Groningen research database (University of Groningen / Centre for Information Technology)A coreference corpus and resolution system for Dutch
74 Citations2008Iris Hendrickx, Gosse Bouma +7 more
The main outcomes of the COREA project are a corpus annotated with coreferential relations and a coreference resolution system for Dutch, and an application-oriented evaluation of the system and a standard cross-validation evaluation.
Lecture notes in computer scienceNon-interactive OCR Post-correction for Giga-Scale Digitization Projects
57 Citations2008Martin Reynaert
A non-interactive system for reducing the level of OCR-induced typographical variation in large text collections, contemporary and historical, that focuses on high-frequency words derived from the corpus to be cleaned and gathers all typographical variants for any particular focus word that lie within the predefined Levenshtein distance.
Data Archiving and Networked Services (DANS)Syntactische annotatie voor het Corpus Gesproken Nederlands (CGN)
47 Citations2002A. Hoekstra, Michael Moortgat +3 more
The paper discusses the syntactic annotation for the Spoken Dutch Corpus, a Dutch/Flemish cooperation project to build an annotated corpus of about one thousand hours of continuous speech, which amounts to 10 million words.
International Journal on Document Analysis and Recognition (IJDAR)Character confusion versus focus word-based correction of spelling and OCR variants in corpora
46 Citations2010Martin Reynaert
The character confusion-based prototype of Text-Induced Corpus Clean-up is compared to its focus word-based counterpart and evaluated on 6 years’ worth of digitized Dutch Parliamentary documents, showing that the system is not sensitive to domain variation.
Towards a Corpus Annotated for Metonymies: the Case of Location Names.
41 Citations2002Katja Markert, Malvina Nissim
A framework for annotating metonymies in domain-independent text that considers the regularity, productivity and underspecification of metonymic usage is described and a fully worked out annotation scheme for location names is presented.
International Journal of Geographical Information SystemsOn metonymy recognition for geographic information retrieval
27 Citations2008Johannes Leveling, Sven Hartrumpf
Evaluation results indicate that removing metonymic senses from the index yields a higher mean average precision (MAP) for GIR, and a memory‐based learner is used to train a classifier and standard evaluation measures such as F‐score and accuracy are determined.
Adding semantic role annotation to a corpus of written Dutch
19 Citations2007Paola Monachesi, Gerwert Stevens +1 more
This work presents an approach to automatic semantic role labeling (SRL) carried out in the context of the Dutch Language Corpus Initiative (D-Coi) project, which consists of bootstrapping from a syntactically annotated corpus by means of a rule-based tagger developed for this purpose.
Lirias (KU Leuven)Varro: An Algorithm and Toolkit for Regular Structure Discovery in Treebanks
15 Citations2010Scott N. Martens
Closed canonically ordered trees as a data structure for efficiently discovering frequently recurring unordered subtrees is introduced.
Ghent University Academic Bibliography (Ghent University)Collecting a corpus of Dutch SMS
12 Citations2012Maaske Treurniet, Orphée De Clercq +2 more
This paper presents the first freely available corpus of Dutch text messages containing data originating from the Netherlands and Flanders and shows that especially free publicity in newspapers and on social media networks results in more contributions.
Data Archiving and Networked Services (DANS)Collecting and Analysing Chats and Tweets in SoNaR
10 Citations2012Eric Sanders
A collection of chats and tweets from the Netherlands and Flanders is described and simple text analysis in the form of unigram frequency lists is carried out to illustrate the difference of language use between the various text types.
Ghent University Academic Bibliography (Ghent University)Cross-Domain Dutch Coreference Resolution
9 Citations2011Orphée De Clercq, Véronique Hoste +1 more
User requirements analysis for the design of a reference corpus of written Dutch
9 Citations2006N.H.J. Oostdijk, Lou Boves
The present paper outlines the user requirements analysis and reports the results so far in the D-Coi project, which aims to specify the design of a 500-million-word reference corpus of written Dutch.
Utrecht University Repository (Utrecht University)Which New York, which Monday? The role of background knowledge and intended audience in automatic disambiguation of spatiotemporal expressions
8 Citations2007Ineke Schuurman
The aim of MiniSTEx, a system for automatic spatiotemporal expressions, is to locate eventualities on a time-axis and to disambiguate geospatial entities in such a way that geosp spatial entities can be located on a map.
DSpace repository (University of Tartu)Spatiotemporal Annotation on Top of an Existing Treebank
6 Citations2007Ineke Schuurman
A spatiotemporal layer of annotation to be added to an existing (syntactic) treebank and may need adaptations for specific situations, even when only one language is covered.
Cultural Aspects of Spatiotemporal Analysis in Multilingual Applications
6 Citations2010Ineke Schuurman, Vincent Vandeghinste
Spatiotemporal Annotation: Interaction between Standards and other Formats
3 Citations2011Ineke Schuurman, Vincent Vandeghinste
This paper discusses an approach for situations in which ISOcat is used to mediate between such formats, and describes which conditions should be met and how ISOcat can offer a helping hand.
A Comparative Study of PDF Generation Methods: Measuring Loss of Fidelity When Converting Arabic and Persian MS Word Files to PDF
2 Citations2011Paul M. Herceg, Catherine N. Ball
Evidence is found that using PDF for data exchange of typical Arabic and Persian documents results in a loss of important electronic text content, which confuses human language technologies such as search engines, machine translati.
