login

Exploiting a parallel TEXT - DATA corpus

Published 1 January 2003
Somayajulu Sripada, Ehud Reiter, Jim Hunter, Jin Yu
Citations25

TL;DR

A parallel corpus of naturally occurring weather forecast texts and their corresponding forecast data; data that the human authors inspected while writing the forecast texts is described, to acquire knowledge needed to build a text generator for automatically producing textual weather forecasts from numerical weather prediction data.

Abstract

In this paper, we describe SUMTIME-METEO, a parallel corpus of naturally occurring weather forecast texts and their corresponding forecast data; data that the human authors inspected while writing the forecast texts. We have analysed the corpus to acquire knowledge needed to build a text generator for automatically producing textual weather forecasts from numerical weather prediction data. Although parallel corpora are commonly used for the development and evaluation of machine translation technology, it is fairly novel in the text generation community. Our analyses of the corpus, in some cases, produced ambiguous results that are not useful and reflected inconsistencies in the underlying corpus. Despite the internal inconsistencies, the text-data parallel corpus was helpful in generating initial hypotheses, which were then tested with knowledge from other sources. We also describe how we have used the corpus for evaluating our prototype forecast text generator. 1

Keywords

Computer ScienceBiochemistry, Genetics and Molecular Biology