login

Development of a document summarization system for effective information services

Published 25 June 1997
Dong-Hyun Jang, Sung Hyun Myaeng
Citations13

TL;DR

A system that constructs a summary by extracting sentences that are likely to represent the main theme of a document by using a probabilistic model that takes into account lexical and statistical information obtained from a document corpus.

Abstract

This paper describes a system that constructs a summary by extracting sentences that are likely to represent the main theme of a document. As a way of selecting summary sentences, the system uses a probabilistic model that takes into account lexical and statistical information obtained from a document corpus. As such, the system consists of two parts: the training part and the summarization part. The former processes sentences that have been manually tagged for summary sentences and extracts necessary statistical information of various kinds, and the latter uses the information to calculate the likelihood of each sentence to become part of a summary. There are at least three unique aspects of this research. First of all, the system uses a text model to identify different components of a text and eliminates parts of text that are not likely to contain summary sentences. Second, although the probabilistic model stems from an existing model developed for English texts, it applies the model to compute multiple probability values based on several features, and computes the final value by combining pieces of evidence from different sources (features) with the Dempster-Shafer theory. Finally, the system is the first of this kind for Korean texts.

Keywords

Computer Science