Multi-document Biography Summarization
arXiv (Cornell University)Published 26 January 2005Open access
Liang Zhou, Miruna Ticrea, Eduard Hovy
Citations52
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
In this paper we describe a biography summarization system using sentence classification and ideas from information retrieval. Although the individual techniques are not new, assembling and applying them to generate multi-document biographies is new. Our system was evaluated in DUC2004. It is among the top performers in task 5-short summaries focused by person questions.
Keywords
Computer ScienceBiochemistry, Genetics and Molecular Biology
ACM Transactions on Intelligent Systems and TechnologyLIBSVM
41,340 Citations2011Chih-Chung Chang, Chih‐Jen Lin
Issues such as solving SVM optimization problems theoretical convergence multiclass classification probability estimates and parameter selection are discussed in detail.
C4.5: Programs for Machine Learning
23,665 Citations1992J. R. Quinlan
A complete guide to the C4.5 system as implemented in C for the UNIX environment, which starts from simple core learning methods and shows how they can be elaborated and extended to deal with typical problems such as missing data and over hitting.
Lecture notes in computer scienceText categorization with Support Vector Machines: Learning with many relevant features
7,925 Citations1998Thorsten Joachims
SVMs achieve substantial improvements over the currently best performing methods and behave robustly over a variety of di-erent learning tasks, eliminating the need for manual parameter tuning.
Transformation-based error-driven learning and natural language processing: a case study in part-of-speech tagging
1,531 Citations1995Eric Brill
This paper describes a simple rule-based approach to automated learning of linguistic knowledge that has been shown for a number of tasks to capture information in a clearer and more direct fashion without a compromise in performance.
A trainable document summarizer
1,349 Citations1995Julian Kupiec, Jan Pedersen +1 more
The trends in the results are in agreement with those of Edmundson who used a subjectively weighted combination of features as opposed to training the feature weights using a corpus, which suggests that even shorter extracts may be useful indicative summmies.
Statistics-Based Summarization - Step One: Sentence Compression
405 Citations2000Kevin Knight, Daniel Marcu
This paper focuses on sentence compression, a simpler version of this larger challenge, and aims to achieve two goals simultaneously: the compressions should be grammatical, and they should retain the most important pieces of information.
Information fusion in the context of multi-document summarization
385 Citations1999Regina Barzilay, Kathleen McKeown +1 more
This approach is unique in its usage of language generation to reformulate the wording of the summary by identifying and synthesizing similar elements across related text from a set of multiple documents.
Generating natural language summaries from multiple on-line sources
381 Citations1998Dragomir Radev, Kathleen McKeown
Multi-document summarization by sentence extraction
356 Citations2000Jade Goldstein, Vibhu O. Mittal +2 more
This paper discusses a text extraction approach to multi- document summarization that builds on single-document summarization methods by using additional, available information about the document set as a whole and the relationships between the documents.
Journal of Artificial Intelligence ResearchInferring Strategies for Sentence Ordering in Multidocument News Summarization
333 Citations2002Regina Barzilay, Noémie Elhadad
A strategy for ordering information that combines constraints from chronological order of events and topical relatedness is implemented and Evaluation of the augmented algorithm shows a significant improvement of the ordering over two baseline strategies.
Medical Entomology and ZoologyEvaluating Natural Language Processing Systems: An Analysis and Review
276 Citations1996Karen Spärck Jones, Julia Galliers
The automatic construction of large-scale corpora for summarization research
133 Citations1999Daniel Marcu
An algorithm is developed that constructs corpora automatically and is shown to be close to that of humans by means of an empirical experiment, which suggests extraction strategies that could improve the performance of automatic summarization systems.
Producing biographical summaries
65 Citations2001Barry Schiffman, Inderjeet Mani +1 more
A biographical multi-document summarizer that summarizes information about people described in the news, using corpus statistics along with linguistic knowledge to select and merge descriptions of people from a document collection, removing redundant descriptions.
Automated multi-document summarization in NeATS
60 Citations2002Chin-Yew Lin, Eduard Hovy
This paper describes the multi-document text summarization system NeATS, which was among the top two performers of the DUC-01 evaluation.
Recent developments in text summarization
56 Citations2001Inderjeet Mani
The significance of some recent developments in summarization technology is discussed, often bundled with information retrieval tools, as well as from the need for corporate knowledge management.
Meta-evaluation of summaries in a cross-lingual environment using content-based metrics
54 Citations2002Horacio Saggion, Simone Teufel +2 more
A framework for the evaluation of summaries in English and Chinese using similarity measures that can be used to evaluate extractive, non-extractive, single and multi-document summarization is described.
A web-trained extraction summarization system
32 Citations2003Liang Zhou, Eduard Hovy
This paper presents a summarization system that uses the web as the source of training data and automatically learning to perform the task of extraction-based summarization at a level comparable to the best DUC systems.
