BBN/UMD at DUC-2004: Topiary
Published 1 January 2004
David Zajic, Bonnie J. Dorr, Richard Schwartz
Citations72
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
It is shown that the combination of linguistically motivated sentence compression with statistically selected topic terms performs better than either alone or either alone, according to some automatic summary evaluation measures.
Abstract
This paper reports our results at DUC2004 and describes our approach, implemented in a system called Topiary. We will show that the combination of linguistically motivated sentence compression with statistically selected topic terms performs better than either alone, according to some automatic summary evaluation measures.
Keywords
Computer Science
Machine LearningAn Algorithm that Learns What's in a Name
784 Citations1999Daniel M. Bikel, Richard Schwartz +1 more
IdentiFinderTM, a hidden Markov model that learns to recognize and classify names, dates, times, and numerical quantities, is evaluated and is competitive with approaches based on handcrafted rules on mixed case text and superior on text where case information is not available.
An evaluation of phrasal and clustered representations on a text categorization task
547 Citations1992David Lewis
It is shown that optimal effectiveness occurs when using only a small proportion of the indexing terms available, and that effectiveness peaks at a higher feature set size and lower effectiveness level for a syntactic phrase indexing than for word-based indexing.
Statistics-Based Summarization - Step One: Sentence Compression
405 Citations2000Kevin Knight, Daniel Marcu
This paper focuses on sentence compression, a simpler version of this larger challenge, and aims to achieve two goals simultaneously: the compressions should be grammatical, and they should retain the most important pieces of information.
Hedge Trimmer
231 Citations2003Bonnie J. Dorr, David Zajic +1 more
Hedge Trimmer is presented, a HEaDline GEneration system that creates a headline for a newspaper story using linguistically-motivated heuristics to guide the choice of a potential headline.
Munich Personal RePEc Archive (Ludwig Maximilian University of Munich)Algorithms That Learn to Extract Information BBN: Description of the Sift System as Used for MUC-7
72 Citations1998S.L. Miller, Michael Crystal +5 more
For MUC-7, BBN has for the first time fielded a fully-trained system for NE, TE, and TR; results are all the output of statistical language models trained on annotated data, rather than programs executing handwritten rules.
A maximum likelihood model for topic classification of broadcast news
57 Citations1997Richard Schwartz, Toru Imai +3 more
A new algorithm for topic classification that allows discrimination among thousands of topics, trained by EM, has sharper distributions of words that result in more accurate topic classification.
Headline Summarization at ISI
29 Citations2003Liang Zhou
This work presents a headline summarization system that is built at ISI and is a top performer for DUC2003’s task 1, generating very short summaries (10 words or less).
ACM Transactions on Asian Language Information ProcessingCross-language headline generation for Hindi
25 Citations2003Bonnie J. Dorr, David Zajic +1 more
It is demonstrated in both automatic and human evaluations that the linguistically motivated approach outperforms two other surrogate-generation methods: a statistical system and a topic discovery system.
Tailoring text using topic words: Selection and compression
9 Citations2004Timm Euler
A new method to extract sentences: that deal with a certain topic from a given text that is based on automatically computed lists of words that represent the desired topics, extending previous methods that rely on syntactical clues only.
