QCS
Published 1 January 2003Open access
Daniel Dunlavy, John M. Conroy, Dianne P. O’Leary
Citations24
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The QCS information retrieval system is presented as a tool for querying, clustering, and summarizing document sets and facilitates the inclusion of new technologies targeting these three IR tasks.
Abstract
The QCS information retrieval (IR) system is presented as a tool for querying, clustering, and summarizing document sets. QCS has been developed as a modular development framework, and thus facilitates the inclusion of new technologies targeting these three IR tasks. Details of the system architecture, the QCS interface, and preliminary results are presented.
Keywords
Computer ScienceBiochemistry, Genetics and Molecular Biology
Proceedings of the IEEEA tutorial on hidden Markov models and selected applications in speech recognition
22,785 Citations1989L. R. Rabiner
Journal of the American Society for Information ScienceIndexing by latent semantic analysis
12,677 Citations1990Scott Deerwester, Susan Dumais +3 more
ROUGE: A Package for Automatic Evaluation of Summaries
8,286 Citations2004Chin-Yew Lin
Four different RouGE measures are introduced: ROUGE-N, ROUge-L, R OUGE-W, and ROUAGE-S included in the Rouge summarization evaluation package and their evaluations.
Communications of the ACMA vector space model for automatic indexing
7,434 Citations1975Gerard Salton, Anita M.-Y. Wong +1 more
An approach based on space density computations is used to choose an optimum indexing vocabulary for a collection of documents, demonstating the usefulness of the model.
Accurate methods for the statistics of surprise and coincidence
2,688 Citations1993Ted Dunning
Machine LearningConcept Decompositions for Large Sparse Text Data Using Clustering
1,297 Citations2001Inderjit S. Dhillon, Dharmendra S. Modha
The concept vectors produced by the spherical k-means algorithm constitute a powerful sparse and localized “basis” for text data sets and are localized in the word space, are sparse, and tend towards orthonormality.
Behavior Research Methods, Instruments, & ComputersImproving the retrieval of information from external sources
576 Citations1991Susan Dumais
A statistical method is described called latent semantic indexing, which models the implicit higher order structure in the association of words and objects and improves retrieval performance by up to 30%.
Massive computingData Mining for Scientific and Engineering Applications
280 Citations2001Robert L. Grossman, Chandrika Kamath +2 more
It is shown that the diversity of applications, the richness of the problems faced by practitioners, and the opportunity to borrow ideas from other domains, make scientific data mining an exciting and challenging field.
Massive computingEfficient Clustering of Very Large Document Collections
236 Citations2001Inderjit S. Dhillon, James Fan +1 more
This paper presents a time and memory efficient technique for the entire clustering process, including the creation of the vector space model, and demonstrates how this efficiency is obtained by a memory-efficient multi-threaded preprocessing scheme and a fast clustering algorithm that fully exploits the sparsity of the data set.
Tracking and summarizing news on a daily basis with Columbia's Newsblaster
233 Citations2002Kathleen McKeown, Regina Barzilay +7 more
Columbia's Newsblaster system for online news summarization is presented, a system that crawls the web for news articles, clusters them on specific topics and produces multidocument summaries for each cluster.
ACM Transactions on Information SystemsA semidiscrete matrix decomposition for latent semantic indexing information retrieval
224 Citations1998Tamara G. Kolda, Dianne P. O’Leary
It is shown that SDD- based LSI does as well as SVD-based LSI in terms of document retrieval while requiring only one-twentieth the storage and one-half the time to compute each query.
Statistical Data Mining and Knowledge Discovery
134 Citations2003
Iterative clustering of high dimensional text data augmented by local search
128 Citations2003Inderjit S. Dhillon, Yuqiang Guan +1 more
A local search procedure that refines a given clustering by incrementally moving data points between clusters, thus achieving a higher objective function value and a powerful "ping-pong" strategy that often qualitatively improves k-means clustering and is computationally efficient.
NewsInEssence
64 Citations2001Dragomir Radev, Sasha Blair-Goldensohn +2 more
A system for finding, visualizing and summarizing a topic-based cluster of news stories and producing summaries of a subset of the stories that it finds, according to parameters specified by the user.
ACM Transactions on Information SystemsMultidocument summarization
58 Citations2004Manuel Jesús Maña López, Manuel de Buenaga Rodríguez +1 more
This article proposes in addition to the classification capacity of clustering techniques, the possibility of offering a indicative extract about the contents of several sources by means of multidocument summarization techniques.
Tagging sentence boundaries
57 Citations2000Andrei Mikheev
This paper describes an extension of the traditional POS tagging by combining it with the document-centered approach to proper name identification and abbreviation handling that made the resulting system robust to domain and topic shifts.
University Libraries (University of Maryland)Text Summarization via Hidden Markov Models and Pivoted QR Matrix Decomposition
40 Citations2001John M. Conroy, Dianne P. O’Leary
Two approaches to generating sentence extract summary of a document using a pivoted QR decomposition of the term-sentence matrix and a hidden Markov model that judges the likelihood that each sentence should be contained in the summary are presented.
WebInEssence: A Personalized Web-Based Multi-Document Summarization and Recommendation System
28 Citations2001Dragomir Radev, Weiguo Fan +1 more
This paper addresses some of the design issues to improve the scalability and readability of the multi-document summarizer included in WebInEssence.
Performance of a Three-Stage System for Multi-Document Summarization
22 Citations2003Daniel Dunlavy, John M. Conroy +2 more
The preprocessing of the data for the group's needs consisted of term identification, part-of-speech (POS) tagging, sentence boundary detection and SGML DTD processing, and the post-processing consisted of removing lead adverbs such as “And” or “But” to make the summaries flow more easily.
