login

Dimensionality reduction aids term co-occurrence based multi-document summarization

Published 1 January 2006Open access
Ben Hachey, Gabriel Murray, David Reitter
Citations14
View PDF

TL;DR

This work uses a representation derived from the singular value decomposition of a term co-occurrence matrix in the Embra system to model text semantics and finds that Embra performs better with dimensionality reduction.

Abstract

A key task in an extraction system for query-oriented multi-document summarisation, necessary for computing relevance and redundancy, is modelling text semantics. In the Embra system, we use a representation derived from the singular value decomposition of a term co-occurrence matrix. We present methods to show the reliability of performance improvements. We find that Embra performs better with dimensionality reduction.

Keywords

Computer Science