login

Segmentation and Detection at IBM

˜The œKluwer international series on information retrievalPublished 1 January 2002
S. Dharanipragada, Martin Franz, Jason S. McCarley, Todd J. Ward, W.-J. Zhu
Citations21

TL;DR

This work investigates the importance of merging microclusters together, and proposes a merging strategy which improves the performance of IBM’s story segmentation models.

Abstract

IBM's story segmentation uses a combination of decision tree and maximum entropy models. They take a variety of lexical, prosodic, semantic, and structural features as their inputs. Both types of models are source-specific, and we substantially lower C seg by combining them. IBM's topic detection system introduces a minimal hierarchy into the clustering: each cluster is comprised of one or more microclusters. We investigate the importance of merging microclusters together, and propose a merging strategy which improves our performance.

Keywords

Computer Science