Information extraction from voicemail transcripts
Published 1 January 2002Open access
Martin Jansche, Steven Abney
Citations39
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
Techniques for obtaining basic information about the caller/sender or a phone number for returning calls must be obtained from message transcripts or other sources.
Abstract
Voicemail is not like email. Even such basic information as the name of the caller/ sender or a phone number for returning calls is not represented explicitly and must be obtained from message transcripts or other sources. We discuss techniques for doing this and the challenges these tasks present.
Keywords
Computer Science
Maximum Entropy Model for Part-Of-Speech Tagging
1,276 Citations1996Adwait Ratnaparkhi
A statistical model which trains from a corpus annotated with Part Of Speech tags and assigns them to previously unseen text with state of the art accuracy and discusses the corpus consistency problems discovered during the implementation of these features.
PNrule: A New Framework for Learning Classifier Models in Data Mining (A Case-Study in Network Intrusion Detection)
140 Citations2001Ramesh K. Agarwal, Mahesh V. Joshi
Although no single technique is proven to be the best in all situations, techniques that learn rule-based models are especially popular in the domain of data mining, and can be contributed to the easy interpretability of the rules by humans, and competitive performance exhibited by rule- based models in many application domains.
Mining needle in a haystack
127 Citations2001Mahesh V. Joshi, Ramesh C. Agarwal +1 more
This paper designs various synthetic data models to identify and analyze the situations in which two state-of-the-art methods, RIPPER and C4.5 rules, either fail to learn a model or learn a very poor model, and learns a model with significantly better recall and precision levels.
SCANMail: browsing and searching speech data by content
33 Citations2001Julia Hirschberg, Michiel Bacchiani +7 more
SCANMail is described, a system that employs automatic speech recognition, information retrieval, information extraction, and human computer interaction technology to permit users to browse and search their voicemail messages by content through a graphical user interface interface.
Information extraction from voicemail
29 Citations2001Jing Huang, Geoffrey Zweig +1 more
This work presents three information extraction methods, one based on hand-crafted rules, one Based on maximum entropy tagging, and oneBased on probabilistic transducer induction, that perform on both manually transcribed messages and on the output of a speech recognition system.
Philosophical Transactions of the Royal Society A Mathematical Physical and Engineering SciencesInformation extraction from broadcast news
23 Citations2000Yoshihiko Gotoh, Steve Renals
This paper discusses the development of trainable statistical models for extracting content from television and radio news broadcasts, and focuses on statistical finite–state models for identifying proper names and other named entities in broadcast speech.
Automatic transcription of voicemail at AT&T
15 Citations2002Michiel Bacchiani
Reports on the automatic transcription accuracy of voicemail messages shows that vocal tract length normalization and adaptation using linear transformations, proven to improve accuracy on the Switchboard task, provide similar accuracy improvements on this task.
Caller identification for the SCANMail voicemail browser
6 Citations2001A. E. Rosenberg, Julia Hirschberg +4 more
Describing CallerID, the server tool attached to SCANMail for the purpose of providing caller labels for voicemail messages, and some results of performance evaluations of the caller identification capability are provided.
