Extracting personal names from email
Published 1 January 2005Open access
Einat Minkov, Richard C. Wang, William W. Cohen
Citations161
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
Two methods for improving performance of person name recognizers for email are presented: email-specific structural features and a recall-enhancing method which exploits name repetition across multiple documents.
Abstract
There has been little prior work on Named Entity Recognition for "informal" documents like email. We present two methods for improving performance of person name recognizers for email: email-specific structural features and a recall-enhancing method which exploits name repetition across multiple documents.
Keywords
Computer ScienceDecision Sciences
ScholarlyCommons (University of Pennsylvania)Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data
12,978 Citations2001John Lafferty, Andrew McCallum +1 more
This work presents iterative parameter estimation algorithms for conditional random fields and compares the performance of the resulting models to HMMs and MEMMs on synthetic and natural-language data.
Mathematical ProgrammingOn the limited memory BFGS method for large scale optimization
8,529 Citations1989Dong C. Liu, Jorge Nocedal
The numerical tests indicate that the L-BFGS method is faster than the method of Buckley and LeNir, and is better able to use additional storage to accelerate convergence, and the convergence properties are studied to prove global convergence on uniformly convex problems.
Biological sequence analysis probabilistic models of proteins and nucleic acids
3,654 Citations2010Richard Durbin
Unsupervised word sense disambiguation rivaling supervised methods
2,408 Citations1995David Yarowsky
An unsupervised learning algorithm for sense disambiguation that, when trained on unannotated English text, rivals the performance of supervised techniques that require time-consuming hand annotations.
Discriminative training methods for hidden Markov models
1,889 Citations2002Michael Collins
Experimental results on part-of-speech tagging and base noun phrase chunking are given, in both cases showing improvements over results for a maximum-entropy tagger.
Transformation-based error-driven learning and natural language processing: a case study in part-of-speech tagging
1,531 Citations1995Eric Brill
This paper describes a simple rule-based approach to automated learning of linguistic knowledge that has been shown for a number of tasks to capture information in a clearer and more direct fashion without a compromise in performance.
Information RetrievalAutomating the Construction of Internet Portals with Machine Learning
1,265 Citations2000Andrew Kachites McCallum, Kamal Nigam +2 more
New research in reinforcement learning, information extraction and text classification that enables efficient spidering, the identification of informative text segments, and the population of topic hierarchies are described.
Shallow parsing with conditional random fields
1,251 Citations2003Fei Sha, Fernando Pereira
This work shows how to train a conditional random field to achieve performance as good as any reported base noun-phrase chunking method on the CoNLL task, and better than any reported single model.
Early results for named entity recognition with conditional random fields, feature induction and web-enhanced lexicons
1,161 Citations2003Andrew McCallum, Wei Li
This work has shown that conditionally-trained models, such as conditional maximum entropy models, handle inter-dependent features of greedy sequence modeling in NLP well.
Large margin classification using the perceptron algorithm
1,099 Citations1998Yoav Freund, Robert E. Schapire
Machine LearningAn Algorithm that Learns What's in a Name
784 Citations1999Daniel M. Bikel, Richard Schwartz +1 more
IdentiFinderTM, a hidden Markov model that learns to recognize and classify names, dates, times, and numerical quantities, is evaluated and is competitive with approaches based on handcrafted rules on mixed case text and superior on text where case information is not available.
PubMedConstructing biological knowledge bases by extracting information from text sources.
597 Citations1999Mark Craven, Johan Kumlien
A research effort aimed at automatically mapping information from text sources into structured representations, such as knowledge bases, is begun, to use machine-learning methods to induce routines for extracting facts from text.
Conference on Email and Anti-SpamIntroducing the Enron Corpus.
542 Citations2004Bryan Klimt, Yiming Yang
The goal in this paper is to analyze the suitability of this corpus for exploring how to classify messages as organized by a human, so these folders would have likely been misleading.
Artificial IntelligenceLearning to construct knowledge bases from the World Wide Web
472 Citations2000Mark Craven, Dan DiPasquo +5 more
The goal of the research described here is to automatically create a computer understandable knowledge base whose content mirrors that of the World Wide Web, and several machine learning algorithms for this task are described, and promising initial results with a prototype system that has created a knowledge base describing university people, courses, and research projects.
Hidden Markov support vector machines
469 Citations2003Yasemin Altün, Ioannis Tsochantaridis +1 more
This paper presents a novel discriminative learning technique for label sequences based on a combination of the two most successful learning algorithms, Support Vector Machines and Hidden Markov Models which it is called HM-SVMs and handles dependencies between neighboring labels using Viterbi decoding.
Ranking algorithms for named-entity extraction
241 Citations2001Michael Collins
Algorithms which rerank the top N hypotheses from a maximum-entropy tagger, the application being the recovery of named-entity boundaries in a corpus of web data, using the voted perceptron algorithm.
Information extraction from HTML: application of a general machine learning approach
237 Citations1998Dayne Freitag
This work shows how information extraction can be cast as a standard machine learning problem, and argues for the suitability of relational learning in solving it, and the implementation of a general-purpose relational learner for information extraction, SRV.
ScholarWorks@UMassAmherst (University of Massachusetts Amherst)Extracting Social Networks and Contact Information From Email and the Web
225 Citations2004Aron Culotta, Ron Bekkerman +1 more
An end-to-end system that extracts a user's social network and its members' contact information given the user's email inbox and discusses the capabilities of the system for address book population, expert-finding, and social network analysis.
Empirical Methods in Natural Language ProcessingLearning to Classify Email into "Speech Acts".
214 Citations2004William W. Cohen, Vitor R. Carvalho +1 more
It is demonstrated that, although this categorization problem is quite different from “topical” text classification, certain categories of messages can nonetheless be detected with high precision and reasonable recall using existing text-classification learning methods.
Exploiting diverse knowledge sources via maximum entropy in named entity recognition
210 Citations1998Andrew Borthwick, John Sterling +2 more
This paper describes a novel statistical namedentity recognition system built around a maximum entity framework using the framework of maximum entropy theory and utilizing a flexible object-based architecture to make use of an extraordinarily diverse range of knowledge sources in making its tagging decisions.
CORE Scholar (Wright State University)Collective Segmentation and Labeling of Distant Entities in Information Extraction
148 Citations2004Charles Sutton, Andrew McCallum
This work presents a CRF that explicitly represents dependencies between the labels of pairs of similar words in a document, and shows that learning these dependencies leads to a 13.7% reduction in error on the field that had caused the most repetition errors.
Confidence estimation for information extraction
115 Citations2004Aron Culotta, Andrew McCallum
This work evaluates a information extraction system based on a linear-chain conditional random field (CRF), a probabilistic model which has performed well on information extraction tasks because of its ability to capture arbitrary, overlapping features of the input in a Markov model.
Methods for domain-independent information extraction from the web: an experimental comparison
114 Citations2004Oren Etzioni, Michael Cafarella +6 more
Three distinct ways to improve KNOWITALL's recall and extraction rate without sacrificing precision are presented and evaluated and their performance is evaluated.
INQUERY does battle with TREC-6
83 Citations1997James Allan, James P. Callan +5 more
Conference on Email and Anti-SpamLearning to Extract Signature and Reply Lines from Email.
82 Citations2004Vitor R. Carvalho, William W. Cohen
Methods for automatically identifying signature blocks and reply lines in plain-text email messages are described, based on applying machine learning methods to a sequential representation of an email message, in which each email is represented as a sequence of lines.
Markov models for language-independent named entity recognition
78 Citations2002Robert Malouf
This report describes the application of Markov models to the problem of language-independent named entity recognition for the CoNLL-2002 shared task.
OPAL (Open@LaTrobe) (La Trobe University)Coordination in Teams: Evidence from a Simulated Management Game
43 Citations2018Robert E. Kraut, Susan R. Fussell +2 more
Using corpus-derived name lists for named entity recognition
41 Citations2000Mark Stevenson, Robert Gaizauskas
This paper describes experiments to establish the performance of a named entity recognition system which builds categorized lists of names from manually annotated training data and shows that by using simple filtering techniques for improving the automatically acquired lists, substantial performance benefits can be achieved.
Memory-based named entity recognition using unannotated data
40 Citations2003Fien De Meulder, Walter Daelemans
The memory-based learner Timbl (Daelemans et al., 2002) was used to find names in English and German newspaper text to show that gazetteers are not beneficial in the English case, while they are beneficial for the German data.
Information extraction from voicemail transcripts
39 Citations2002Martin Jansche, Steven Abney
Techniques for obtaining basic information about the caller/sender or a phone number for returning calls must be obtained from message transcripts or other sources.
ScholarWorks@UMassAmherst (University of Massachusetts Amherst)An Exploration of Entity Models, Collective Classification and Relation Description
36 Citations2004Hema Raghavan, James Allan +1 more
This paper explores the middle ground using a representation which is term entity models, in which questions about structured data may be posed and answered, but the complexities and task-specific restrictions of ontologies are avoided.
Information extraction from voicemail
29 Citations2001Jing Huang, Geoffrey Zweig +1 more
This work presents three information extraction methods, one based on hand-crafted rules, one Based on maximum entropy tagging, and oneBased on probabilistic transducer induction, that perform on both manually transcribed messages and on the output of a speech recognition system.
Relational Markov Networks for Collective Information Extraction
21 Citations2004Razvan Bunescu and Raymond J. Mooney
A new IE method is presented that employs Relational Markov Networks, which can represent arbitrary dependencies between extractions, which allows for “collective information extraction” that exploits the mutual influence between possible extractions.
ACM SIGCAS Computers and SocietyFinding lists of people on the web
16 Citations2004Latanya Sweeney
RosterFinder works by identifying rosters from candidate Web pages based on the ratio of distinct known names to distinct words appearing in the page, which is a simple algorithm for locating Web pages that consist predominately of a list of names.
Research Showcase @ Carnegie Mellon University (Carnegie Mellon University)Learning to Understand Web Site Update Requests
12 Citations2018William W. Cohen, Einat Minkov +1 more
An intelligent system that can process natural language website update requests semi-automatically is proposed, which can analyze requests, posted via email, to update the factual content of individual tuples in a database-backed website.
