Comparative experiments on learning information extractors for proteins and their interactions
Artificial Intelligence in MedicinePublished 14 December 2004
Răzvan Bunescu, Ruifang Ge, Rohit J. Kate, Edward M. Marcotte, Raymond J. Mooney, Arun Ramani
Citations432
SJR quartileQ1
SJR score1.40
SNIP1.93
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The results show that it is promising to use machine learning to automatically build systems for extracting information from biomedical text with higher precision than manually-developed rules.
Abstract
Our results show that it is promising to use machine learning to automatically build systems for extracting information from biomedical text. The results also give a broad picture of the relative strengths of a wide variety of methods when tested on a reasonably large human-annotated corpus.
Keywords
Computer ScienceBiochemistry, Genetics and Molecular Biology
The Nature of Statistical Learning Theory
39,279 Citations1995Vladimir Vapnik
TechnometricsStatistical Learning Theory
26,913 Citations1999Yuhai Wu, Vladimir Vapnik
Presenting a method for determining the necessary and sufficient conditions for consistency of learning process, the author covers function estimates from small data pools, applying these estimations to real-life problems, and much more.
NatureInitial sequencing and analysis of the human genome
24,579 Citations2001Eric S. Lander, Lauren Linton +250 more
The results of an international collaboration to produce and make freely available a draft sequence of the human genome are reported and an initial analysis is presented, describing some of the insights that can be gleaned from the sequence.
Proceedings of the IEEEA tutorial on hidden Markov models and selected applications in speech recognition
22,785 Citations1989L. R. Rabiner
Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference
16,927 Citations1988Judea Pearl
The author provides a coherent explication of probability as a language for reasoning with partial belief and offers a unifying perspective on other AI approaches to uncertainty, such as the Dempster-Shafer formalism, truth maintenance systems, and nonmonotonic logic.
ScholarlyCommons (University of Pennsylvania)Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data
12,978 Citations2001John Lafferty, Andrew McCallum +1 more
This work presents iterative parameter estimation algorithms for conditional random fields and compares the performance of the resulting models to HMMs and MEMMs on synthetic and natural-language data.
Pattern classification and scene analysis
12,643 Citations1973Richard O. Duda, Peter E. Hart
Lecture notes in computer scienceText categorization with Support Vector Machines: Learning with many relevant features
7,925 Citations1998Thorsten Joachims
SVMs achieve substantial improvements over the currently best performing methods and behave robustly over a variety of di-erent learning tasks, eliminating the need for manual parameter tuning.
Experiments with a new boosting algorithm
7,585 Citations1996Yoav Freund, Robert E. Schapire
Building a Large Annotated Corpus of English: The Penn Treebank
7,528 Citations1993Mitchell P. Marcus
Algorithms on strings, trees, and sequences: computer science and computational biology
3,529 Citations1997Dan Gusfield
Computational LinguisticsA maximum entropy approach to natural language processing
3,120 Citations1996Adam Berger, Vincent J. Della Pietra +1 more
A maximum-likelihood approach for automatically constructing maximum entropy models is presented and how to implement this approach efficiently is described, using as examples several problems in natural language processing.
Cambridge University Press eBooksAlgorithms on Strings, Trees and Sequences
3,045 Citations1997Dan Gusfield
Ukkonen’s method is the method of choice for most problems requiring the construction of a suffix tree, and it will be presented first because it is easier to understand.
Nucleic Acids ResearchThe Database of Interacting Proteins: 2004 update
2,298 Citations2003Łukasz Salwiński
This CORE set can be used as a reference when evaluating the reliability of high-throughput protein-protein interaction data sets, for development of prediction methods, as well as in the studies of the properties of protein interaction networks.
Algorithms on strings, trees, and sequences
1,723 Citations1997Dan Gusfield
Transformation-based error-driven learning and natural language processing: a case study in part-of-speech tagging
1,531 Citations1995Eric Brill
This paper describes a simple rule-based approach to automated learning of linguistic knowledge that has been shown for a number of tasks to capture information in a clearer and more direct fashion without a compromise in performance.
Maximum Entropy Markov Models for Information Extraction and Segmentation
1,333 Citations2000Andrew McCallum, Dayne Freitag +1 more
A new Markovian sequence model is presented that allows observations to be represented as arbitrary overlapping features (such as word, capitalization, formatting, part-of-speech), and defines the conditional probability of state sequences given observation sequences.
Text, speech and language technologyText Chunking Using Transformation-Based Learning
1,257 Citations1999Lance Ramshaw, Mitchell P. Marcus
This work has shown that the transformation-based learning approach can be applied at a higher level of textual interpretation for locating chunks in the tagged text, including non-recursive “baseNP” chunks.
Machine LearningMachine Learning for the Detection of Oil Spills in Satellite Radar Images
1,253 Citations1998Miroslav Kubát, Robert C. Holte +1 more
This case study relates issues as problem formulation, selection of evaluation measures, and data preparation to properties of the oil spill application, such as its imbalanced class distribution, that are shown to be common to many applications.
The Annals of Mathematical StatisticsGeneralized Iterative Scaling for Log-Linear Models
1,203 Citations1972J. N. Darroch, D. Ratcliff
Nucleic Acids ResearchDIP: the Database of Interacting Proteins
1,130 Citations2000Ioannis Xénarios
The Database of Interacting Proteins (DIP; http://dip.doe-mbi.ucla.edu) is a database that documents experimentally determined protein-protein interactions to provide the scientific community with a comprehensive and integrated tool for browsing and efficiently extracting information about protein interactions and interaction networks in biological processes.
Wrapper induction for information extraction
1,044 Citations1997Nicholas Kushmerick, Daniel S. Weld
This work introduces wrapper induction, a method for automatically constructing wrappers, and identifies hlrt, a wrapper class that is e(cid:14)ciently learnable, yet expressive enough to handle 48% of a recently surveyed sample of Internet resources.
Unsupervised Models for Named Entity Classification
817 Citations1999Michael Collins, Yoram Singer
It is shown that the use of unlabeled data can reduce the requirements for supervision to just 7 simple "seed" rules, gaining leverage from natural redundancy in the data.
Machine LearningAn Algorithm that Learns What's in a Name
784 Citations1999Daniel M. Bikel, Richard Schwartz +1 more
IdentiFinderTM, a hidden Markov model that learns to recognize and classify names, dates, times, and numerical quantities, is evaluated and is competitive with approaches based on handcrafted rules on mixed case text and superior on text where case information is not available.
Communications of the ACMInformation extraction
726 Citations1996Jim Cowie, Wendy G. Lehnert
A relatively new development—information extraction (IE)—is the subject of this article and can transform the raw material, refining and reducing it to a germ of the original text.
PubMedConstructing biological knowledge bases by extracting information from text sources.
597 Citations1999Mark Craven, Johan Kumlien
A research effort aimed at automatically mapping information from text sources into structured representations, such as knowledge bases, is begun, to use machine-learning methods to induce routines for extracting facts from text.
BioinformaticsGENIES: a natural-language processing system for the extraction of molecular pathways from journal articles
509 Citations2001Carol Friedman, Pauline Kra +3 more
A system is presented that extracts and structures information about cellular pathways from the biological literature in accordance with a knowledge model that was developed earlier and implemented by modifying an existing medical natural language processing system.
PubMedToward information extraction: identifying protein names from biological papers.
494 Citations1998Ken Fukuda, Akihiro Tamura +2 more
A new method of extracting material names, PROPER, using surface clue on character strings is proposed, which extracts material names in the sentence with 94.70% precision and 98.84% recall, regardless of whether it is already known or newly defined.
Relational Learning of Pattern-Match Rules for Information Extraction.
494 Citations1997Mary Elaine Califf, Raymond J. Mooney
EDGAR: Extraction of Drugs, Genes And Relations from the Biomedical Literature
425 Citations1999Thomas C. Rindflesch, Lorraine Tanabe +2 more
The mechanisms for automatically generating assertions about drugs and genes relevant to cancer and on a simple application, conceptual clustering of documents are reported on.
Nature GeneticsAssociation of genes to genetically inherited diseases using data mining
357 Citations2002Carolina Perez‐Iratxeta, Peer Bork +1 more
A scoring system for the possible functional relationships of human genes to 455 genetically inherited diseases that have been mapped to chromosomal regions without assignment of a particular gene indicates that for some diseases, the chance of identifying the underlying gene is higher.
Engineering Applications of Artificial IntelligenceProceedings of the 10th European conference on artificial intelligence
329 Citations1993Laurent Siklóssy, Jaak Tepandi
BioinformaticsTagging gene and protein names in biomedical text
310 Citations2002Lorraine Tanabe, W. John Wilbur
This work proposes to approach the detection of gene and protein names in scientific abstracts as part-of-speech tagging, the most basic form of linguistic corpus annotation, and demonstrates that this method can be applied to large sets of MEDLINE abstracts, without the need for special conditions or human experts to predetermine relevant subsets.
BioinformaticsMining literature for protein–protein interactions
295 Citations2001Edward M. Marcotte, Ioannis Xénarios +1 more
It is shown that the frequencies of words in Medline abstracts can be used to determine whether or not a given paper discusses protein-protein interactions, and the relevant information can be captured for the Database of Interacting Proteins.
Active Learning for Natural Language Parsing and Information Extraction
290 Citations1999Cynthia A. Thompson, Mary Elaine Califf +1 more
It is shown that active learning can signicantly reduce the number of annotated examples required to achieve a given level of performance for these complex tasks: semantic parsing and information extraction.
Use of support vector learning for chunk identification
285 Citations2000Taku Kudoh, Yūji Matsumoto
This paper investigates how SVMs with a very large number of features perform with the classification task of chunk labelling, CoNLL-2000 shared task, chunk identification.
Extracting the names of genes and gene products with a hidden Markov model
249 Citations2000Nigel Collier, Chikashi Nobata +1 more
A study into the use of a linear interpolating hidden Markov model (HMM) for the task of extracting technical terminology from MEDLINE abstracts and texts in the molecular-biology domain, the first stage in a system that will extract event information for automatically updating biology databases.
Automatic Extraction of Protein Interactions from Scientific Abstracts
247 Citations1999James Thomas, David Milward +3 more
This paper motivates the use of Information Extraction for gathering data on protein interactions, describes the customization of an existing IE system, SRI's Highlight, for this task and presents the results of an experiment on unseen Medline abstracts which show that customization to a new domain can be fast, reliable and cost-effective.
Ranking algorithms for named-entity extraction
241 Citations2001Michael Collins
Algorithms which rerank the top N hypotheses from a maximum-entropy tagger, the application being the recovery of named-entity boundaries in a corpus of web data, using the voted perceptron algorithm.
Boosted Wrapper Induction
232 Citations2000Dayne Freitag, Nicholas Kushmerick
This work describes an algorithm that learns simple, low-coverage wrapper-like extraction patterns, which it then applies to conventional information extraction problems using boosting, resulting in BWI, a trainable information extraction system with a strong precision bias and F1 performance better than state-of-the-art techniques in many domains.
Robust Relational Parsing over Biomedical Literature: Extracting Inhibit Relation
211 Citations2001James Pustejovsky, José M. Castaño +3 more
The design of a robust parser for identifying and extracting biomolecular relations from the biomedical literature is described, and the use of text-based anaphora resolution to enhance the results of argument binding in relational extraction is used.
Genome ResearchAssociating Genes with Gene Ontology Codes Using a Maximum Entropy Analysis of Biomedical Literature
200 Citations2002Soumya Raychaudhuri, Jeffrey T. Chang +2 more
It is concluded that statistical methods may be used to assign GO codes and may be useful for the difficult task of reassignment as terminology standards evolve over time.
Nucleic Acids ResearchDIP: The Database of Interacting Proteins: 2001 update
180 Citations2001Ioannis Xénarios
The Database of Interacting Proteins (DIP; http://dip.doe-mbi.ucla. edu) is a database that documents experimentally determined protein-protein interactions.
Two Applications of Information Extraction to Biological Science Journal Articles: Enzyme Interactions and Protein Structures
169 Citations1999Kevin Humphreys, George Demetriou +1 more
This paper describes how an information extraction system designed to participate in the MUC exercises has been modified for two bioinformatics applications: EMPathIE, concerns with enzyme and metabolic pathways; and PASTA, concerned with protein structure.
PubMedDetecting Gene Symbols and Names in Biological Texts: A First Step toward Pertinent Information Extraction.
151 Citations1998Denys Proux, François Rechenmann +3 more
A program for the identification of gene symbols and names inside sentences has been devised, made up of a series of sieves of different natures, lexical, morphological and semantic, to distinguish among the words of a sentence those which can only be potential gene symbols or names.
BIDIRECTIONAL INCREMENTAL PARSING FOR AUTOMATIC PATHWAY IDENTIFICATION WITH COMBINATORY CATEGORIAL GRAMMAR
142 Citations2000Jong Cheol Park, Hyun Sook Kim +1 more
An implemented system that utilizes combinatory categorical grammar known to be competent in modeling natural language, with a controlled mechanism for the parser to operate bidirectionally and incrementally is described.
Representing sentence structure in hidden Markov models for information extraction
123 Citations2001Soumya Ray, Mark Craven
An approach to representing the grammatical structure of sentences in the states of the model by using an objective function during HMM training which maximizes the ability of the learned models to identify the phrases of interest is proposed.
IEEE Intelligent Systems and their ApplicationsThe frame-based module of the SUISEKI information extraction system
118 Citations2002Christian Blaschke, Alfonso Valencia
CREATING KNOWLEDGE REPOSITORIES FROM BIOMEDICAL REPORTS: THE MEDSYNDIKATE TEXT MINING SYSTEM
97 Citations2001Udo Hahn, Martin Romacker +1 more
The strong demands MEDSYNDIKATE poses to the availability of expressive knowledge sources are accounted for by two alternative approaches to (semi)automatic ontology engineering.
Comparative and Functional GenomicsCan bibliographic pointers for known biological data be found automatically? Protein interactions as a case study
66 Citations2001Christian Blaschke, Alfonso Valencia
The DIP data set is proposed as a biological reference to benchmark IE systems for detecting previously known interactions and a positive finding is the capacity of the IE system to identify new relations between proteins, even in a set of proteins previously characterized by human experts.
PubMedA pragmatic information extraction strategy for gathering data on genetic interactions.
62 Citations2000Denys Proux, François Rechenmann +1 more
This work is using a combination of existing linguistic and knowledge processing tools to automatically extract information about gene interactions in the literature using a pragmatic strategy to perform information extraction from biologic texts.
BioinformaticsFinding relevant references to genes and proteins in Medlineusing a Bayesian approach
19 Citations2002Julie E. Leonard, Jeffrey B. Colombe +1 more
Results show that Medline entries discussing the same gene or protein have similar word usage, and that the method of assessing this similarity using EP values is valid, and enables an EP cutoff value to be determined that accurately and reproducibly balances precision and recall, allowing automated analysis of literature mining results.
A Theory-Refinement Approach to Information Extraction
14 Citations2001Tina Eliassi‐Rad, Jude Shavlik
This work investigates applying theory renemen t to the task of extracting information from text by using generate-and-test to address the IE task and produces candidate extractions by intelligently searching the space of possible extractions.
