Induced lexico-syntactic patterns improve information extraction from online medical forums
Journal of the American Medical Informatics AssociationPublished 27 June 2014Open access
Sonal Gupta, D. L. MacLean, Jeffrey Heer, Christopher D. Manning
Citations51
SJR quartileQ1
SJR score2.04
SNIP1.95
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This is the first paper to extract SC and DT entities from PAT by learning lexico-syntactic patterns from data annotated with seed dictionaries, and exhibits learning of informal terms often used in PAT but missing from typical dictionaries.
Abstract
Our entity extractor based on lexico-syntactic patterns is a successful and preferable technique for identifying specific entity types in PAT. To the best of our knowledge, this is the first paper to extract SC and DT entities from PAT. We exhibit learning of informal terms often used in PAT but missing from typical dictionaries.
Keywords
Computer ScienceBiochemistry, Genetics and Molecular Biology
Automatic acquisition of hyponyms from large text corpora
3,283 Citations1992Marti A. Hearst
A set of lexico-syntactic patterns that are easily recognizable, that occur frequently and across text genre boundaries, and that indisputably indicate the lexical relation of interest are identified.
Incorporating non-local information into information extraction systems by Gibbs sampling
3,035 Citations2005Jenny Rose Finkel, Trond Grenager +1 more
By using simulated annealing in place of Viterbi decoding in sequence models such as HMMs, CMMs, and CRFs, it is possible to incorporate non-local structure while preserving tractable inference.
Class-based n -gram models of natural language
2,893 Citations1992Peter F. Brown, P.V. deSouza +3 more
This work addresses the problem of predicting a word from previous words in a sample of text and discusses n-gram models based on classes of words, finding that these models are able to extract classes that have the flavor of either syntactically based groupings or semanticallybased groupings, depending on the nature of the underlying statistics.
PubMedEffective mapping of biomedical text to the UMLS Metathesaurus: the MetaMap program.
2,011 Citations2001Alan R. Aronson
This study sought to determine whether an association study using information contained in clinical notes could identify known and potentially novel risk factors for nonadherence to antihypertensive medications.
Journal of the American Medical Informatics AssociationAn overview of MetaMap: historical perspective and recent advances
1,540 Citations2010Alan R. Aronson, François-Michel Lang
This study reports on MetaMap's evolution over more than a decade, concentrating on those features arising out of the research needs of the biomedical informatics community both within and outside of the National Library of Medicine.
Diabetes CareCinnamon Improves Glucose and Lipids of People With Type 2 Diabetes
980 Citations2003Alam Khan, Mahpara Safdar +3 more
The results of this study demonstrate that intake of 1, 3, or 6 g of cinnamon per day reduces serum glucose, triglyceride, LDL cholesterol, and total cholesterol in people with type 2 diabetes and suggest that the inclusion of cinnamon in the diet of people withtype 2 diabetes will reduce risk factors associated with diabetes and cardiovascular diseases.
Clinical Infectious DiseasesGoogle Trends: A Web‐Based Tool for Real‐Time Surveillance of Disease Outbreaks
836 Citations2009Herman Carneiro, Eleftherios Mylonakis
Google Flu Trends can detect regional outbreaks of influenza 7-10 days before conventional Centers for Disease Control and Prevention surveillance systems and should work with public health care practitioners to develop specialized tools, using Google Flu Trends as a blueprint, to track infectious diseases.
International Journal on Digital LibrariesAutomatic recognition of multi-word terms:. the C-value/NC-value method
818 Citations2000Katerina T. Frantzi, Sophia Ananiadou +1 more
This paper presents a domain-independent method for the automatic extraction of multi-word terms, from machine-readable special language corpora, using C-value/NC-value, which enhances the common statistical measure of frequency of occurrence for term extraction, making it sensitive to a particular type ofMulti- word terms, the nested terms.
A bootstrapping method for learning semantic lexicons using extraction pattern contexts
369 Citations2002M. Thelen, Ellen Riloff
The semantic lexicons produced by Basilisk have higher precision than those produced by previous techniques, with several categories showing substantial improvement.
Nature BiotechnologyAccelerated clinical discovery using self-reported patient data collected online and a patient-matching algorithm
367 Citations2011Paul Wicks, Timothy E. Vaughan +2 more
DSpace@MIT (Massachusetts Institute of Technology)Semi-Supervised Learning for Natural Language
304 Citations2005Percy Liang
This thesis focuses on two segmentation tasks, named-entity recognition and Chinese word segmentation, and shows that features derived from unlabeled data substantially improves performance, both in terms of reducing the amount of labeled data needed to achieve a certain performance level and in termsof reducing the error using a fixed amount of labeling data.
Journal of the American Medical Informatics AssociationExploring and Developing Consumer Health Vocabularies
292 Citations2005Qiang Zeng, Tony Tse
This paper presents the point of view that CHV development is practical and necessary for extending research on informatics-based tools to facilitate consumer health information seeking, retrieval, and understanding and briefly describes a distributed, bottom-up approach.
Towards Internet-Age Pharmacovigilance: Extracting Adverse Drug Reactions from User Posts in Health-Related Social Networks
257 Citations2010Robert Leaman, L Wojtulewicz +4 more
It is concluded that user comments pose a significant natural language processing challenge, but do contain useful extractable information which merits further exploration and is evaluated on a manually annotated set of user comments with promising performance.
PubMedThe open biomedical annotator.
234 Citations2009Clément Jonquet, Nigam H. Shah +1 more
The Open Biomedical Annotator (OBA) is presented, an ontology-based Web service that annotates public datasets with biomedical ontology concepts based on their textual metadata ( www.bioontology.org).
Journal of the American Medical Informatics AssociationA novel signal detection algorithm for identifying hidden drug-drug interactions in adverse event reports
209 Citations2011Nicholas P. Tatonetti, Guy Haskin Fernald +1 more
This work presents a novel method to identify latent drug interaction signals in spontaneous reporting systems by using side effect profiles to infer the presence of unreported adverse events.
Journal of the American Medical Informatics AssociationWeb-scale pharmacovigilance: listening to signals from the crowd
201 Citations2013Ryen W. White, Nicholas P. Tatonetti +3 more
It is found that anonymized signals on drug interactions can be mined from search logs, and logs of the search activities of populations of computer users can contribute to drug safety surveillance.
Journal of the American Medical Informatics AssociationUsing rule-based natural language processing to improve disease normalization in biomedical text
136 Citations2012Kang Ning, Bharat Singh +3 more
The added value of NLP for the recognition and normalization of diseases with MetaMap and Peregrine is shown and the NLP module is general and can be applied in combination with any concept normalization system.
Diabetes CareHypoglycemic Effect of <i>Opuntia streptacantha</i> Lemaire in NIDDM
111 Citations1988A C Frati-Munari, Blanca E Gordillo +2 more
This study shows that the stems of O. streptacantha Lem.
PubMedA study of biomedical concept identification: MetaMap vs. people.
94 Citations2003Wanda Pratt, Meliha Yetisgen-Yildiz
A study compares MetaMap's performance against that of six people and finds that for those concepts that subjects generally agreed on, MetaMap was able to identify most concepts, if they were represented in the UMLS.
PubMedPatientsLikeMe: Consumer health vocabulary as a folksonomy.
80 Citations2008Catherine Arnott Smith, Paul Wicks
Analysis of the failed matches of PatientsLikeMe symptom terms reveals challenges for online patient communication, not only with healthcare professionals, but with other patients.
Journal of the American Medical Informatics AssociationIdentifying medical terms in patient-authored text: a crowdsourcing-based approach
79 Citations2013Diana MacLean, Jeffrey Heer
It is demonstrated that crowdsourcing PAT medical term identification tasks to non-experts is a viable method for creating large, accurately-labeled PAT datasets; moreover, such datasets can be used to train classifiers that outperform existing medicalterm identification tools.
PubMedUnsupervised method for automatic construction of a disease dictionary from a large free text collection.
50 Citations2008Rong Xu, Kaustubh Supekar +3 more
An automated, unsupervised, iterative pattern learning approach for constructing a comprehensive medical dictionary of disease terms from randomized clinical trial (RCT) abstracts is developed, and different ranking methods for automatically extracting con-textual patterns and concept terms are compared.
Journal of the American Medical Informatics AssociationAutomated identification of drug and food allergies entered using non-standard terminology
35 Citations2013Richard H. Epstein, Paul St. Jacques +4 more
A high performing, easily maintained algorithm can successfully identify medication and food allergies from free text entries in EHR systems.
