POSBIOTM--NER: a trainable biomedical named-entity recognition system
Computer applications in the biosciencesPublished 6 April 2005Open access
Yu Song, Edward Kim, G. G. Lee, Byoung-Kee Yi
Citations35
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
SUMMARY: POSBIOTM-NER is a trainable biomedical named-entity recognition system. POSBIOTM-NER can be automatically trained and adapted to new datasets without performance degradation, using CRF (conditional random field) machine learning techniques and automatic linguistic feature analysis. Currently, we have trained our system on three different datasets. GENIA-NER was trained based on GENIA Corpus, GENE-NER based on BioCreative data and GPCR-NER based on our own POSBIOTM/NE corpus, respectively, which would be used in GPCR-related pathway extraction.
Keywords
Computer ScienceBiochemistry, Genetics and Molecular Biology
ScholarlyCommons (University of Pennsylvania)Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data
12,978 Citations2001John Lafferty, Andrew McCallum +1 more
This work presents iterative parameter estimation algorithms for conditional random fields and compares the performance of the resulting models to HMMs and MEMMs on synthetic and natural-language data.
BioinformaticsGENIA corpus—a semantically annotated corpus for bio-textmining
1,243 Citations2003JD Kim, Tomoko Ohta +2 more
Introduction to the bio-entity recognition task at JNLPBA
579 Citations2004Jin-Dong Kim, Tomoko Ohta +3 more
The JNLPBA shared task of bio-entity recognition using an extended version of the GENIA version 3 named entity corpus of MEDLINE abstracts is described and a general discussion of the approaches taken by participating systems is presented.
BioinformaticsRecognizing names in biomedical texts: a machine learning approach
253 Citations2004Guodong Zhou, Jie Zhang +3 more
The PowerBioNE system is the first system which deals with the cascaded entity name phenomenon and the HMM and the k-NN algorithm outperform other models, such as back-off HMM, linear interpolated H MM, support vector machines, C4.5 rules and RIPPER, by effectively capturing the local context dependency and resolving the data sparseness problem.
Extracting the names of genes and gene products with a hidden Markov model
249 Citations2000Nigel Collier, Chikashi Nobata +1 more
A study into the use of a linear interpolating hidden Markov model (HMM) for the task of extracting technical terminology from MEDLINE abstracts and texts in the molecular-biology domain, the first stage in a system that will extract event information for automatically updating biology databases.
BMC BioinformaticsExploring the boundaries: gene and protein identification in biomedical text
108 Citations2005Jenny Rose Finkel, Shipra Dingare +4 more
Central contributions are rich use of features derived from the training data at multiple levels of granularity, a focus on correctly identifying entity boundaries, and the innovative use of several external knowledge sources including full MEDLINE abstracts and web searches.
