login

A Distributional Semantics Approach to Simultaneous Recognition of Multiple Classes of Named Entities

Lecture notes in computer sciencePublished 1 January 2010
Siddhartha Jonnalagadda, Robert Leaman, Trevor Cohen, Graciela Gonzalez‐Hernandez
Citations13
SJR quartileQ2
SJR score0.35
SNIP0.55

TL;DR

Sahlgren et al's permutation-based variant of the Random Indexing model is used to create a scalable and efficient system to simultaneously recognize multiple entity classes mentioned in natural language, which is validated on the GENIA corpus.

Abstract

Named Entity Recognition and Classification is being studied for last two decades. Since semantic features take huge amount of training time and are slow in inference, the existing tools apply features and rules mainly at the word level or use lexicons. Recent advances in distributional semantics allow us to efficiently create paradigmatic models that encode word order. We used Sahlgren et al's permutation-based variant of the Random Indexing model to create a scalable and efficient system to simultaneously recognize multiple entity classes mentioned in natural language, which is validated on the GENIA corpus which has annotations for 46 biomedical entity classes and supports nested entities. Using distributional semantics features only, it achieves an overall micro-averaged F-measure of 67.3% based on fragment matching with performance ranging from 7.4% for "DNA substructure" to 80.7% for "Bioentity".

Keywords

Computer ScienceBiochemistry, Genetics and Molecular Biology