login

From a children's first dictionary to a lexical knowledge base of conceptual graphs

Published 1 January 1997
Fred Popowich, Caroline Barrière
Citations33

TL;DR

The paper deals with small interventions that potentially have a strong impact on one of the society’s most urgent health issues when it comes to children, and apparently the flaws in the interpretation of the analyses by Zampollo et al. did not receive enough attention prior to publication.

Abstract

This thesis aims at building a Lexical Knowledge Base (LKB) that will be useful to a Natural Language Processing (NLP) system by extracting information from a Machine Readable Dictionary (MRD). Our source of knowledge is the American Heritage First Dictionary$\sp1$ (AHFD) which contains 1800 entries and is designed for children of age six to eight learning the structure and the basic vocabulary of their language. Using a children's dictionary allows us to restrict our vocabulary, but still work on general knowledge about day to day concepts and actions. Our Lexical Knowledge Base contains information extracted from the AHFD and represented using the Conceptual Graph (CG) formalism. The graph definitions explicitly give the information contained in all the noun and verb definitions from the AHFD. Each sentence of each definition is tagged, parsed and automatically transformed into a conceptual graph. The type hierarchy, extracted automatically from the definitions, groups all the nouns and verbs in the dictionary into a taxonomy. Covert categories will be discovered among the definitions and will complement the type hierarchy in its role for establishing concept similarity. Covert categories can be thought of as concepts not associated to a dictionary entry, such as writing instrument or device giving time. They allow grouping of words based on different criteria than a common hypernym, and therefore augment the space to explore for finding similarity among concepts. The relation hierarchy is built manually which groups into subclasses/superclasses the relations used in our CG representation of definitions. The relations can be prepositions such as in, on or with or deeper semantic relations such as part-of, material or instrument. Concept clusters are constructed automatically around a trigger word to put it into a larger context. Its graph representation is joined to the graph representations of other words in the dictionary that are related to it. The set of related words forms a concept cluster and their graph representation, showing all the relations between them and other related words, is a Concept Clustering Knowledge Graph. One important aspect of the thesis is the underlying thread of finding similarity through concept and graph comparison as a general way of processing information. The ideas presented in this thesis are implemented in a system ARC-Concept. We present and discuss the results obtained. ftn$\sp1$Copyright$\sp\circler$ 1994 by Houghton Mifflin Company. Reproduced by permission from THE AMERICAN HERITAGE FIRST DICTIONARY.

Keywords

Computer Science