Toward semantic understanding: an approach based on information extraction ontologies
Published 1 January 2004
David W. Embley
Citations112
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This paper proffers the use of information-extraction ontologies as an approach that may lead to semantic understanding.
Abstract
Information is ubiquitous, and we are flooded with more than we can process. Somehow, we must rely less on visual processing, point-and-click navigation, and manual decision making and more on computer sifting and organization of information and automated negotiation and decision making. A resolution of these problems requires software with semantic understanding---a grand challenge of our time.
Keywords
Computer Science
Elsevier eBooksA FRAMEWORK FOR REPRESENTING KNOWLEDGE
4,545 Citations1988Marvin Minsky
The enormous problem of the volume of background common sense knowledge required to understand even very simple natural language texts is discussed and it is suggested that networks of frames are a reasonable approach to represent such knowledge.
Computer NetworksFocused crawling: a new approach to topic-specific Web resource discovery
1,492 Citations1999Soumen Chakrabarti, Martin van den Berg +1 more
A new hypertext resource discovery system called a Focused Crawler that is robust against large perturbations in the starting set of URLs, and capable of exploring out and discovering valuable resources that are dozens of links away from the start set, while carefully pruning the millions of pages that may lie within this same radius.
KQML as an agent communication language
1,470 Citations1994Tim Finin, Richard Fritzson +2 more
The design of and experimentation with the Knowledge Query and Manipulation Language (KQML), a new language and protocol for exchanging information and knowledge, which is aimed at developing techniques and methodology for building large-scale knowledge bases which are sharable and reusable.
Generic Schema Matching with Cupid
1,271 Citations2001Jayant Madhavan, Philip A. Bernstein +1 more
This paper proposes a new algorithm, Cupid, that discovers mappings between schema elements based on their names, data types, constraints, and schema structure, using a broader set of techniques than past approaches.
Communications of the ACMSoftware agents
1,089 Citations1994Michael Genesereth, Steven P. Ketchpel
A relatively loose notion of an agent as a self-contained program capable of controlling its own decision making and acting, based on its perception of its environment, in pursuit of one or more objectives will be used here.
Communications of the ACMEvaluating ontological decisions with OntoClean
765 Citations2002Nicola Guarino, Christopher Welty
Explosing common misuses of the subsumption relationship and the formal basis for why they are wrong and how to stop them.
Programming with abstract data types
736 Citations1974Barbara Liskov, Stephen N. Zilles
An approach which allows the set of built-in abstractions to be augmented when the need for a new data abstraction is discovered and is an outgrowth of work on designing a language for structured programming.
Reconciling schemas of disparate data sources
733 Citations2001AnHai Doan, Pedro Domingos +1 more
LSD is a system that employs and extends current machine-learning techniques to semi-automatically find semantic mappings between the source schemas and the mediated schema, and its architecture is extensible to additional learners that may exploit new kinds of information.
ACM SIGMOD RecordA brief survey of web data extraction tools
690 Citations2002Alberto H. F. Laender, Berthier Ribeiro‐Neto +2 more
A taxonomy for characterizing Web data extraction fools is proposed, a survey of major web data extraction tools described in the literature is briefly surveyed, and a qualitative analysis of them is provided.
John Wiley & Sons, Inc. eBooksTowards the Semantic Web: Ontology-driven Knowledge Management
564 Citations2002John Davies, Dieter Fensel +1 more
Towards theSemantic Web focuses on the application of Semantic Web technology and ontologies in particular to electronically available information to improve the quality of knowledge management in large and distributed organizations.
ACM Transactions on Database SystemsAn ontological analysis of the relationship construct in conceptual modeling
415 Citations1999Yair Wand, Veda C. Storey +1 more
This analysis uses ontology, which is the branch of philosophy dealing with models of reality, to analyze the meaning of common conceptual modeling constructs and derives rules for the use of relationships in entity-relationship conceptual modeling.
Communications of the ACMA collaborative approach to ontology design
388 Citations2002Clyde W. Holsapple, K.D. Joshi
Creating a general ontology characterizing the conduct of knowledge management and its implications for knowledge management is described.
Data & Knowledge EngineeringSupporting ontological analysis of taxonomic relationships
356 Citations2001Christopher Welty, Nicola Guarino
This work has adopted several notions from the philosophical practice of formal ontology, and adapted them for use in information systems, to provide a solid logical framework within which the properties that form a taxonomy can be analyzed.
ACM SIGMOD RecordData modelling versus ontology engineering
353 Citations2002Peter Spyns, Robert Meersman +1 more
The DOGMA ontology engineering approach is introduced that separates "atomic" conceptual relations from "predicative" domain rules and a layer of "relatively generic" ontological commitments that hold the domain rules.
Text Information Retrieval Systems
340 Citations1992Charles T. Meadow, Donald H. Kraft +1 more
This book covers the nature of information, how it is organized for use by a computer, how search functions are carried out, and some of the theory underlying these functions, and how retrieved items, users, and complete systems are evaluated.
Data & Knowledge EngineeringConceptual-model-based data extraction from multiple-record Web pages
311 Citations1999David W. Embley, Douglas M. Campbell +5 more
Experiments show that it is possible to achieve good recall and precision ratios for documents that are rich in recognizable constants and narrow in ontological breadth in a conceptual-modeling approach.
ACM SIGMOD RecordRecord-boundary discovery in Web documents
234 Citations1999David W. Embley, Yi Jiang +1 more
BT Technology JournalFIPA — Towards a Standard for Software Agents
227 Citations1998P. D. O'Brien, Richard Nicol
FIPA (Foundation for Intelligent Physical Agents) is described and an overview and guide to the FIPA97 specification is provided, which discusses how FIPa relates to other agent standards activities and concludes with FipA's plans for 1998.
ACM Transactions on Information SystemsInformation extraction as a basis for high-precision text classification
206 Citations1994Ellen Riloff, Wendy G. Lehnert
An approach to text classification that represents a compromise between traditional word-based techniques and in-depth natural language processing and an automated method for empirically deriving appropriate threshold values is described.
A machine learning approach to building domain-specific search engines
183 Citations1999Andrew McCallum, Kamal Nigam +2 more
The use of machine learning techniques are proposed to greatly automate the creation and maintenance of domain-specific search engines and new research in reinforcement learning, text classification and information extraction that enables efficient spidering, populates topic hierarchies, and identifies informative text segments is described.
Building Domain-Specific Search Engines with Machine Learning Techniques
130 Citations1999Andrew McCallum, Kamal Nigam +2 more
New research in reinforcement learning, information extraction and text classification that enables efficient spidering, identifying informative text segments, and populating topic hierarchies is described.
Data & Knowledge EngineeringAutomating the extraction of data from HTML tables with unknown structure
81 Citations2004David W. Embley, Cui Tao +1 more
Experimental results show that the solution entails elements of table understanding, data integration, and wrapper creation and can successfully locate data of interest in tables and map the data from source HTML tables with unknown structure to a given target database schema.
Nano LettersSuperimposed Information for the Internet.
66 Citations1999David Maier, Lois Delcambre
Two techniques to enhance the Q of a nanomechanical beam are explored, showing that by embedding a nanobeam in a 1D phononic crystal (PnC), it is possible to localize its flexural motion and shield it against radiation loss, and taking advantage of the mode-shape dependence of stress-induced "loss dilution.
Programming with data frames for everyday data items
43 Citations1980David W. Embley
Today's programmers confront the drudgery of writing routines to recognize, validate, transform, store, retrieve, manipulate, and display these items and also the challenge to develop user-friendly data-entry systems and insure data integrity.
Information SystemsExtracting information from heterogeneous information sources using ontologically specified target views
41 Citations2003Joachim Biskup, David W. Embley
This paper proposes a framework for addressing the issues involved in deluging volumes of structured and unstructured data contained in databases, data warehouses, and the global Internet, and is able to prove that when a source has a valid interpretation, the generated mapping produces avalid interpretation for the part of the target loaded from the source.
Ontology generation from tables
38 Citations2004Yuri A. Tijerino, David W. Embley +2 more
A new framework to table understanding is developed that applies an ontology-based conceptual modeling extraction approach to understand a table's structure and conceptual content to the extent possible and discover the constraints that hold between concepts extracted from the table.
Source discovery and schema mapping for data integration
19 Citations2003David W. Embley, Li Xu
A Target-based Integration Query System (TIQS) is offered as an alternative point of view that is neither GAV nor LAV and it is proven that query reformulation in TIQS reduces to rule unfolding and the reformulated user queries extract all the query answers available from sources with respect to the definition of TIZS for the proposed queries.
Record Location and Reconfiguration in Unstructured Multiple-Record Web Documents
18 Citations2000David W. Embley, Li Xu
Lecture notes in computer scienceRecognizing Ontology-Applicable Multiple-Record Web Documents
14 Citations2001David W. Embley, Yiu‐Kai Ng +1 more
A technique for recognizing which multiplere-cord Web documents apply to an ontologically specified application and, based on machine-learned rules over these heuristic measurements, determines whether a Web document is applicable for a given ontology.
OIL and DAML + OIL: Ontology Languages for the Semantic Web
13 Citations2002Dieter Fensel, Frank van Harmelen +1 more
ACM SIGMOD RecordJim Gray speaks out
12 Citations2003Marianne Winslett
Jim Gray, who received a Turing award in 1998 for his contributions to computer science, especially in the area of transaction processing, is interviewed in Madison, Wisconsin, the site of the 2002 PODS and SIGMOD conferences.
ScholarsArchive (Brigham Young University)Ontology-Based Extraction of RDF Data from the World Wide Web
11 Citations2003Timothy Adam Chartrand
This project applies existing information-extraction techniques to extract information from the WWW based on a Semantic Web ontology to produceSemantic Web data with respect to that ontology, and experiments with ontologies in four application domains show that this approach can indeed extract SemanticWeb data from theWWW with precision and recall similar to that achieved by the underlying information extraction system.
Improving the quality of systems and domain analysis through object class congruency
5 Citations2002Stephen W. Clyde, David W. Embley +1 more
A new concept for assessing the quality of object classes in analysis models, called object-class congruency, is formally defined and discussed, and two semantic-preserving transformations that convert incongruent classes into congruent classes are given.
