Web mining research
ACM SIGKDD Explorations NewsletterPublished 1 June 2000Open access
Raymond Kosala, Hendrik Blockeel
Citations1,439
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This paper surveys the research in the area of Web mining, point out some confusions regarded the usage of the term Web mining and suggest three Web mining categories, which are then situate some of the research with respect to these three categories.
Abstract
status: Published
Keywords
Computer SciencePhysics and Astronomy
Computer Networks and ISDN SystemsThe anatomy of a large-scale hypertextual Web search engine
15,828 Citations1998Sergey Brin, Lawrence M. Page
This paper provides an in-depth description of Google, a prototype of a large-scale search engine which makes heavy use of the structure present in hypertext and looks at the problem of how to effectively deal with uncontrolled hypertext collections where anyone can publish anything they want.
Journal of the American Society for Information ScienceIndexing by latent semantic analysis
12,677 Citations1990Scott Deerwester, Susan Dumais +3 more
Journal of the ACMAuthoritative sources in a hyperlinked environment
9,060 Citations1999Jon Kleinberg
This work proposes and test an algorithmic formulation of the notion of authority, based on the relationship between a set of relevant authoritative pages and the set of “hub pages” that join them together in the link structure, and has connections to the eigenvectors of certain matrices associated with the link graph.
Communications of the ACMFab
2,927 Citations1997Marko Balabanović, Yoav Shoham
It is explained how a hybrid system can incorporate the advantages of both methods while inheriting the disadvantages of neither, and how the particular design of the Fab architecture brings two additional benefits.
Communications of the ACMAgents that reduce work and information overload
2,321 Citations1994Pattie Maes
Results from several prototype agents that have been built using an approach to building interface agents are presented, including agents that provide personalized assistance with meeting scheduling, email handling, electronic news filtering, and selection of entertainment.
ACM SIGKDD Explorations NewsletterWeb usage mining
2,113 Citations2000Jaideep Srivastava, Robert Cooley +2 more
A detailed taxonomy of the work in this area, including research efforts as well as commercial offerings is provided, and a brief overview of the WebSIFT system as an example of a prototypical Web usage mining system is given.
Symposium on Discrete AlgorithmsAuthoritative sources in a hyperlinked environment
1,836 Citations1998Jon Kleinberg
Inductive learning algorithms and representations for text categorization
1,465 Citations1998Susan Dumais, John Platt +2 more
A comparison of the effectiveness of five different automatic learning algorithms for text categorization in terms of learning speed, realtime classification speed, and classification accuracy is compared.
Knowledge and Information SystemsData Preparation for Mining World Wide Web Browsing Patterns
1,464 Citations1999Robert Cooley, Bamshad Mobasher +1 more
This paper presents several data preparation techniques in order to identify unique users and user sessions and Transactions identified by the proposed methods are used to discover association rules from real world data using the WEBMINER system.
NatureAccessibility of information on the web
1,356 Citations1999Steve Lawrence, C. Lee Giles
As the web becomes a major communications medium, the data on it must be made more accessible, and search engines need to make the data more accessible.
Web mining: information and pattern discovery on the World Wide Web
1,170 Citations2002Robert Cooley, Bamshad Mobasher +1 more
This paper defines Web mining and presents an overview of the various research issues, techniques, and development efforts, and briefly describes WEBMINER, a system for Web usage mining, and concludes the paper by listing research issues.
DataGuides: Enabling Query Formulation and Optimization in Semistructured Databases
1,147 Citations1997Roy Goldman, Jennifer Widom
The theoretical foundations of DataGuides are presented along with an algorithm for their creation and an overview of incremental maintenance, and performance results based on the implementation of dataGuides in the Lore DBMS for semistructured data are provided.
International Journal on Digital LibrariesThe Lorel query language for semistructured data
1,067 Citations1997Serge Abiteboul, Dallan Quass +3 more
The main novelties of the Lorel language are the extensive use of coercion to relieve the user from the strict typing of OQL, which is inappropriate for semistructured data; and powerful path expressions, which permit a flexible form of declarative navigational access and are particularly suitable when the details of the structure are not known to the user.
Wrapper induction for information extraction
1,044 Citations1997Nicholas Kushmerick, Daniel S. Weld
This work introduces wrapper induction, a method for automatically constructing wrappers, and identifies hlrt, a wrapper class that is e(cid:14)ciently learnable, yet expressive enough to handle 48% of a recently surveyed sample of Internet resources.
Computer NetworksTrawling the Web for emerging cyber-communities
1,010 Citations1999Ravi Kumar, Prabhakar Raghavan +2 more
The subject of this paper is the systematic enumeration of over 100,000 emerging communities from a Web crawl, motivating a graph-theoretic approach to locating such communities, and describing the algorithms and algorithmic engineering necessary to find structures that subscribe to this notion.
Research Showcase @ Carnegie Mellon University (Carnegie Mellon University)Topic Detection and Tracking Pilot Study Final Report
951 Citations2018James Allan, Jaime Carbonell +3 more
The TSIMMIS project: Integration of heterogeneous information sources
930 Citations1994Sudarshan S. Chawathe, Héctor García-Molina +5 more
Machine LearningLearning Information Extraction Rules for Semi-Structured and Free Text
928 Citations1999Stephen Soderland
WHISK is designed to handle text styles ranging from highly structured to free text, including text that is neither rigidly formatted nor composed of grammatical sentences, and can also handle extraction from free text such as news stories.
Untangling text data mining
856 Citations1999Marti A. Hearst
Data mining, information access, and corpus-based computational linguistics are defined and the relationship of these to text data mining is discussed, and the intent behind these contrasts is to draw attention to exciting new kinds of problems for computational linguists.
Knowledge discovery and data mining: towards a unifying framework
847 Citations1996Usama M. Fayyad, Gregory Piatetsky-Shapiro +1 more
The KDD process and basic data mining algorithms are defined, links between data mining, knowledge discovery, and other related fields are described, and an analysis of challenges facing practitioners in the field is analyzed.
Enhanced hypertext categorization using hyperlinks
775 Citations1998Soumen Chakrabarti, Byron Dom +1 more
This work has developed a text classifier that misclassified only 13% of the documents in the well-known Reuters benchmark; this was comparable to the best results ever obtained and its technique also adapts gracefully to the fraction of neighboring documents having known topics.
Using Maximum Entropy for Text Classification
756 Citations1999Kamal Nigam, John Lafferty +1 more
This paper uses maximum entropy techniques for text classification by estimating the conditional distribution of the class variable given the document by comparing accuracy to naive Bayes and showing that maximum entropy is sometimes significantly better, but also sometimes worse.
Communications of the ACMInformation extraction
726 Citations1996Jim Cowie, Wendy G. Lehnert
A relatively new development—information extraction (IE)—is the subject of this article and can transform the raw material, refining and reducing it to a germ of the original text.
Web Watcher: A Tour Guide for the World Wide Web.
681 Citations1997Thorsten Joachims, Dayne Freitag +1 more
The learning algorithms used by WebWatcher, experimental results showing their e ectiveness, and lessons learned from this case study in Web tour guide agents are described.
ACM SIGIR ForumImproved Algorithms for Topic Distillation in a Hyperlinked Environment
677 Citations2017Krishna Bharat, Monika Henzinger
This paper addresses the problem of topic distillation on the World Wide Web, namely, given a typical user query to find quality documents related to the query topic, by augmenting a previous connectivity analysis based algorithm with content analysis.
Communications of the ACMMachine learning and data mining
675 Citations1999Tom M. Mitchell
The eld of data mining addresses the question of how best to use this historical data to discover general regularities and to improve future decisions.
Learning to extract symbolic knowledge from the World Wide Web
675 Citations1998Mark Craven, Dan DiPasquo +5 more
The goal of the research described here is to automatically create a computer understandable world wide knowledge base whose content mirrors that of the World Wide Web, and several machine learning algorithms for this task are described.
ACM SIGIR ForumOn-Line New Event Detection and Tracking
653 Citations2017James Allan, Ron Papka +1 more
Research Commons (University of Waikato)Domain-specific keyphrase extraction
623 Citations1999Eibe Frank, Gordon W. Paynter +3 more
This paper shows that a simple procedure for keyphrase extraction based on the naive Bayes learning scheme performs comparably to the state of the art, and explains how this procedure's performance can be boosted by automatically tailoring the extraction process to the particular document collection at hand.
ACM SIGMOD RecordDatabase techniques for the World-Wide Web
560 Citations1998Daniela Florescu, Alon Y. Levy +1 more
The primary goal of this survey is to classify the different tasks to which database concepts have been applied, and to emphasize the technical innovations that were required to do so.
A query language and optimization techniques for unstructured data
535 Citations1996Peter Buneman, Susan B. Davidson +2 more
Here a simple language UnQL is proposed for querying data organized as a rooted, edge-labeled graph and it is shown that known optimization techniques for operators on flat relations apply to the "horizontal" dimension of UnQL.
On-line new event detection and tracking
524 Citations1998James Allan, Ron Papka +1 more
The approach to detection uses a single pass clustering algo-rithm and a novel thresholding model that incorporates the properties of events as a major component and the value of \surprising" features that have unusual occurrence characteristics are discussed.
Communications of the ACMThe World-Wide Web
524 Citations1996Oren Etzioni
Information on the Web is sufficiently structured to facilitate effective Web mining, according to the structured Web hypothesis, and preliminary Web mining successes are surveyed and directions for future work are suggested.
ComputerMining the Web's link structure
502 Citations1999Soumen Chakrabarti, Byron Dom +6 more
Clever is a search engine that analyzes hyperlinks to uncover two types of pages: authorities, which provide the best source of information on a given topic; and hubs, which provides collections of links to authorities.
NeurocomputingWEBSOM – Self-organizing maps of document collections
493 Citations1998Samuel Kaski, Timo Honkela +2 more
Special consideration is given to the computation of very large document maps which is possible with general-purpose computers if the dimensionality of the word category histograms is first reduced with a random mapping method and if computationally efficient algorithms are used in computing the SOMs.
Semistructured data
443 Citations1997Peter Buneman
A number of issues surrounding semi-structured data are covered: finding a concise formulation, building a sufficiently expressive language for querying and transformation, and opti-mizat,ion problems.
Improved algorithms for topic distillation in a hyperlinked environment
439 Citations1998Krishna Bharat, Monika Henzinger
This paper addresses the problem of topic distillation on the World Wide Web, namely, given a typical user query to find quality documents related to the query topic, by augmenting a previous connectivity analysis based algorithm with content analysis.
Knowledge discovery in Textual Databases (KDT)
437 Citations1995Ronen Feldman, Ido Dagan
This research combines the KDD and text categorization paradigms and suggests advances to the state of the art in both areas.
Feature Selection for Unbalanced Class Distribution and Naive Bayes
432 Citations1999Dunja Mladenić, Marko Grobelnik
This paper describes an approach to feature subset selection that takes into account problem speciics and learning algorithm characteristics, and shows that considering domain and algorithm characteristics signiicantly improves the results of classiication.
Information SystemsGenerating finite-state transducers for semi-structured data extraction from the Web
414 Citations1998Chun‐Nan Hsu, Ming-Tzung Dung
This paper presents SoftMealy, a novel wrapper representation formalism based on a finite-state transducer and contextual rules that can wrap a wide range of semistructured Web pages because FSTs can encode each different attribute permutation as a path.
Courses and lecturesA Hybrid User Model for News Story Classification
352 Citations1999Daniel Billsus, Michael J. Pazzani
An intelligent agent designed to compile a daily news program for individual users, which motivates the use of a multi-strategy machine learning approach that allows for the induction of user models that consist of separate models for long-term and short-term interests.
IEEE Intelligent Systems and their ApplicationsLearning approaches for detecting and tracking news events
338 Citations1999Yi Yang, Jaime Carbonell +4 more
The authors extend existing supervised-learning and unsupervised-clustering algorithms to allow document classification based on the information content and temporal aspects of news events to be classified using manually segmented documents.
Advances in Inductive Logic Programming
313 Citations1996Luc De Raedt
A state-of-the-art overview of Inductive Logic Programming is provided, based on the succesful ESPRIT basic research project no. 6020, which can be used as a thorough introduction to the field.
Extracting Semistructured Information from the Web.
304 Citations1997J. Hammer, Héctor García-Molina +3 more
A configurable tool for extracting semistructured data from a set of HTML pages and for converting the extracted information into database objects and various ways of improving the functionality of the current prototype are described.
Feature Engineering for Text Classification
300 Citations1999Sam Scott, Stan Matwin
More sophisticated Natural Language Processing techniques need to be developed before better text representations can be produced for classification.
ACM SIGKDD Explorations NewsletterData mining for hypertext
291 Citations2000Soumen Chakrabarti
Recent advances in learning and mining problems related to hypertext in general and the Web in particular are surveyed and the continuum of supervised to semi-supervised to unsupervised learning problems is reviewed.
Information Extraction with HMM Structures Learned by Stochastic Optimization
286 Citations2000Dayne Freitag, Andrew McCallum
This paper demonstrates that extraction accuracy strongly depends on the selection of structure, and presents an algorithm for automatically finding good structures by stochastic optimization, which finds HMM models that almost always out-perform a fixed model, and have superior average performance across tasks.
IEEE Intelligent Systems and their ApplicationsText-learning and related intelligent agents: a survey
280 Citations1999Dunja Mladenić
Personal WebWatcher is described, a content-based intelligent agent that uses text-learning for user-customized Web browsing that focuses on three key criteria: what representation the particular application uses for documents, how it selects features, and what learning algorithm it uses.
Information extraction from HTML: application of a general machine learning approach
237 Citations1998Dayne Freitag
This work shows how information extraction can be cast as a standard machine learning problem, and argues for the suitability of relational learning in solving it, and the implementation of a general-purpose relational learner for information extraction, SRV.
Information Extraction with HMMs and Shrinkage
235 Citations1999Dayne Freitag, Andrew Kachites McCallum
A statistical technique called shrinkage is used that significantly improves parameter estimation of the HMM emission probabilities in the face of sparse training data and the resulting HMM outperforms a state-of-the-art rule-learning system.
Extraction Patterns for Information Extraction Tasks: A Survey
227 Citations1999Ion Muslea
This paper surveys the various types of extraction patterns that are generated by machine learning algorithms and identifies three main categories of patterns, which cover a variety of application domains, and compares and contrast the patterns from each category.
Extracting schema from semistructured data
224 Citations1998Svetlozar Nestorov, Serge Abiteboul +1 more
It is established that the general problem of finding an optimal form of semistructured data based on labeled, directed graphs is NP-hard, but some heuristics and techniques based on clustering that allow efficient and near-optimal treatment of the problem are presented.
Using Reinforcement Learning to Spider the Web Efficiently
213 Citations1999Jason D. M. Rennie, Andrew McCallum
This paper presents an algorithm for learning a value function that maps hyperlinks to future discounted reward using a naive Bayes text classifier and shows a threefold improvement in spidering efficiency over traditional breadth-first search, and up to a two-fold improvement over reinforcement learning with immediate reward.
A declarative language for querying and restructuring the Web
212 Citations2002Laks V. S. Lakshmanan, Fereidoon Sadri +1 more
This work develops a simple logic called WebLog that is capable of retrieving information from HTML (Hypertext Markup Language) documents in the Web, inspired by SchemaLog, a logic for multidatabase interoperability.
IEEE Intelligent Systems and their ApplicationsMaximizing text-mining performance
211 Citations1999Sabine Weiß, Chid Apte +5 more
Lecture notes in computer scienceResearch Issues in Web Data Mining
207 Citations1999Sanjay Madria, Sourav S. Bhowmick +2 more
This paper focuses on web data mining research in context of the authors' web warehousing project called WHOWEDA (Warehouse of Web Data), and categorized web datamining into threes areas; web content mining, web structure mining and web usage mining.
Discovering trends in text databases
207 Citations1997Brian Lent, Rakesh Agrawal +1 more
This work addresses the problem of discovering trends in text databases by defining a trend, a specific subsequence of the history of a phrase that satisfies the users’ query over the histories.
Querying the World Wide Web
198 Citations2002Alberto O. Mendelzon, George A. Mihaila +1 more
Lecture notes in computer scienceMining association rules in multiple relations
188 Citations1997Luc De Raedt, Luc De Raedt
The system Warmr is presented, which extends Apriori to mine association rules in multiple relations, and is applied to the natural language processing task of mining part-of-speech tagging rules in a large corpus of English.
ACM SIGMOD RecordA query language for a Web-site management system
185 Citations1997Mary Fernández, Daniela Florescu +2 more
The syntax and semantics of STRUQL, the query language at the core of STRUDEL, are described and it is believed that STRuQL is a language of independent interest, and is useful for other applications involving the management of semistructured data, as well as a view definition language for such data.
A machine learning approach to building domain-specific search engines
183 Citations1999Andrew McCallum, Kamal Nigam +2 more
The use of machine learning techniques are proposed to greatly automate the creation and maintenance of domain-specific search engines and new research in reinforcement learning, text classification and information extraction that enables efficient spidering, populates topic hierarchies, and identifies informative text segments is described.
Cut and paste
182 Citations1997Paolo Atzeni, Giansalvatore Mecca
The paper develops Editor, a language for manipulating semistructured documents, such as those typically available on the Web, that is computationally complete, in the sense that any computable document restructuring can be expressed in Editor.
Courses and lecturesUser Modeling in Adaptive Interface
170 Citations1999Pat Langley
The notion of adaptive user interfaces, interactive systems that invoke machine learning to improve their interaction with humans, is examined and three ongoing research efforts that extend this framework in new directions are described.
Lecture notes in computer scienceText mining at the term level
165 Citations1998Ronen Feldman, Moshe Fresko +6 more
This paper describes the Term Extraction module of the Document Explorer system, and provides experimental evaluation performed on a set of 52,000 documents published by Reuters in the years 1995–1996.
Theory and Practice of Object SystemsWebOQL: Restructuring documents, databases, and webs
164 Citations1999Gustavo O. Arocena, Alberto O. Mendelzon
Medical Entomology and ZoologyPrinciples of Multimedia Database Systems
142 Citations1998V. S. Subrahmanian
The author reveals how the design and architecture of a Multimedia Database and Query Languages for Retrieving Multimedia Data based on the Principle of Uniformity changed over time from simple to complex to efficient and effective.
IEEE Transactions on Knowledge and Data EngineeringDiscovering structural association of semistructured data
139 Citations2000Ke Wang, Huiqing Liu
The discovery task is affected by structural features of semistructured data in a nontrivial way and traditional data mining frameworks are inapplicable.
Lecture notes in computer scienceExploiting Structural Information for Text Classification on the WWW
123 Citations1999Johannes Fürnkranz
Experimental evidence is presented that confirms the working hypothesis that it is often easier to classify a hypertext page using information provided on pages that point to it instead of using information that is provided on the page itself.
ACM SIGMOD RecordInferring structure in semistructured data
119 Citations1997Svetlozer Nestorov, Serge Abiteboul +1 more
A notion of a type hierarchy for such data is proposed, and a method for deriving the type hierarchy is outlined, and rules for assigning types to data elements are outlined.
Little words can make a big difference for text classification
117 Citations1995Ellen Riloff
This work presents results from text classification experiments that compare relevancy signatures, which use local linguistic context, with corresponding indexing terms that do not, and suggests that stopword lists and stemming algorithms may remove or conflate many words that could be used to create more effective indexing Terms.
The Cluster-Abstraction Model: Unsupervised Learning of Topic Hierarchies from Text Data
117 Citations1999Thomas Hofmann
This paper presents a novel statistical latent class model for text mining and interactive information access, called Cluster-Abstraction Model (CAM), which is purely data driven and utilizes contact-specific word occurrence statistics.
Data mining and the Web
105 Citations1999Minos Garofalakis, Rajeev Rastogi +2 more
Popular data mining techniques like association rules, classification, clustering and outlier detection are reviewed as well as efficient algorithms for implementing the technique, that have been proposed by researchers in recent years are discussed.
Research Portal (King's College London)Navigation Pattern Discovery from Internet Data
90 Citations1999AG Büchner, Matthias Baumgarten +3 more
A new algorithm called MiDAS is introduced that extends traditional sequence discovery with a wide range of web-specific features and allows the detection of sequences across monitored attributes, such as URLs and http referrers.
MultiMediaMiner
89 Citations1998Osmar R. Zai͏̈ane, Jiawei Han +3 more
The construction of a multimedia data cube which facilitates multiple dimensional analysis of multimedia data, primarily based on visual content, and the mining of multiple kinds of knowledge, including summarization, comparison, classification, association, and clustering.
A Mutually Beneficial Integration of Data Mining and Information Extraction
86 Citations2000Un Yong Nahm, Raymond J. Mooney
A system called DISCOTEX is described, that combines IE and data mining methodologies to perform text mining as well as improve the performance of the underlying extraction system.
Applying data mining techniques for descriptive phrase extraction in digital document collections
81 Citations2002Helena Ahonen-Myka, O. P. Heinonen +2 more
This paper shows that general data mining methods are applicable to text analysis tasks such as descriptive phrase extraction and presents a general framework for text mining, based on generalized episodes and episode rules.
Studies in classification, data analysis, and knowledge organizationText Mining - Knowledge extraction from unstructured textual data
72 Citations1998Martin Rajman, Romaric Besançon
This paper presents two examples of information that can be automatically extracted from text collections: probabilistic associations of key-words and prototypical document instances and the Natural Language Processing tools necessary for such extractions.
An overview of issues in developing industrial data mining and knowledge discovery applications
72 Citations1996Gregory Piatetsky-Shapiro, Ron Brachman +3 more
This paper surveys the growing number of industrial applications of data mining and knowledge discovery, and describes some representative applications, and examines how to assess the potential of a knowledge discovery application.
Schema discovery for semistructured data
71 Citations1997Ke Wang, Huiqing Liu
This work motivates the schema discovery in this general setting and proposes a framework and algorithm for it and applies the framework to a real Web database, the Internet Movies Database, to discover typical schema of most voted movies.
Medical Entomology and ZoologyMultimedia and imaging databases
70 Citations1995Setrag Khoshfian, Andrea Baker
This book provides information essential to the incorporation of multimedia databases that will improve the quantity and quality of information manipulated by computer users in many areas including medicine, computer aided design, and information retrieval systems.
Text mining: a new frontier for lossless compression
68 Citations1999Ian H. Witten, Z. Bray +2 more
This paper aims to promote text compression as a key technology for text mining, allowing databases to be created from formatted tables such as stock-market information on Web pages.
Lecture notes in computer scienceData Mining for the Web
56 Citations1999Myra Spiliopoulou
The web is being increasingly used as a borderless marketplace for the purchase and exchange of goods, the most prominent among them being information.
Lecture notes in computer scienceAutomatic Labeling of Self-Organizing Maps: Making a Treasure-Map Reveal Its Secrets
55 Citations1999Andreas Rauber, Dieter Merkl
The LabelSOM approach for automatically labeling a trained self-organizing map with the features of the input data that are the most relevant ones for the assignment of a set of input data to a particular cluster is presented.
Mining association rules in hypertext databases
54 Citations1998José Borges, Mark Levene
The concepts of confidence and support for composite association rules, and two algorithms to mine such rules are proposed, show that, in spite of the worst-case complexity analysis which indicates exponential behaviour, in practice the algorithms' complexity is linear in the number of nodes traversed.
Artificial Intelligence ReviewMedical Data Mining on the Internet: Research on a Cancer Information System
49 Citations1999Andrea L. Houston, Hsinchun Chen +5 more
An architecture for medical knowledge information systems that will permit data mining across several medical information sources is proposed and a suite of data mining tools that are developing to assist NCI in improving public access to and use of the vast cancer information collections are discussed.
IEEE Intelligent Systems and their ApplicationsTetraFusion: information discovery on the Internet
39 Citations1999Francis Crimmins, Alan F. Smeaton +2 more
The TetraFusion system supports knowledge discovery from the World Wide Web by helping users perform data mining operations on sets of harvested URLs.
Lecture notes in computer scienceLearning for Text Categorization and Information Extraction with ILP
38 Citations2000Markus Junker, Michael Sintek +1 more
This paper introduces three basic types and three simple predicate definitions over these types which enable us to write text categorization and information extraction rules as logic programs and presents an approach to the problem of learning rules for TC and IE in terms of ILP.
WebML: Querying the World-Wide Web for Resources and Knowledge.
37 Citations1998Osmar R. Zaı̈ane, Jiawei Han
A declarative query language that would allow resource discovery on the Internet with interactive and progressively interactive inquiries and consents to the discovery of knowledge within the content of the documents and the structure of the hyperspace is proposed.
Lecture notes in computer scienceIn Search of the Lost Schema
33 Citations1999Stéphane Grumbach, Giansalvatore Mecca
Depending upon the encoding of empty sets, two polynomial on-line algorithms are proposed for solving the schema finding problem, and it is proved that with a high probability, both algorithms find the schema after examining a fixed number of tuples, thus leading in practice to a linear time behavior with respect to the database size for wrapping the data.
IEEE Intelligent Systems and their ApplicationsGleaning the Web
32 Citations1999N. Kushmerik
This work describes these IE tasks and explains how machine learning yields highly scalable IE systems, and discusses remaining challenges and argues that scaling up AI applications on the Internet is an important challenge to machine learning.
Intelligent Agents for Web-based Tasks: An Advice-Taking Approach
25 Citations1998Jude Shavlik, Tina Eliassi‐Rad
The architecture provides an appealing middle ground between nonadaptive agent programming languages and systems that solely learn user preferences from the user’s ratings of pages, and how advice is mapped into neural network implementations of the two functions.
Some Practical Observations on Integration of Web Information.
21 Citations1999William W. Cohen
The patient showed an activated fibrinolytic system and was given dextran in order to prevent further platelet aggregation and fibrin deposition, and freshly frozen plasma and fresh blood were given to replace the coagulation factors.
Large-Scale Mining of Usage Data on Web Sites
20 Citations2000Γεώργιος Παλιούρας, Christos Papatheodorou +3 more
An approach to the discovery of trends in the usage of large Web-based information systems based on the empirical analysis of the users interaction with the system and the construction of user groups with common interests (user communities).
A robust system architecture for mining semi-structured data
12 Citations1998Lisa Singh, Bin Chen +3 more
A versatile system architecture for text mining that maintains structured data components in a relational database and unstructured concepts in a concept library is proposed.
Finding Co-occurring Text Phrases by Combining Sequence and Frequent Set Discovery
10 Citations1999Helena Ahonen-Myka, Oskari Heinonen +2 more
This work considers nding multi-term text phrases that tend to co-occur in the documents of a document collection that are frequent in the document collection and that are not contained in any other longer frequent sequence.
IEEE Intelligent Systems and their ApplicationsIntegrating and using large databases of text, images, video, and audio
6 Citations1999Alex Hauptmann
There is no clear categorization or organization of the various research efforts concerning mixed-media databases concerning TIVA sources, but intelligent, content-understanding systems can greatly improve the usefulness of the huge quantities of existing material from these sources.
Research Showcase @ Carnegie Mellon University (Carnegie Mellon University)CONALD Report on the Workshop on Learning from Text and the Web
3 Citations2018Jaime Carbonell, Mark Craven +3 more
…
