Mining the Web: Discovering Knowledge from Hypertext Data
Published 1 January 2002
Soumen Chakrabarti
Citations695
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
Preface. Introduction. I Infrastructure: Crawling the Web. Web search. II Learning: Similarity and clustering. Supervised learning for text. Semi-supervised learning. III Applications: Social network analysis. Resource discovery. The future of Web mining.
Keywords
Computer Science
Elements of Information Theory
37,533 Citations2001Thomas M. Cover, Joy A. Thomas
ScienceEmergence of Scaling in Random Networks
36,366 Citations1999Albert-Ĺaszló Barabási, Réka Albert
A model based on these two ingredients reproduces the observed stationary scale-free distributions, which indicates that the development of large networks is governed by robust self-organizing phenomena that go beyond the particulars of the individual systems.
Choice Reviews OnlineData mining: concepts and techniques
28,877 Citations2012Jiawei Han, Micheline Kamber +1 more
Data mining is the search for new, valuable, and nontrivial information in large volumes of data, a cooperative effort of humans and computers that is possible to put data-mining activities into one of two categories: Predictive data mining, which produces the model of the system described by the given data set, or Descriptive data mining which produces new, nontrivials information based on the available data set.
TechnometricsStatistical Learning Theory
26,913 Citations1999Yuhai Wu, Vladimir Vapnik
Presenting a method for determining the necessary and sufficient conditions for consistency of learning process, the author covers function estimates from small data pools, applying these estimations to real-life problems, and much more.
Social network analysis methods and applications
18,142 Citations2007Stanley Wasserman, Katherine Faust
Computer Networks and ISDN SystemsThe anatomy of a large-scale hypertextual Web search engine
15,828 Citations1998Sergey Brin, Lawrence M. Page
This paper provides an in-depth description of Google, a prototype of a large-scale search engine which makes heavy use of the structure present in hypertext and looks at the problem of how to effectively deal with uncontrolled hypertext collections where anyone can publish anything they want.
Journal of the American Society for Information ScienceIndexing by latent semantic analysis
12,677 Citations1990Scott Deerwester, Susan Dumais +3 more
The PageRank Citation Ranking : Bringing Order to the Web
12,645 Citations1999Lawrence M. Page, Sergey Brin +2 more
This paper describes PageRank, a mathod for rating Web pages objectively and mechanically, effectively measuring the human interest and attention devoted to them, and shows how to efficiently compute PageRank for large numbers of pages.
Pattern classification and scene analysis
12,643 Citations1973Richard O. Duda, Peter E. Hart
The Selfish Gene
10,510 Citations1976Richard Dawkins
Journal of the ACMAuthoritative sources in a hyperlinked environment
9,060 Citations1999Jon Kleinberg
This work proposes and test an algorithmic formulation of the notion of authority, based on the relationship between a set of relevant authoritative pages and the set of “hub pages” that join them together in the link structure, and has connections to the eigenvectors of certain matrices associated with the link graph.
Lecture notes in computer scienceText categorization with Support Vector Machines: Learning with many relevant features
7,925 Citations1998Thorsten Joachims
SVMs achieve substantial improvements over the currently best performing methods and behave robustly over a variety of di-erent learning tasks, eliminating the need for manual parameter tuning.
Combining labeled and unlabeled data with co-training
5,604 Citations1998Avrim Blum, Tom M. Mitchell
Proceedings of the IEEEThe viterbi algorithm
5,595 Citations1973G. David Forney
This paper gives a tutorial exposition of the Viterbi algorithm and of how it is implemented and analyzed, and increasing use of the algorithm in a widening variety of areas is foreseen.
A Comparative Study on Feature Selection in Text Categorization
4,766 Citations1997Yiming Yang, Jan Pedersen
DF thresholding, the simplest method with the lowest cost in computation, can be reliably used instead of IG or CHI when the computation of these measures are too expensive, and strong correlations between the DF, IG and CHI values of a term are found.
Technical reportsMaking Large-Scale SVM Learning Practical
4,317 Citations2006Thorsten Joachims
This chapter presents algorithmic and computational results developed for SVM light V 2.0, which make large-scale SVM training more practical and give guidelines for the application of SVMs to large domains.
A theory of the learnable
4,242 Citations1984Leslie G. Valiant
This paper regards learning as the phenomenon of knowledge acquisition in the absence of explicit programming, and gives a precise methodology for studying this phenomenon from a computational viewpoint.
Cambridge University Press eBooksRandomized Algorithms
4,069 Citations1995Rajeev Motwani, Prabhakar Raghavan
BIRCH
3,906 Citations1996Tian Zhang, Raghu Ramakrishnan +1 more
A data clustering method named BIRCH (Balanced Iterative Reducing and Clustering using Hierarchies) is presented, and it is demonstrated that it is especially suitable for very large databases.
Elsevier eBooksFast Effective Rule Induction
3,767 Citations1995William W. Cohen
This paper evaluates the recently-proposed rule learning algorithm IREP on a large and diverse collection of benchmark problems, and proposes a number of modifications resulting in an algorithm RIPPERk that is very competitive with C4.5 and C 4.5rules with respect to error rates, but much more efficient on large samples.
A comparison of event models for naive bayes text classification
3,224 Citations1998Andrew McCallum, Kamal Nigam
It is found that the multi-variate Bernoulli performs well with small vocabulary sizes, but that the multinomial performs usually performs even better at larger vocabulary sizes--providing on average a 27% reduction in error over the multi -variateBernoulli model at any vocabulary size.
Similarity Search in High Dimensions via Hashing
3,108 Citations1999Aristides Gionis, Piotr Indyk +1 more
Experimental results indicate that the novel scheme for approximate similarity search based on hashing scales well even for a relatively large number of dimensions, and provides experimental evidence that the method gives improvement in running time over other methods for searching in highdimensional spaces based on hierarchical tree decomposition.
Machine LearningOn the Optimality of the Simple Bayesian Classifier under Zero-One Loss
3,072 Citations1997Pedro Domingos, Michael J. Pazzani
The Bayesian classifier is shown to be optimal for learning conjunctions and disjunctions, even though they violate the independence assumption, and will often outperform more powerful classifiers for common training set sizes and numbers of attributes, even if its bias is a priori much less appropriate to the domain.
As We May Think
3,008 Citations1945Vannevar Bush
Machine LearningText Classification from Labeled and Unlabeled Documents using EM
2,749 Citations2000Kamal Nigam, Andrew Kachites McCallum +2 more
This paper shows that the accuracy of learned text classifiers can be improved by augmenting a small number of labeled training documents with a large pool of unlabeled documents, and presents two extensions to the algorithm that improve classification accuracy under these conditions.
Accurate methods for the statistics of surprise and coincidence
2,688 Citations1993Ted Dunning
Sequential Minimal Optimization : A Fast Algorithm for Training Support Vector Machines
2,685 Citations1998John Platt
A re-examination of text categorization methods
2,660 Citations1999Yiming Yang, Xin Liu
The results show that SVM, kNN and LLSF signi cantly outperform NNet and NB when the number of positive training instances per category are small, and that all the methods perform comparably when the categories are over 300 instances.
ACM SIGIR ForumA Language Modeling Approach to Information Retrieval
2,532 Citations2017Jay Ponte, W. Bruce Croft
It will be shown that probabilistic methods can be used to predict topic changes in the context of the task of new event detection and provide further proof of concept for the use of language models for retrieval tasks.
Information Retrieval: Data Structures and Algorithms
2,428 Citations1992William B. Frakes, Ricardo Baeza‐Yates
For programmers and students interested in parsing text, automated indexing, its the first collection in book form of the basic data structures and algorithms that are critical to the storage and retrieval of documents.
Principles of Data Mining
2,398 Citations2001David J. Hand, Heikki Mannila +1 more
Automatic subspace clustering of high dimensional data for data mining applications
2,386 Citations1998Rakesh Agrawal, Johannes Gehrke +2 more
CLIQUE is presented, a clustering algorithm that satisfies each of these requirements of data mining applications including the ability to find clusters embedded in subspaces of high dimensional data, scalability, end-user comprehensibility of the results, non-presumption of any canonical data distribution, and insensitivity to the order of input records.
Springer series in statisticsEstimation with Quadratic Loss
2,306 Citations1992William James, Charles Stein
arXiv (Cornell University)Probabilistic Latent Semantic Analysis
2,092 Citations2013Thomas Hofmann
This work proposes a widely applicable generalization of maximum likelihood model fitting by tempered EM, based on a mixture decomposition derived from a latent class model which results in a more principled approach which has a solid foundation in statistics.
Lecture notes in computer scienceNaive (Bayes) at forty: The independence assumption in information retrieval
2,092 Citations1998David Lewis
The naive Bayes classifier, currently experiencing a renaissance in machine learning, has long been a core technique in information retrieval, and some of the variations used for text retrieval and classification are reviewed.
Communications of the ACMCYC
1,947 Citations1995Douglas B. Lenat
The fundamental assumptions of doing such a large-scale project are examined, the technical lessons learned by the developers are reviewed, and the range of applications that are or soon will be enabled by the technology is surveyed.
INADMISSIBILITY OF THE USUAL ESTIMATOR FOR THE MEAN OF A MULTIVARIATE NORMAL DISTRIBUTION
1,849 Citations1956Charles Stein
Studies in computational intelligenceA Tutorial on Learning with Bayesian Networks
1,699 Citations2008David Heckerman
Methods for constructing Bayesian networks from prior knowledge are discussed and methods for using data to improve these models are summarized, including techniques for learning with incomplete data.
Introduction to Government and Binding Theory
1,594 Citations1991Liliane Haegeman
The Chomskyan Perspective on Language Study presents a perspective on language study from the perspective of a Chomskyan linguist, focusing on Germanic Languages.
Computer NetworksFocused crawling: a new approach to topic-specific Web resource discovery
1,492 Citations1999Soumen Chakrabarti, Martin van den Berg +1 more
A new hypertext resource discovery system called a Focused Crawler that is robust against large perturbations in the starting set of URLs, and capable of exploring out and discovering valuable resources that are dozens of links away from the start set, while carefully pruning the millions of pages that may lie within this same radius.
SIAM ReviewUsing Linear Algebra for Intelligent Information Retrieval
1,484 Citations1995Michael W. Berry, Susan Dumais +1 more
A lexical match between words in users’ requests and those in or assigned to documents in a database helps retrieve textual materials from scientific databases.
Inductive learning algorithms and representations for text categorization
1,465 Citations1998Susan Dumais, John Platt +2 more
A comparison of the effectiveness of five different automatic learning algorithms for text categorization in terms of learning speed, realtime classification speed, and classification accuracy is compared.
Toward optimal feature selection
1,444 Citations1996Daphne Koller, Mehran Sahami
An efficient algorithm for feature selection which computes an approximation to the optimal feature selection criterion is given, showing that the algorithm effectively handles datasets with a very large number of features.
A simple rule-based part of speech tagger
1,418 Citations1992Eric Brill
This work presents a simple rule-based part of speech tagger which automatically acquires its rules and tags with accuracy comparable to stochastic taggers, demonstrating that the stochastics method is not the only viable method for part ofspeech tagging.
NatureAccessibility of information on the web
1,356 Citations1999Steve Lawrence, C. Lee Giles
As the web becomes a major communications medium, the data on it must be made more accessible, and search engines need to make the data more accessible.
Computer Networks and ISDN SystemsSyntactic clustering of the Web
1,347 Citations1997Andrei Broder, S. Glassman +2 more
An efficient way to determine the syntactic similarity of files is developed and applied to every document on the World Wide Web, and a clustering of all the documents that are syntactically similar is built.
The MIT Press eBooksNatural Language Understanding
1,279 Citations2021
The text features a new chapter on statistically-based methods using large corpora and an appendix on speech recognition and spoken language understanding and information on semantics that was covered in the first edition has been largely expanded in this edition.
A multilevel algorithm for partitioning graphs
1,073 Citations1995Bruce Hendrickson, Robert W. Leland
A multilevel algorithm for graph partitioning in which the graph is approximated by a sequence of increasingly smaller graphs, and the smallest graph is then partitioned using a spectral method, and this partition is propagated back through the hierarchy of graphs.
IEEE Transactions on Pattern Analysis and Machine IntelligenceInducing features of random fields
1,044 Citations1997S. Della Pietra, V. Della Pietra +1 more
The random field models and techniques introduced in this paper differ from those common to much of the computer vision literature in that the underlying random fields are non-Markovian and have a large number of parameters that must be estimated.
Drug SafetyPrinciples of Data Mining
1,034 Citations2007David J. Hand
The book consists of three sections and provides a tutorial overview of the principles underlying data mining algorithms and their application, and shows how all of the preceding analysis fits together when applied to real-world data mining problems.
Data Mining and Knowledge DiscoveryOn Bias, Variance, 0/1—Loss, and the Curse-of-Dimensionality
1,029 Citations1997Jerome H. Friedman
This work candramatically mitigate the effect of the bias associated with some simpleestimators like “naive” Bayes, and the bias induced by the curse-of-dimensionality on nearest-neighbor procedures.
Analyzing the effectiveness and applicability of co-training
1,028 Citations2000Kamal Nigam, Rayid Ghani
It is demonstrated that when learning from labeled and unlabeled data, algorithms explicitly leveraging a natural independent split of the features outperform algorithms that do not and may out-perform algorithms not using a split.
The MIT Press eBooksProbabilities for SV Machines
1,015 Citations2000John C. Platt
This chapter contains sections titled: Introduction, Fitting a Sigmoid After the SVM, Empirical Tests, Conclusions, Appendix: Pseudo-code for the Sigmoids Training.
Computer NetworksTrawling the Web for emerging cyber-communities
1,010 Citations1999Ravi Kumar, Prabhakar Raghavan +2 more
The subject of this paper is the systematic enumeration of over 100,000 emerging communities from a Web crawl, motivating a graph-theoretic approach to locating such communities, and describing the algorithms and algorithmic engineering necessary to find structures that subscribe to this notion.
ComputerSelf-organization and identification of Web communities
1,007 Citations2002Gary William Flake, Sandra Lawrence +2 more
This work shows that the Web self-organizes and its link structure allows efficient identification of communities and is significant because no central authority or process governs the formation and structure of hyperlinks.
ScienceSearching the World Wide Web
974 Citations1998Steve Lawrence, C. Lee Giles
The coverage and recency of the major World Wide Web search engines was analyzed, yielding some surprising results, including a lower bound on the size of the indexable Web of 320 million pages.
Bayesian classification (AutoClass): theory and results
972 Citations1996Peter Cheeseman, John Stutz
It is emphasized that no current unsupervised classi(cid:12)cation system can produce maximally useful results when operated alone, and that it is the interaction between domain experts and the machine searching over the model space, that generates new knowledge.
Information Processing & ManagementA probabilistic model of information retrieval: development and comparative experiments
947 Citations2000Karen Spärck Jones, Steve Walker +1 more
Machine LearningLearning Information Extraction Rules for Semi-Structured and Free Text
928 Citations1999Stephen Soderland
WHISK is designed to handle text styles ranging from highly structured to free text, including text that is neither rigidly formatted nor composed of grammatical sentences, and can also handle extraction from free text such as news stories.
Medical Entomology and ZoologyInference and disputed authorship : The Federalist
916 Citations1964Frederick Mosteller, David L. Wallace
IEEE Transactions on Neural NetworksSelf organization of a massive document collection
911 Citations2000Teuvo Kohonen, Samuel Kaski +5 more
A system that is able to organize vast document collections according to textual similarities based on the self-organizing map (SOM) algorithm, based on 500-dimensional vectors of stochastic figures obtained as random projections of weighted word histograms.
Latent semantic indexing
870 Citations1998Christos H. Papadimitriou, Hisao Tamaki +2 more
It is proved that, under certain conditions, LSI does succeed in capturing the underlying semantics of the corpus and achieves improved retrieval performance, and the technique of random projection is proposed as a way of speeding up LSI.
ACM Transactions on Information SystemsAutomated learning of decision rules for text categorization
867 Citations1994Chidanand Apté, Fred J. Damerau +1 more
It is shown that machine-generated decision rules appear comparable to human performance, while using the identical rule-based representation, and compared with other machine-learning techniques.
Computer Networks and ISDN SystemsEfficient crawling through URL ordering
840 Citations1998Junghoo Cho, Héctor García-Molina +1 more
This paper studies in what order a crawler should visit the URLs it has seen, in order to obtain more "important" pages first, and shows that a Crawler with a good ordering scheme can obtain important pages significantly faster than one without.
Hierarchically Classifying Documents Using Very Few Words
840 Citations1997Daphne Koller, Mehran Sahami
This work proposes an approach that utilizes the hierarchical topic structure to decompose the classification task into a set of simpler problems, one at each node in the classification tree, which can be solved accurately by focusing only on a very small set of features, those relevant to the task at hand.
Medical Entomology and ZoologyAn Introduction to Machine Translation
796 Citations1992W. John Hutchins, Harold Somers
General introduction and brief history linguistic background computational aspects basic strategies analysis problems of transfer and interlingua generation the practical use of MT systems evaluation Systran SUSY Meteo (TAUM) Ariane (GETA) Eurotra METAL Rosetta DLT some other systems and directions of research.
Enhanced hypertext categorization using hyperlinks
775 Citations1998Soumen Chakrabarti, Byron Dom +1 more
This work has developed a text classifier that misclassified only 13% of the documents in the well-known Reuters benchmark; this was comparable to the best results ever obtained and its technique also adapts gracefully to the fraction of neighboring documents having known topics.
Inferring Web communities from link topology
772 Citations1998David Gibson, Jon Kleinberg +1 more
This investigation shows that although the process by which users of the Web create pages and links is very difficult to understand at a “local” level, it results in a much greater degree of orderly high-level structure than has typically been assumed.
Using Maximum Entropy for Text Classification
756 Citations1999Kamal Nigam, John Lafferty +1 more
This paper uses maximum entropy techniques for text classification by estimating the conditional distribution of the class variable given the document by comparing accuracy to naive Bayes and showing that maximum entropy is sometimes significantly better, but also sometimes worse.
Scaling clustering algorithms to large databases
709 Citations1998Patricia Bradley, Usama M. Fayyad +1 more
A scalable clustering framework applicable to a wide class of iterative clustering that requires at most one scan of the database and is instantiated and numerically justified with the popular K-Means clustering algorithm.
Information Processing & ManagementRecent trends in hierarchic document clustering: A critical review
707 Citations1988Peter Willett
Algorithms that can be used to allow the implementation of hierarchic agglomerative clustering methods for document retrieval, and experimental evidence suggests that nearest neighbor clusters provide a reasonably efficient and effective means of including interdocument similarity information in document retrieval systems.
Computer Networks and ISDN SystemsAutomatic resource compilation by analyzing hyperlink structure and associated text
700 Citations1998Soumen Chakrabarti, Byron Dom +4 more
An evaluation of ARC suggests that the resources found by ARC frequently fare almost as well as, and sometimes better than, lists of resources that are manually compiled or classified into a topic.
ACM SIGIR ForumImproved Algorithms for Topic Distillation in a Hyperlinked Environment
677 Citations2017Krishna Bharat, Monika Henzinger
This paper addresses the problem of topic distillation on the World Wide Web, namely, given a typical user query to find quality documents related to the query topic, by augmenting a previous connectivity analysis based algorithm with content analysis.
Language<b>Survey of the state of the art in human language technology</b> . Ed. by Giovanni Varile, Antonio Zampolli, Ronald Cole, Joseph Mariani, Hans Uszkoreit, Annie Zaenen, and Victor Zue. (Studies in natural language processing.) Cambridge: Cambridge University Press. Pp. xx, 513. Cloth $49.95.
604 Citations2000Kay Cohen
World Wide WebMercator: A scalable, extensible Web crawler
578 Citations1999Allan Heydon, Marc Najork
This paper describes Mercator, a scalable, extensible Web crawler written entirely in Java, and comments on Mercator's performance, which is found to be comparable to that of other crawlers for which performance numbers have been published.
ACM Transactions on Information SystemsEvaluation of an inference network-based retrieval model
577 Citations1991Howard R. Turtle, W. Bruce Croft
Network representations show promise as mechanisms for inferring probable relationships between documents and queries and have been used in information retrieval since at least the early 1960s.
Focused Crawling Using Context Graphs
521 Citations2000Michelangelo Diligenti, Frans Coetzee +3 more
A focused crawling algorithm is presented that builds a model for the context within which topically relevant pages occur on the web that can capture typical link hierarchies within which valuable pages occur, as well as model content on documents that frequently cooccur with relevant pages.
ComputerMining the Web's link structure
502 Citations1999Soumen Chakrabarti, Byron Dom +6 more
Clever is a search engine that analyzes hyperlinks to uncover two types of pages: authorities, which provide the best source of information on a given topic; and hubs, which provides collections of links to authorities.
Relational Learning of Pattern-Match Rules for Information Extraction.
494 Citations1997Mary Elaine Califf, Raymond J. Mooney
Computer NetworksFinding related pages in the World Wide Web
491 Citations1999Jay B. Dean, Monika Henzinger
This paper discusses a different approach to Web searching where the input to the search process is not a set of query terms, but instead is the URL of a page, and the output is aSet of related Web pages.
Improving Text Classification by Shrinkage in a Hierarchy of Classes
478 Citations2022Andrew McCallum, Roni Rosenfeld +2 more
This paper shows that the accuracy of a naive Bayes text classi(cid:12)er can be significantly improved by taking advantage of a hierarchy of classes, and adopts an established statistical technique called shrinkage that smoothes parameter estimates of a data-sparse child with its parent in order to obtain more robust parameter estimates.
Proceedings of the National Academy of SciencesWinners don't take all: Characterizing the competition for links on the web
461 Citations2002David M. Pennock, Gary William Flake +3 more
A simple generative model quantifies the degree to which the rich nodes grow richer, and how new (and poorly connected) nodes can compete, and accurately accounts for the true connectivity distributions of category-specific web pages, the web as a whole, and other social networks.
arXiv (Cornell University)Probabilistic Models for Unified Collaborative and Content-Based Recommendation in Sparse-Data Environments
440 Citations2013Alexandrin Popescul, Lyle Ungar +2 more
It is shown that secondary content information can often be used to overcome sparsity and appropriate mixture models incorporating secondary data produce significantly better quality recommenders than k-nearest neighbors (k-NN).
Integrating multiple knowledge sources to disambiguate word sense
410 Citations1996Hwee Tou Ng, Hian Beng Lee
This approach integrates a diverse set of knowledge sources to disambiguate word sense, including part of speech of neighboring words, morphological form, the unordered set of surrounding words, local collocations, and verb-object syntactic relation.
Journal of the ACMApproximation algorithms for classification problems with pairwise relationships
405 Citations2002Jon Kleinberg, Éva Tardos
Computer Networks and ISDN SystemsA technique for measuring the relative size and overlap of public Web search engines
393 Citations1998Krishna Bharat, Andrei Broder
A standardized, statistical way of measuring search engine coverage and overlap through random queries is described that can be implemented by third-party evaluators using only public query interfaces and suggests the size of the static, public Web as of November was over 200 million pages.
Computational LinguisticsGrammatical category disambiguation by statistical optimization
387 Citations1988Steven J. DeRose
An algorithm for disambiguation that is similar to CLAWS but that operates in linear rather than in exponential time and space, and which minimizes the unsystematic augments is presented.
Research Showcase @ Carnegie Mellon University (Carnegie Mellon University)Learning Hidden Markov Model Structure for Information Extraction
384 Citations2018Kristie Seymore, Roni Rosenfeld
It is demonstrated that a manually-constructed model that contains multiple states per extraction field outperforms a model with one state per field, and the use of distantly-labeled data to set model parameters provides a significant improvement in extraction accuracy.
Silk from a sow's ear
382 Citations1996Peter Pirolli, James E. Pitkow +1 more
This paper presents the exploration into techniques that utilize both the topology and textual similarity between items as well as usage data collected by servers and page meta-information lke title and size.
ACM Transactions on Information SystemsSALSA
359 Citations2001Ronny Lempel, Shlomo Moran
It is proved that SALSA is quivalent to a weighted in degree analysis of the link-sturcutre of WWW subgraphs, making it computationally more efficient than the Mutual reinforcement approach, and comparisions reveal a topological Phenomenon called the TKC effect which prevents the Mutual Reinforcement approach from identifying meaningful authorities.
Two algorithms for nearest-neighbor search in high dimensions
347 Citations1997Jon Kleinberg
A new approach to the nearest-neighbor problem is developed, based on a method for combining randomly chosen one-dimensional projections of the underlying point set, which results in an algorithm for finding e-approximate nearest neighbors with a query time of O((d log d)(d + log n)).
Scaling question answering to the Web
327 Citations2001Cody C. T. Kwok, Oren Etzioni +1 more
Mulder is introduced, which is believed to be the first general-purpose, fully-automated question-answering system available on the web, and its architecture is described, which relies on multiple search-engine queries, natural-language parsing, and a novel voting procedure to yield reliable answers coupled with high recall.
…
