Document Classification Using a Finite Mixture Model
arXiv (Cornell University)Published 6 May 1997Open access
Hang Li, Kenji Yamanishi
Citations4
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
Americanae nace como un proyecto conjunto que surge dentro de la Red Europea de Información y Documentación sobre América Latina (REDIAL), y que ha afrontado la Biblioteca de la Agencia Española de Cooperación Internacional para el Desarrollo (AECID). Esta nueva biblioteca virtual hace más accesibles los libros digitales de tema americanista a los investigadores y usuarios interesados de cualquier parte del mundo.
Keywords
Computer Science
Journal of the Royal Statistical Society Series B (Statistical Methodology)Maximum Likelihood from Incomplete Data Via the <i>EM</i> Algorithm
49,657 Citations1977A. P. Dempster, N. M. Laird +1 more
Journal of the American Society for Information ScienceIndexing by latent semantic analysis
12,677 Citations1990Scott Deerwester, Susan Dumais +3 more
Journal of the American Statistical AssociationThe Calculation of Posterior Distributions by Data Augmentation
3,782 Citations1987Martin A. Tanner, Wing Hung Wong
If data augmentation can be used in the calculation of the maximum likelihood estimate, then in the same cases one ought to be able to use it in the computation of the posterior distribution of parameters of interest.
Journal of the American Society for Information ScienceRelevance weighting of search terms
2,068 Citations1976Stephen Robertson, Karen Spärck Jones
This paper examines statistical techniques for exploiting relevance information to weight search terms using information about the distribution of index terms in documents in general and shows that specific weighted search methods are implied by a general probabilistic theory of retrieval.
Encyclopedia of Statistics in Behavioral ScienceFinite Mixture Distributions
1,287 Citations2005Barry J. Everitt
Distributional clustering of English words
994 Citations1993Fernando Pereira, Naftali Tishby +1 more
ACM Transactions on Information SystemsAutomated learning of decision rules for text categorization
867 Citations1994Chidanand Apté, Fred J. Damerau +1 more
It is shown that machine-generated decision rules appear comparable to human performance, while using the identical rule-based representation, and compared with other machine-learning techniques.
An evaluation of phrasal and clustered representations on a text categorization task
547 Citations1992David Lewis
It is shown that optimal effectiveness occurs when using only a small proportion of the indexing terms available, and that effectiveness peaks at a higher feature set size and lower effectiveness level for a syntactic phrase indexing than for word-based indexing.
A comparison of classifiers and document representations for the routing problem
451 Citations1995Hinrich Schütze, David A. Hull +1 more
This paper considers three classification techniques which have decision rules that are derived via explicit error minimization linear discriminant analysis, logistic regression, and neuraf networks, and finds that features based on latent semantic indexing are more effective for techniques such aslinear discriminant anaf-ysis and logistic regressors, which have no way to protect against overfitting.
ACM Transactions on Information SystemsAn example-based mapping method for text categorization and retrieval
405 Citations1994Yiming Yang, Christopher G. Chute
It is evident that the LLSF approach uses the relevance information effectively within human decisions of categorization and retrieval, and achieves a semantic mapping of free texts to their representations in an indexing language.
Context-sensitive learning methods for text categorization
253 Citations1996William W. Cohen, Yoram Singer
Information Processing & ManagementModels for retrieval with probabilistic indexing
177 Citations1989Norbert Fuhr
Three retrieval models for probabilistic indexing are described along with evaluation results for each, including the binary independence indexing (BII) model, which is a generalized version of the Maron and Kuhns indexing model.
A comparison of new and old algorithms for a mixture estimation problem
78 Citations1995David P. Helmbold, Yoram Singer +2 more
Poor estimates of context are worse than none
62 Citations1990William A. Gale, Kenneth Church
It is found that it is possible to make useful estimates of contextual probabilities that improve performance in a spelling correction application, and it is shown that the Good-Turing method makes the use of contextual information practical for a spelling corrector.
Information Processing & ManagementA probability distribution model for information retrieval
44 Citations1989S. K. M. Wong, Yiyu Yao
A probability distribution model for information retrieval is proposed that not only enhances retrieval effectiveness as demonstrated by experiments, but also provides valuable insight into many fundamental concepts introduced over the years in a variety of retrieval models.
A randomized approximation of the MDL for stochastic models with hidden variables
10 Citations1996Kenji Yamanishi
A randomized algorithm, RAMDL, which efficiently approximates the SC for stochastic models with hidden variables, and quantifies the tradeoff relation between the statistical approximation accuracy and computational complexity, in general forms.
