Classifying the Hungarian web
Published 1 January 2003Open access
András Kornai, Marc Krellenstein, M Mulligan, David Twomey, Fruzsina Veress, Alec Wysoker
Citations5
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
Based on a simple statistical language, model, and the large-scale supporting evidence from vizsla, it is argued that in topic classification only positive evidence matters.
Abstract
In this paper we present some lessons learned from building vizsla, the keyword search and topic classification system used on the largest Hungarian portal, [origo.hu]. Based on a simple statistical language, model, and the large-scale supporting evidence from vizsla, we argue that in topic classification only positive evidence matters.
Keywords
Computer ScienceDecision Sciences
ACM SIGIR ForumA Language Modeling Approach to Information Retrieval
2,532 Citations2017Jay Ponte, W. Bruce Croft
It will be shown that probabilistic methods can be used to predict topic changes in the context of the task of new event detection and provide further proof of concept for the use of language models for retrieval tasks.
ACM SIGIR ForumA Study of Smoothing Methods for Language Models Applied to Ad Hoc Information Retrieval
1,571 Citations2017ChengXiang Zhai, John Lafferty
This paper examines the sensitivity of retrieval performance to the smoothing parameters and compares several popular smoothing methods on different test collection.
A hidden Markov model information retrieval system
469 Citations1999David R. Miller, Tim Leek +1 more
A novel method for performing blind feedback in the HMM framework, a more complex HMM that models bigram production, and several other algorithmic re nements form a state-of-the-art retrieval system that ranked among the best on the TREC-7 ad hoc retrieval task.
Proceedings of the IRELinear Decision Functions, with Application to Pattern Recognition
147 Citations1962W. H. Highleyman
This paper is concerned with the study of a particular class of categorizers, the linear decision function, which can be empirically designed without making any assumptions whatsoever about either the distribution of the receptor measurements or the a priori probabilities of occurrence of the pattern classes, providing an appropriate pattern source is available.
On relevance weights with little relevance information
144 Citations1997Stephen Robertson, Steve Walker
This research highlights the need to understand more fully the role of emotion in the development of Syetema, as well as the role that language and social media have in this process.
Twenty-One at TREC-7: ad-hoc and cross-language track
143 Citations1998Djoerd Hiemstra, Wessel Kraaij
This paper describes the official runs of the Twenty-One group for TREC-7 and develops a new weighting algorithm, which outperforms the popular Cornell version of BM25 on the ad-hoc collection and developed a fuzzy matching algorithm to recover from missing translations and spelling variants of proper names.
ACM Computing SurveysRough'n'Ready
30 Citations1999Francis Kubala, Sean Colbath +2 more
A software system consisting of a meeting recorder and browser was designed and developed to provide a higher level view of collaborative meetings, co-locational or distributed and a way to browse through and listen to those parts which are most relevant to the user.
Linear Discriminant Text Classification in High Dimension
6 Citations2002András Kornai, J. Richards
This paper trained and tested LD-based systems for a variety of classification schemes widely used in the clinical drug trial process and obtained significant reduction in the rate of misclassification compared both to generic Bayesian machine-learning techniques and to the current generation of domain-specific autocoders based on string matching.
