Learning with positive and unlabeled examples using weighted logistic regression
Published 21 August 2003Open access
Wee Sun Lee, Bing Liu
Citations334
SJR quartileQ4
SJR score0.11
SNIP0.06
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A performance measure that can be estimated from positive and unlabeled examples for evaluating retrieval performance, which is proportional to the product of precision and recall, can be used with a validation set to select regularization parameters for logistic regression.
Abstract
The problem of learning with positive and unlabeled examples arises frequently in retrieval applications.
Keywords
Computer Science
ACM Transactions on Intelligent Systems and TechnologyLIBSVM
41,340 Citations2011Chih-Chung Chang, Chih‐Jen Lin
Issues such as solving SVM optimization problems theoretical convergence multiclass classification probability estimates and parameter selection are discussed in detail.
Combining labeled and unlabeled data with co-training
5,604 Citations1998Avrim Blum, Tom M. Mitchell
Technical reportsMaking Large-Scale SVM Learning Practical
4,317 Citations2006Thorsten Joachims
This chapter presents algorithmic and computational results developed for SVM light V 2.0, which make large-scale SVM training more practical and give guidelines for the application of SVMs to large domains.
Transductive Inference for Text Classification using Support Vector Machines
2,717 Citations1999Thorsten Joachims
An analysis of why TSVMs are well suited for text classi(cid:12)cation is presented, and an algorithm for training TSVMs e(cid:14)-ciently, handling 10,000 examples and more is proposed.
Elsevier eBooksNewsWeeder: Learning to Filter Netnews
2,022 Citations1995Ken Lang
The results show that a learning algorithm based on the Minimum Description Length (MDL) principle was able to raise the percentage of interesting articles to be shown to users from 14% to 52% on average.
One-class svms for document classification
1,225 Citations2002Larry M. Manevitz, Malik Yousef
The SVM approach as represented by Schoelkopf was superior to all the methods except the neural network one, where it was, although occasionally worse, essentially comparable.
Journal of the ACMEfficient noise-tolerant learning from statistical queries
712 Citations1998Michael Kearns
This paper formalizes a new but related model of learning from statistical queries, and demonstrates the generality of the statistical query model, showing that practically every class learnable in Valiant's model and its variants can also be learned in the new model (and thus can be learning in the presence of noise).
Partially Supervised Classification of Text Documents
516 Citations2002Bing Liu, Wee Sun Lee +2 more
This paper shows that the problem of identifying documents from a set of documents of a particular topic or class P and a large set M of mixed documents, and that under appropriate conditions, solutions to the constrained optimization problem will give good solution to the partially supervised classification problem.
Learning to classify text from labeled and unlabeled documents
330 Citations1998Kamal Nigam, Andrew McCallum +2 more
It is shown that the accuracy of text classifiers trained with a small number of labeled documents can be improved by augmenting this small training set with a large pool of unlabeled documents, and an algorithm is introduced based on the combination of Expectation-Maximization with a naive Bayes classifier.
PEBL
277 Citations2002Hwanjo Yu, Jiawei Han +1 more
The Positive Example Based Learning (PEBL) framework for Web page classification is introduced which eliminates the need for manually collecting negative training examples in pre-processing and an algorithm called Mapping-Convergence (M-C) is presented that achieves classification accuracy as high as that of traditional SVM (with positive and negative data).
Lecture notes in computer scienceLearning from positive data
200 Citations1997Stephen Muggleton
New results are presented which show that within a Bayesian framework not only grammars, but also logic programs are learnable with arbitrarily low expected error from positive examples only and the upper bound for expected error of a learner which maximises the Bayes' posterior probability is within a small additive term of one which does the same from a mixture of positive and negative examples.
Lecture notes in computer sciencePAC Learning from Positive Statistical Queries
177 Citations1998François Denis
This work defines a PAC learning model from positive and unlabeled examples, and defines a PAC learning model from positive and unlabeled statistical queries and shows that k-DNF and k-decision lists are learnable in both models.
Journal of Computer and System SciencesRobust Trainability of Single Neurons
153 Citations1995K.-U. Hoffgen, Hans Ulrich Simon +1 more
It is shown that the problem of learning a probably almost optimal weight vector for a neuron is so difficult that the minimum error cannot even be approximated to within a constant factor in polynomial time (unless RP = NP); the same hardness result is obtained for several variants of this problem.
A polynomial-time algorithm for learning noisy linear threshold functions
121 Citations2002Avrim Blum, Alan Frieze +2 more
Lecture notes in computer scienceAthena: Mining-Based Interactive Management of Text Databases
87 Citations2000Rakesh Agrawal, Roberto J. Bayardo +1 more
Athena is described: a system for creating, exploiting, and maintaining a hierarchy of textual documents through interactive mining-based operations through linear-time classification and clustering engines which are applied interactively to speed the development of accurate models.
International Conference Information ProcessingText Classification from Positive and Unlabeled Examples
82 Citations2002François Denis, Rémi Gilleron +1 more
Experimental results show that performan e of the naive Bayes algorithm for learning from positive and unlabeled examples is omparable with naive Bayesian algorithm forLearning from labeled data.
Learning linear threshold functions in the presence of classification noise
77 Citations1994Tom Bylander
It is shown that the linear threshold functions are polynomially learnable in the presence of classification noise, i.e., polynomial in n, where n is the number of Boolean attributes, ε and δ are the usual accuracy and confidence parameters, and &sgr; indicates the minimum distance of any example from the target hyperplane.
