Sensitive webpage classification for content advertising
Published 12 August 2007
Xin Jin, Ying Li, Teresa Mah, Jie Tong
Citations53
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This paper takes a webpage classification approach to solve the problem of how to detect whether a publisher webpage contains sensitive content and is appropriate for showing advertisement(s) on it, and designs a unique sensitive content taxonomy.
Abstract
Online advertising has been a popular topic in recent years. In this paper, we address one of the important problems in online advertising, i.e., how to detect whether a publisher webpage contains sensitive content and is appropriate for showing advertisement(s) on it.
Keywords
Computer Science
A Comparative Study on Feature Selection in Text Categorization
4,766 Citations1997Yiming Yang, Jan Pedersen
DF thresholding, the simplest method with the lowest cost in computation, can be reliably used instead of IG or CHI when the computation of these measures are too expensive, and strong correlations between the DF, IG and CHI values of a term are found.
Machine LearningText Classification from Labeled and Unlabeled Documents using EM
2,749 Citations2000Kamal Nigam, Andrew Kachites McCallum +2 more
This paper shows that the accuracy of learned text classifiers can be improved by augmenting a small number of labeled training documents with a large pool of unlabeled documents, and presents two extensions to the algorithm that improve classification accuracy under these conditions.
Hierarchical classification of Web content
805 Citations2000Susan Dumais, Hao Chen
This paper explores the use of hierarchical structure for classifying a large, heterogeneous collection of web content using support vector machine (SVM) classifiers, which have been shown to be efficient and effective for classification, but not previously explored in the context of hierarchical classification.
Employing EM and Pool-Based Active Learning for Text Classification
681 Citations1998Andrew McCallum, Kamal Nigam
The Query-by-Committee method of active learning is modified to use the unlabeled pool for explicitly estimating document density when selecting examples for labeling, and the combination of EM and active learning requires only slightly more than half as many labeled training examples to achieve the same accuracy as either the improved active learning or EM alone.
Hierarchical document categorization with support vector machines
379 Citations2004Lijuan Cai, Thomas Hofmann
A novel hierarchical classification method that generalizes Support Vector Machine learning and that is based on discriminant functions that are structured in a way that mirrors the class hierarchy is proposed.
Finding advertising keywords on web pages
336 Citations2006Wen-tau Yih, Joshua Goodman +1 more
A system that learns how to extract keywords from web pages for advertisement targeting, using a number of features, such as term frequency of each potential keyword, inverse document frequency, presence in meta-data, and how often the term occurs in search query logs.
ACM SIGKDD Explorations NewsletterSupport vector machines classification with a very large-scale taxonomy
225 Citations2005Tie‐Yan Liu, Yiming Yang +4 more
The first evaluation of Support Vector Machines in web-page classification over the full taxonomy of the Yahoo! categories found that the hierarchical use of SVMs is efficient enough for very large-scale classification; however, in terms of effectiveness, the performance of SVM over the Yahoo!. Directory is still far from satisfactory, which indicates that more substantial investigation is needed.
Sequential conditional Generalized Iterative Scaling
79 Citations2001Joshua Goodman
A speedup for training conditional maximum entropy models is described, a simple variation on Generalized Iterative Scaling, but converges roughly an order of magnitude faster, depending on the number of constraints, and the way speed is measured.
Reducing the human overhead in text categorization
57 Citations2006Arnd Christian König, Eric Brill
This work constructs a hybrid classifier that utilizes human reasoning over automatically discovered text patterns to complement machine learning, and demonstrates that the resulting technique results in significant reduction of the human effort required to obtain a given classification accuracy.
