Exploring linguistic features for web spam detection
Published 22 April 2008
Jakub Piskorski, Marcin Sydow, Dawid Weiss
Citations61
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
Preliminary analysis seems to indicate that certain linguistic features may be useful for the spam-detection task when combined with features studied elsewhere.
Abstract
We study the usability of linguistic features in the Web spam classification task. The features were computed on two Web spam corpora: Webspam-Uk2006 and Webspam-Uk2007, we make them publicly available for other researchers. Preliminary analysis seems to indicate that certain linguistic features may be useful for the spam-detection task when combined with features studied elsewhere.
Keywords
Computer Science
SENTIWORDNET: A Publicly Available Lexical Resource for Opinion Mining
2,489 Citations2006Andrea Esuli, Fabrizio Sebastiani
SENTIWORDNET is a lexical resource in which each WORDNET synset is associated to three numerical scores Obj, Pos and Neg, describing how objective, positive, and negative the terms contained in the synset are.
Detecting spam web pages through content analysis
612 Citations2006Alexandros Ntoulas, Marc Najork +2 more
Some previously-undescribed techniques for automatically detecting spam pages are considered, and the effectiveness of these techniques in isolation and when aggregated using classification algorithms is examined.
Group Decision and NegotiationAutomating Linguistics-Based Cues for Detecting Deception in Text-Based Asynchronous Computer-Mediated Communications
441 Citations2004Lina Zhou, Judee K. Burgoon +2 more
A test of the selected LBC in a simulated TA-CMC experiment showed that a systematic analysis of linguistic information could be useful in the detection of deception and some newly discovered linguistic constructs and their component LBC were helpful in differentiating deception from truth.
Know your neighbors
313 Citations2007Carlos Castillo, Debora Donato +3 more
A spam detection system that combines link-based and content-based features, and uses the topology of the Web graph by exploiting the link dependencies among the Web pages, which finds that linked hosts tend to belong to the same class.
Spam, damn spam, and statistics
293 Citations2004Dennis Fetterly, Mark S. Manasse +1 more
This paper proposes that some spam web pages can be identified through statistical analysis, and examines a variety of properties, including linkage structure, page content, and page evolution, and finds that outliers in the statistical distribution of these properties are highly likely to be caused by web spam.
Blocking blog spam with language model disagreement
210 Citations2005Gilad Mishne, David Carmel +1 more
An approach for detecting link spam common in blog comments by comparing the language models used in the blog post, the comment, and pages linked by the comments, which requires no training, no hard-coded rule sets, and no knowledge of complete-web connectivity.
ACM SIGIR ForumA reference collection for web spam
205 Citations2006Carlos Castillo, Debora Donato +5 more
This is the first publicly available Web spam collection that includes page contents and links, and that has been labelled by a large and diverse set of judges.
Detecting phrase-level duplication on the world wide web
127 Citations2005Dennis Fetterly, Mark S. Manasse +1 more
The algorithms used to discover a number of other instances of large-scale phrase-level replication within the two data sets collected in December 2002 and June 2004 are described.
Lecture notes in computer scienceThwarting the Nigritude Ultramarine: Learning to Identify Link Spam
71 Citations2005Isabel Drost, Tobias Scheffer
The problem of identifying link spam is formulated and a methodology for generating training data is discussed and experiments reveal the effectiveness of classes of intrinsic and relational attributes and shed light on the robustness of classifiers against obfuscation of attributes by an adversarial spammer.
Adversarial Information Retrieval on the WebTracking Web Spam with Hidden Style Similarity
44 Citations2006Tanguy Urvoy, Thomas Lavergne +1 more
This paper presents a (hidden) style similarity based on extra-textual features in html source code and describes a method to clusterize a large collection of documents according to this measure.
Web spam detection via commercial intent analysis
33 Citations2007András A. Benczúr, István Bíró +2 more
A number of features for Web spam filtering based on the occurrence of keywords that are either of high advertisement value or highly spammed are proposed, which improve the classification accuracy of the publicly available WEBSPAM-UK2006 features by 3%.
WITCH: A NEW APPROACH TO WEB SPAM DETECTION
33 Citations2008Jacob Abernethy
An algorithm that learns to detect spam hosts or pages on the Web that simultaneously exploits the structure of the Web graph as well as page contents and features and provides state-of-the-art accuracy on a standard Web spam benchmark.
