A reference collection for web spam
ACM SIGIR ForumPublished 1 December 2006
Carlos Castillo, Debora Donato, Luca Becchetti, Paolo Boldi, Stefano Leonardi, Massimo Santini
Citations205
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This is the first publicly available Web spam collection that includes page contents and links, and that has been labelled by a large and diverse set of judges.
Abstract
We describe the WEBSPAM-UK2006 collection, a large set of Web pages that have been manually annotated with labels indicating if the hosts are include Web spam aspects or not. This is the first publicly available Web spam collection that includes page contents and links, and that has been labelled by a large and diverse set of judges.
Keywords
Computer Science
The PageRank Citation Ranking : Bringing Order to the Web
12,645 Citations1999Lawrence M. Page, Sergey Brin +2 more
This paper describes PageRank, a mathod for rating Web pages objectively and mechanically, effectively measuring the human interest and attention devoted to them, and shows how to efficiently compute PageRank for large numbers of pages.
The webgraph framework I
1,215 Citations2004Paolo Boldi, Sebastiano Vigna
This papers presents the compression techniques used in WebGraph, which are centred around referentiation and intervalisation (which in turn are dual to each other).
Elsevier eBooksCombating Web Spam with TrustRank
1,025 Citations2004Zoltán Gyöngyi, Héctor García-Molina +1 more
This paper proposes techniques to semi-automatically separate reputable, good pages from spam, and shows that they can effectively filter out spam from a significant fraction of the web, based on a good seed set of less than 200 sites.
Detecting spam web pages through content analysis
612 Citations2006Alexandros Ntoulas, Marc Najork +2 more
Some previously-undescribed techniques for automatically detecting spam pages are considered, and the effectiveness of these techniques in isolation and when aggregated using classification algorithms is examined.
Software Practice and ExperienceUbiCrawler: a scalable fully distributed Web crawler
564 Citations2004Paolo Boldi, Bruno Codenotti +2 more
The main features of UbiCrawler are platform independence, linear scalability, graceful degradation in the presence of faults, a very effective assignment function for partitioning the domain to crawl, and more in general the complete decentralization of every task.
Adversarial Information Retrieval on the WebWeb Spam Taxonomy
488 Citations2005Zoltán Gyöngyi, Héctor García-Molina
This paper presents a comprehensive taxonomy of current spamming techniques, which it is believed can help in developing appropriate countermeasures.
Ranking the web frontier
212 Citations2004Nadav Eiron, Kevin S. McCurley +1 more
This paper analyzes features of the rapidly growing "frontier" of the web, namely the part of theweb that crawlers are unable to cover for one reason or another, and suggests ways to improve the quality of ranking by modeling the growing presence of "link rot" on the web as more sites and pages fall out of maintenance.
SpamRank -- Fully Automatic Link Spam Detection.
171 Citations2005András A. Benczúr, Károly Csalogány +2 more
Recognizing Nepotistic Links on the Web
150 Citations2000Brian D. Davison
High accuracy in initial experiments is reported to show the potential for using a machine learning tool to automatically recognize and eliminate nepotistic links— links between pages that are present for reasons other than merit.
Using rank propagation and Probabilistic counting for Link-Based Spam Detection
101 Citations2006Luca Becchetti, Carlos Castillo +3 more
This paper proposes spam detection techniques that only consider the link structure of Web, regardless of page contents, and compute statistics of the links in the vicinity of every Web page applying rank propagation and probabilistic counting over the Web graph.
Link-Based Similarity Search to Fight Web Spam
60 Citations2006András A. Benczúr, Károly Csalogány +1 more
This work forms classifiers by investigating similarity top lists of an unknown page along various measures such as co-citation, companion, nearest neighbors in low dimensional projections and SimRank and test the method over two data sets previously used to measure spam filtering algorithms.
Communications of the ACMThe bubble of web visibility
52 Citations2005Marco Gori, Ian H. Witten
Promoting visibility as seen through the unique lens of search engines will help improve visibility in the rapidly changing environment.
Detecting nepotistic links by language model disagreement
25 Citations2006András A. Benczúr, István Bíró +2 more
The applicability of hyperlink downweighting by means of language model disagreement is demonstrated and various forms of nepotism such as common maintainers, ads, link exchanges or misused affiliate programs are fought.
