Modeling and Data Mining in Blogosphere
Synthesis lectures on data mining and knowledge discoveryPublished 1 January 2009
Nitin Agarwal, Huan Liu
Citations72
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
Spam blogs or "splogs" are an increasing concern in Blogosphere and are discussed in detail with the approaches leveraging supervised machine learning algorithms and interaction patterns.
Abstract
This book offers a comprehensive overview of the various concepts and research issues about blogs or weblogs. It introduces techniques and approaches, tools and applications, and evaluation methodolog
Keywords
Computer SciencePhysics and Astronomy
Journal of Machine Learning ResearchLatent dirichlet allocation
27,049 Citations2003David M. Blei, Andrew Y. Ng +1 more
Computer Networks and ISDN SystemsThe anatomy of a large-scale hypertextual Web search engine
15,828 Citations1998Sergey Brin, Lawrence M. Page
This paper provides an in-depth description of Google, a prototype of a large-scale search engine which makes heavy use of the structure present in hypertext and looks at the problem of how to effectively deal with uncontrolled hypertext collections where anyone can publish anything they want.
IEEE Transactions on Pattern Analysis and Machine IntelligenceNormalized cuts and image segmentation
15,696 Citations2000Jianbo Shi, Jitendra Malik
Journal of the American Society for Information ScienceIndexing by latent semantic analysis
12,677 Citations1990Scott Deerwester, Susan Dumais +3 more
Statistics and ComputingA tutorial on spectral clustering
10,269 Citations2007Ulrike von Luxburg
This tutorial describes different graph Laplacians and their basic properties, present the most common spectral clustering algorithms, and derive those algorithms from scratch by several different approaches.
Information Processing & ManagementTerm-weighting approaches in automatic text retrieval
9,532 Citations1988Gerard Salton, Chris Buckley
This paper summarizes the insights gained in automatic term weighting, and provides baseline single term indexing models with which other more elaborate content analysis procedures can be compared.
On Spectral Clustering: Analysis and an algorithm
7,751 Citations2001Andrew Y. Ng, Michael I. Jordan +1 more
A simple spectral clustering algorithm that can be implemented using a few lines of Matlab is presented, and tools from matrix perturbation theory are used to analyze the algorithm, and give conditions under which it can be expected to do well.
Maximizing the spread of influence through a social network
7,314 Citations2003David Kempe, Jon Kleinberg +1 more
An analysis framework based on submodular functions shows that a natural greedy strategy obtains a solution that is provably within 63% of optimal for several classes of models, and suggests a general approach for reasoning about the performance guarantees of algorithms for these types of influence problems in social networks.
New York University Press eBooks4 What Is Web 2.0? Design Patterns and Business Models for the Next Generation of Software
5,654 Citations2012Tim O’Reilly
On power-law relationships of the Internet topology
4,222 Citations1999Michalis Faloutsos, Petros Faloutsos +1 more
These power-laws hold for three snapshots of the Internet, between November 1997 and December 1998, despite a 45% growth of its size during that period, and can be used to generate and select realistic topologies for simulation purposes.
Cambridge University Press eBooksRandomized Algorithms
4,069 Citations1995Rajeev Motwani, Prabhakar Raghavan
Choice Reviews OnlineThe long tail: why the future of business is selling less of more
2,933 Citations2007Chris Anderson
The central premise of the book is that the combination of the Pareto or Zipf distribution that is characteristic of Web traffic and the direct access to consumers via Web technology has opened up new business opportunities in the ''long tail''.
Cost-effective outbreak detection in networks
2,446 Citations2007Jure Leskovec, Andreas Krause +4 more
This work exploits submodularity to develop an efficient algorithm that scales to large problems, achieving near optimal placements, while being 700 times faster than a simple greedy algorithm and achieving speedups and savings in storage of several orders of magnitude.
Graphs over time
2,264 Citations2005Jure Leskovec, Jon Kleinberg +1 more
A new graph generator is provided, based on a "forest fire" spreading process, that has a simple, intuitive justification, requires very few parameters (like the "flammability" of nodes), and produces graphs exhibiting the full range of properties observed both in prior work and in the present study.
Contemporary Sociology A Journal of ReviewsTrust: A Sociological Theory
2,251 Citations2001Edward W. Lehman, Piotr Sztompka
Feature Selection for Knowledge Discovery and Data Mining
2,191 Citations1998Huan Liu, Hiroshi Motoda
Feature Selection for Knowledge Discovery and Data Mining offers an overview of the methods developed since the 1970's and provides a general framework in order to examine these methods and categorize them and suggests guidelines for how to use different methods under various circumstances.
Marketing LettersTalk of the Network: A Complex Systems Look at the Underlying Process of Word-of-Mouth
1,995 Citations2001Jacob Goldenberg, Barak Libai +1 more
The results clearly indicate that information dissemination is dominated by both weak and strong w-o-m, rather than by advertising, which means that strong and weak ties become the main forces propelling growth.
Journal of Consumer ResearchInfluentials, Networks, and Public Opinion Formation
1,899 Citations2007Duncan J. Watts, Peter Sheridan Dodds
Symposium on Discrete AlgorithmsAuthoritative sources in a hyperlinked environment
1,836 Citations1998Jon Kleinberg
Mining knowledge-sharing sites for viral marketing
1,635 Citations2002Matthew Richardson, Pedro Domingos
This research optimize the amount of marketing funds spent on each customer, rather than just making a binary decision on whether to market to him, and takes into account the fact that knowledge of the network is partial, and that gathering that knowledge can itself have a cost.
Propagation of trust and distrust
1,471 Citations2004R. Guha, Ravi Kumar +2 more
It is shown that a small number of expressed trusts/distrust per individual allows us to predict trust between any two people in the system with high accuracy.
IEEE Transactions on Computer-Aided Design of Integrated Circuits and SystemsNew spectral methods for ratio cut partitioning and clustering
1,279 Citations1992L. Hagen, Andrew B. Kahng
It is shown that the second smallest eigenvalue of a matrix derived from the netlist gives a provably good approximation of the optimal ratio cut partition cost.
We the media grassroots journalism by the people, for the people
1,249 Citations2006Dan Gillmor
Elsevier eBooksCombating Web Spam with TrustRank
1,025 Citations2004Zoltán Gyöngyi, Héctor García-Molina +1 more
This paper proposes techniques to semi-automatically separate reputable, good pages from spam, and shows that they can effectively filter out spam from a significant fraction of the web, based on a good seed set of less than 200 sites.
ACM SIGKDD Explorations NewsletterInformation diffusion through blogspace
951 Citations2004Daniel Gruhl, David Liben‐Nowell +2 more
A macroscopic characterization of topic propagation through the authors' corpus, formalizing the notion of long-running "chatter" topics consisting recursively of "spike" topics generated by outside world events, or more rarely, by resonances within the community.
Regional conference series in mathematicsComplex Graphs and Networks
702 Citations2006Fan Chung, Linyuan Lü
Detecting spam web pages through content analysis
612 Citations2006Alexandros Ntoulas, Marc Najork +2 more
Some previously-undescribed techniques for automatically detecting spam pages are considered, and the effectiveness of these techniques in isolation and when aggregated using classification algorithms is examined.
Identifying the influential bloggers in a community
511 Citations2008Nitin Agarwal, Huan Liu +2 more
The challenges of identifying influential bloggers are discussed, what constitutes influential bloggers is investigated, a preliminary model attempting to quantify an influential blogger is presented, and the way for building a robust model that allows for finding various types of the influentials is paved.
ACM Transactions on Internet TechnologyInferring binary trust relationships in Web-based social networks
434 Citations2006Jennifer Golbeck, James Hendler
A definition of trust suitable for use in Web-based social networks with a discussion of the properties that will influence its use in computation is introduced and two algorithms for inferring trust relationships between individuals that are not directly connected in the network are presented.
The predictive power of online chatter
411 Citations2005Daniel Gruhl, Ramanathan V. Guha +3 more
First, carefully hand-crafted queries produce matching postings whose volume predicts sales ranks, and even though sales rank motion might be difficult to predict in general, algorithmic predictors can use online postings to successfully predict spikes in sales rank.
Evolutionary spectral clustering by incorporating temporal smoothness
369 Citations2007Yün Chi, Xiaodan Song +3 more
This paper proposes two frameworks that incorporate temporal smoothness in evolutionary spectral clustering and demonstrates that their methods provide the optimal solutions to the relaxed versions of the corresponding evolutionary k-means clustering problems.
ACM Transactions on Computer-Human InteractionSocial matching
322 Citations2005Loren Terveen, David W. McDonald
The scope of social matching systems is clarified by distinguishing them from other recommender systems and related systems and techniques and shows how existing systems explore different points within the design space.
Improved annotation of the blogosphere via autotagging and hierarchical clustering
304 Citations2006Christopher Brooks, Nancy Montanez
It is found that tags are useful for grouping articles into broad categories, but less effective in indicating the particular content of an article, and clustering algorithms can be used to reconstruct a topical hierarchy among tags.
Statistical analysis of the social network and discussion threads in slashdot
298 Citations2008Vicenç Gómez, Andreas Kaltenbrunner +1 more
The social network emerging from the user comment activity on the website Slashdot is analyzed, showing common features of traditional social networks such as a giant component, small average path length and high clustering but differs from them showing moderate reciprocity and neutral assortativity by degree.
Identifying opinion leaders in the blogosphere
238 Citations2007Xiaodan Song, Yün Chi +2 more
The InfluenceRank algorithm ranks blogs according to not only how important they are as compared to other blogs, but also how novel the information they can contribute to the network.
Blocking blog spam with language model disagreement
210 Citations2005Gilad Mishne, David Carmel +1 more
An approach for detecting link spam common in blog comments by comparing the language models used in the blog post, the comment, and pages linked by the comments, which requires no training, no hard-coded rule sets, and no knowledge of complete-web connectivity.
Naked Conversations: How Blogs are Changing the Way Businesses Talk with Customers
182 Citations2006Robert Scoble, Shel Israel
Maryland Shared Open Access Repository (USMAI Consortium)Modeling the Spread of Influence on the Blogosphere
139 Citations2006Akshay Java, Pranam Kolari +2 more
This paper validate the effectiveness of some of the influence models on the blogosphere, and shows how PageRank based heuristics could be used to select an influential set of bloggers such that the authors could maximize the spread of information on theBlogosphere.
Cascading Behavior in Large Blog Graphs
132 Citations2007Jure Leskovec, Mary McGlohon +3 more
The ICWSM 2009 Spinn3r Dataset
126 Citations2009Kevin Burton, Akshay Java +1 more
The dataset, provided by Spinn3r.com, is a set of 44 million blog posts made between August 1st and October 1st, 2008, which spans a number of big news events as well as everything else you might expect to find posted to blogs.
Maryland Shared Open Access Repository (USMAI Consortium)Detecting Spam Blogs: A Machine Learning Approach
124 Citations2006Pranam Kolari, Akshay Java +3 more
It is discussed how SVM models based on local and link-based features can be used to detect splogs and an evaluation of learned models and their utility to blog search engines; systems that employ techniques differing from those of conventional web search engines.
Incremental Spectral Clustering With Application to Monitoring of Evolving Blog Communities
112 Citations2007Huazhong Ning, Wei Xu +3 more
The incremental algorithm, initialized by a standard spectral clustering, continuously and efficiently updates the eigenvalue system and generates instant cluster labels, as the data set is evolving, and achieves similar accuracy but with much lower computational cost.
BlogRank
102 Citations2006Apostolos Kritikopoulos, Martha Sideri +1 more
This work presents a method for ranking weblogs utilizing both link graph and similarity, and based on an enhanced and weighted graph of weblogs capturing crucial weblog features, which is a modified version of PageRank.
Exploring in the weblog space by detecting informative and affective articles
76 Citations2007Xiaochuan Ni, Gui-Rong Xue +3 more
This paper presents a machine learning method for classifying informative and affective articles among weblogs, and develops an intent-driven weblog-search engine based on the classification techniques to improve the satisfaction of Web users.
Modeling Trust and Influence on Blogosphere using Link Polarity
73 Citations2007Anubhav Kale
This paper uses links in the blog graph to associate sentiments with the links connecting two blogs, and uses trust propagation models to “spread” the initial polarity values to all possible pairs of nodes.
BlogScope
72 Citations2007Nilesh Bansal, Nick Koudas
BlogScope is an information discovery and text analysis system that offers a set of unique features that include, spatio-temporal analysis of blogs, flexible navigation of the Blogosphere through information bursts, keyword correlations and burst synopsis, as well as enhanced ranking functions for improved query answer relevance.
Structural and temporal analysis of the blogosphere through community factorization
71 Citations2007Yün Chi, Shenghuo Zhu +3 more
This paper proposes a novel technique that captures the structure and temporal dynamics of blog communities, formulated as a factorization problem in the framework of constrained optimization, in which the objective is to best explain the observed interactions in the blogosphere over time.
Discovering Important Bloggers based on Analyzing Blog Threads
66 Citations2005Shinsuke Nakajima, Junichi Tatemura +3 more
Models of blogs and blog thread data, and methods of extracting blog threads, discovering important bloggers, and acquiring important content from their entries are discussed.
Lecture notes in computer scienceActive Constrained Clustering by Examining Spectral Eigenvectors
65 Citations2005Qianjun Xu, Marie desJardins +1 more
The ACCESS method uses an analysis based on the theoretical properties of spectral decomposition to identify data items that are likely to be located on the boundaries of clusters, and for which providing constraints can resolve ambiguity in the cluster descriptions.
Enhancing clustering blog documents by utilizing author/reader comments
51 Citations2007Beibei Li, Shuting Xu +1 more
It is argued that the author/reader comments of the blog pages may have more discriminating effect in clustering blog documents, and a word-page matrix constructed by downloading blog pages from a well-known website and a k-means clustering algorithm with different weights assigned to the title, body, and comment parts confirms this hypothesis.
ACM Transactions on the WebDetecting splogs via temporal dynamics using self-similarity analysis
50 Citations2008Yu‐Ru Lin, Hari Sundaram +3 more
This article addresses the problem of spam blog (splog) detection using temporal and structural regularity of content, post time and links using a new technique based on the observation that a blog is a dynamic, growing sequence of entries (or posts) rather than a collection of individual pages.
Using Tags and Clustering to Identify Topic-Relevant Blogs
46 Citations2007Conor Hayes
It is demonstrated how tags provide useful, discriminating information where the blog corpus is initially partitioned using a conventional clustering technique, and concluded that tags have a key auxiliary role in refining and confirming the information produced using typical knowledge discovery techniques.
An information structural approach to spoken language generation
46 Citations1996Scott Prevost
A two-tiered information structure representation is used in the high-level content planning and sentence planning stages of generation to produce efficient, coherent speech that makes certain discourse relationships, such as explicit contrasts, appropriately salient.
Building Trust with Corporate Blogs.
45 Citations2007Paul Dwyer
It is suggested that companies can use blogging to complement customer relationship management processes to the extent their customers exhibit an organic desire to commune by combining provocative informational content with expressions of benevolent intent.
Dynamic classification of groups through social network analysis and HMMs
44 Citations2005Thayne Coffman, Sherry Marcus
This work motivates and presents results from a case study using a simulation of suspicious groups communicating in a normal background population, achieving 96% classification accuracy on novel synthetic data using two 35-state univariate HMMs trained to model normal and suspicious evolutions of the characteristic path length metric.
Proceedings of the International AAAI Conference on Web and Social MediaA Social Identity Approach to Identify Familiar Strangers in a Social Network
42 Citations2009Nitin Agarwal, Huan Liu +3 more
This work forms the problem, shows why it is significant to address the challenge, and presents an approach that innovatively employs the social identities of the individuals with competitive approaches.
Maryland Shared Open Access Repository (USMAI Consortium)Detecting Commmunities via Simultaneous Clustering of Graphs and Folksonomies
31 Citations2008Akshay Java, Anupam Joshi +1 more
The Simultaneous Cut (SimCut) method has the advantage that it can group related tags and cluster the nodes simultaneously and can be easily and efficiently implemented.
Chinese Physics LettersTopological Properties and Transition Features Generated by a New Hybrid Preferential Model
27 Citations2005Fang Jin-Qing, Liang Yong
A new hybrid preferential model (HPM) is proposed for generating both scale-free and small world properties and it is found that when the ratio d/r increases, the average path length L is decreased, while the average clustering coefficient C is increased.
XRDS Crossroads The ACM Magazine for StudentsExploring global terrorism data
26 Citations2008Joong-Hoon Lee
The GTD Explorer is introduced, a web-based interactive visual exploratory tool that deals with this data, which counts the number of incidents grouped over a certain criteria, and stack charts on top of each other to see both individual and accumulated patterns of incidents over time.
Deriving wishlists from blogs show us your blog, and we'll tell you what books to buy
23 Citations2006Gilad Mishne, Maarten de Rijke
Text analysis and external knowledge sources are used to estimate the commercial taste of bloggers from their text, showing that valuable insights can be mined from blogs not just at the aggregate but also at the individual blog level.
Clustering Blogs with Collective Wisdom
23 Citations2008Nitin Agarwal, Magdiel Galan +2 more
This work proposes to tap into collective wisdom in clustering blog sites, present statistical and visual results, report findings, and suggest future work extending to many real-world applications.
Characterizing Network Motifs to Identify Spam Comments
15 Citations2008Elaheh Kamaliha, Fatemeh Riahi +2 more
The functionality of different network motif profiles in the comment network are looked at, and it is illustrated that some of these patterns and their statistical features could be exploited to classify comments and bloggers to spammers and non-spammers.
Proceedings of the International AAAI Conference on Web and Social MediaBlogTrackers: A Tool for Sociologists to Track and Analyze Blogosphere
15 Citations2009Nitin Agarwal, Shamanth Kumar +2 more
An overview of BlogTrackers is presented, its functions of various components are illustrated, and future work for expansion is outlined in meeting the growing needs of sociologists.
Online spam-blog detection through blog search
13 Citations2008Linhong Zhu, Aixin Sun +1 more
A novel post-indexing spam-blog (or splog) detection method, which capitalizes on the results returned by blog search engines, and builds and maintains Blog profiles for those blogs whose posts frequently appear in the top-ranked search results.
Munich Personal RePEc Archive (Ludwig Maximilian University of Munich)Planning and Assessing Stability Operations: A Proposed Value Focus Thinking Approach
12 Citations2012Gerald D. Fensterer
Representative entry selection for profiling blogs
1 Citations2008Jinfeng Zhuang, Steven C. H. Hoi +2 more
This work forms the entry selection task into a combinatorial optimization problem and proposes a greedy yet effective algorithm for finding a good approximate solution by exploiting the theory of submodular functions.
