Finding your way through blogspace: Using semantics for cross-domain blog analysis
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This paper proposes two methods for semantics-enhanced blogs analysis that allow the analyst to integrate domain-specific as well as general background knowledge, and presents a detailed experimental analysis of a sample of four sets of blogs with different thematic foci.
Abstract
Blogspace is one of the most dynamic areas of today’s Internet, and it is increasingly recognised that blogs are much more than “meaningless chatter”. Many syntax-based approaches exist to analyse the text and the net-work structure between blogs. While this is very helpful for purposes such as the detection of discussion bursts concerning uniquely-named topics (e.g., a book, prod-uct, or person), it is insufficient for understanding blogs discussing new phenomena in different wordings, or for finding and explaining relationships between new dis-course topics or the context of a new topic in a larger domain of discourse. In this paper, we propose two methods for semantics-enhanced blogs analysis that al-low the analyst to integrate domain-specific as well as general background knowledge. The methods rely on the Term Extractor for identifying keyphrases (Navigli & Velardi, 2004), SSI (Structural Semantic Intercon-nections) for disambiguating terms (Navigli & Velardi, 2005), and the taxonomy of domain labels by (Magnini & Cavaglià, 2000). Applications include topic detection and grouping, the proposal of blog tags and the forming of blog directories, and blog recommender systems. To illustrate the usefulness of our approach, we present a detailed experimental analysis of a sample of four sets of blogs with different thematic foci (food, health, law, and weblogs about blogging).
