Exploiting internal and external semantics for the clustering of short texts using world knowledge
Published 2 November 2009Open access
Xia Hu, Nan Sun, Chao Zhang, Tat‐Seng Chua
Citations249
SJR quartileQ4
SJR score0.11
SNIP0.06
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The proposed method employs a hierarchical three-level structure to tackle the data sparsity problem of original short texts and reconstruct the corresponding feature space with the integration of multiple semantic knowledge bases -- Wikipedia and WordNet.
Abstract
10.1145/1645953.1646071
Keywords
Computer Science
Elsevier eBooksData Mining: Practical Machine Learning Tools and Techniques
25,718 Citations2011Ian H. Witten, Eibe Frank
Program electronic library and information systemsAn algorithm for suffix stripping
8,136 Citations1980Martin Porter
An algorithm for suffix stripping is described, which has been implemented as a short, fast program in BCPL, and performs slightly better than a much more elaborate system with which it has been compared.
Computing semantic relatedness using Wikipedia-based explicit semantic analysis
1,990 Citations2007Evgeniy Gabrilovich, Shaul Markovitch
This work proposes Explicit Semantic Analysis (ESA), a novel method that represents the meaning of texts in a high-dimensional space of concepts derived from Wikipedia that results in substantial improvements in correlation of computed relatedness scores with human judgments.
Mining the peanut gallery
1,908 Citations2003Kushal Dave, Steve Lawrence +1 more
This work develops a method for automatically distinguishing between positive and negative reviews and draws on information retrieval techniques for feature extraction and scoring, and the results for various metrics and heuristics vary depending on the testing situation.
Reexamining the cluster hypothesis
828 Citations1996Marti A. Hearst, Jan Pedersen
This work systematically evaluates Scatter/Gather in this context and finds significant improvements over similarity search ranking alone and provides evidence validating the cluster hypothesis which states that relevant documents tend to be more similar to each other than to non-relevant documents.
Three generative, lexicalised models for statistical parsing
749 Citations1997Michael Collins
A new statistical parsing model is proposed, which is a generative model of lexicalised context-free grammar and extended to include a probabilistic treatment of both subcategorisation and wh-movement.
A web-based kernel function for measuring the similarity of short text snippets
743 Citations2006Mehran Sahami, Timothy D. Heilman
This paper defines a similarity kernel function, mathematically analyze some of its properties, and provides examples of its efficacy, and shows the use of this kernel function in a large-scale system for suggesting related queries to search engine users.
Learning to classify short and sparse text & web with hidden topics from large-scale data collections
739 Citations2008Xuan-Hieu Phan, Le-Minh Nguyen +1 more
A general framework for building classifiers that deal with short and sparse text & Web segments by making the most of hidden topics discovered from large-scale data collections that is general enough to be applied to different data domains and genres ranging from Web search results to medical text.
Computer NetworksGrouper: a dynamic clustering interface to Web search results
716 Citations1999Oren Zamir, Oren Etzioni
This paper introduces Grouper, an interface to the results of the HuskySearch meta-search engine, which dynamically groups the search results into clusters labeled by phrases extracted from the snippets, and reports on the first empirical comparison of user Web search behavior on a standard ranked-list presentation versus a clustered presentation.
Learning to cluster web search results
582 Citations2004Hua-Jun Zeng, Qi-Cai He +3 more
This paper reformalizes the clustering problem as a salient phrase ranking problem, and first extracts and ranks salient phrases as candidate cluster names, based on a regression model learned from human labeled training data.
Measuring semantic similarity between words using web search engines
544 Citations2007
A robust semantic similarity measure that uses the information available on the Web to measure similarity between words or entities and a novel approach to compute semantic similarity using automatically extracted lexico-syntactic patterns from text snippets is proposed.
Overcoming the brittleness bottleneck using wikipedia: enhancing text categorization with encyclopedic knowledge
404 Citations2006Evgeniy Gabrilovich, Shaul Markovitch
It is proposed to enrich document representation through automatic use of a vast compendium of human knowledge--an encyclopedia, and empirical results confirm that this knowledge-intensive representation brings text categorization to a qualitatively new level of performance across a diverse collection of datasets.
Text classification and named entities for new event detection
364 Citations2004Giridhar Kumaran, James Allan
This paper shows how performance on New Event Detection (NED) can be improved by the use of text classification techniques as well as by using named entities in a new way, and explores modifications to the document representation in a vector space-based NED system.
Ontologies improve text document clustering
353 Citations2004Andreas Hotho, Steffen Staab +1 more
This work integrates core ontologies as background knowledge into the process of clustering text documents and compares clustering techniques based on pre-categorizations of texts from Reuters newsfeeds and on a smaller domain of an eLearning course about Java.
Clustering short texts using wikipedia
335 Citations2007Somnath Banerjee, Krishnan Ramanathan +1 more
A method of improving the accuracy of clustering short texts by enriching their representation with additional features from Wikipedia is proposed and empirical results indicate that this enriched representation of text items can substantially improve the clustering accuracy when compared to the conventional bag of words representation.
TUbilio (Technical University of Darmstadt)Extracting Lexical Semantic Knowledge from Wikipedia and Wiktionary
326 Citations2008Torsten Zesch, Christof Müller +1 more
This paper presents two application programming interfaces for Wikipedia and Wiktionary which are especially designed for mining the rich lexical semantic information dispersed in the knowledge bases, and provide efficient and structured access to the available knowledge.
Lecture notes in computer scienceSimilarity Measures for Short Segments of Text
304 Citations2007Donald Metzler, Susan Dumais +1 more
This work formally evaluate and analyze the methods on a query-query similarity task using 363,822 queries from a web search log, and provides insights into the strengths and weaknesses of each method, including important tradeoffs between effectiveness and efficiency.
Lingo: Search Results Clustering Algorithm Based on Singular Value Decomposition
281 Citations2004Stanisław Osiński, Jerzy Stefanowski +1 more
This paper presents Lingo—a novel algorithm for clustering search results, which emphasizes cluster description quality, and describes methods used in the algorithm: algebraic transformations of the term-document matrix and frequent phrase extraction using suffix arrays.
WordNet improves text document clustering
272 Citations2003Andreas Hotho, Steffen Staab +1 more
This work integrates background knowledge — in the authors' application Wordnet — into the process of clustering text documents, and clusters the documents by a standard partitional algorithm.
Feature generation for text categorization using world knowledge
232 Citations2005Evgeniy Gabrilovich, Shaul Markovitch
Improved machine learning algorithms for text categorization with generated features based on domain-specific and common-sense knowledge are enhanced, addressing the two main problems of natural language processing--synonymy and polysemy.
Constant interaction-time scatter/gather browsing of very large document collections
206 Citations1993Douglass R. Cutting, David R. Karger +1 more
This work presents a scheme that supports constant interaction-time Scatter/Gather of arbitrarily large collections after near-linear time preprocessing, and involves the construction of a cluster hierarchy.
Frequency estimates for statistical word similarity measures
197 Citations2003Egidio L. Terra, Charles L. A. Clarke
A comparative study of two methods for estimating word co-occurrence frequencies required by word similarity measures, generated from a terabyte-sized corpus of Web data.
Enhancing text clustering by leveraging Wikipedia semantics
196 Citations2008Jian Hu, Lujun Fang +5 more
A way to build a concept thesaurus based on the semantic relations (synonym, hypernym, and associative relation) extracted from Wikipedia is proposed and a unified framework to leverage these semantic relations in order to enhance traditional content similarity measure for text clustering is developed.
ACM SIGIR ForumThe Wikipedia XML corpus
195 Citations2006Ludovic Denoyer, Patrick Gallinari
This encyclopedia is composed of millions of articles in different languages and anyone can edit an article using a wiki markup language that offers a simplified alternative to HTML.
IEEE Transactions on Knowledge and Data EngineeringEfficient Phrase-Based Document Similarity for Clustering
158 Citations2008Hung Chim, Xiaotie Deng
The phrase-based document similarity is applied to the group-average Hierarchical Agglomerative Clustering (HAC) algorithm and the new clustering approach is developed, which is very effective on clustering the documents of two standard document benchmark corpora OHSUMED and RCV1.
Novel association measures using web search with double checking
156 Citations2006Hsin‐Hsi Chen, Ming-Shun Lin +1 more
Five association measures including variants of Dice, Overlap Ratio, Jaccard, and Cosine, as well as Co-Occurrence Double Check (CODC), are presented and show that the five measures are quite useful.
Second language ResearchPsycholinguistic techniques in second language acquisition research
133 Citations2003Theodoros Marinis
The hardware and software packages and other equipment required for the setting-up of a psycholinguistics laboratory, the advantages and disadvantages of the software packages available and what financial costs are involved are discussed.
Term clustering of syntactic phrases
100 Citations1989David Lewis, W. Bruce Croft
This paper discusses the implementation of a syntactic phrase generator, as well as the preliminary experiments with producing phrase clusters, and shows small improvements in retrieval effectiveness resulting from the use of phrase clusters.
Using the web to overcome data sparseness
97 Citations2002Frank Keller, Maria Lapata +1 more
It is shown that the web can be employed to obtain frequencies for bigrams that are unseen in a given corpus by demonstrating that web frequencies and correlate with frequencies obtained from a carefully edited, balanced corpus.
Learning semantic classes for word sense disambiguation
67 Citations2005Upali Kohomban, Wee Sun Lee
It is shown that these general concepts from a sense tagged corpus can be transformed to fine grained word senses using simple heuristics, and applying the technique for recent SENSEVAL data sets shows that this approach can yield state of the art performance.
Applied Physics Letters10.1162/153244302320884533
62 Citations2000James Hammerton, Miles Osborne +2 more
The origins of shallow parsing as a specific task for machine learning of language, the articles accepted for this special issue, a representative sample of current research in this area, and future directions for machine learning of shallow parsing are suggested.
Computers and the HumanitiesIntegrating Linguistic Resources in TC through WSD
41 Citations2001Luís Alfonso Ureña López, Manuel de Buenaga Rodríguez +1 more
An approach to TC based on the integration of a training collection and a lexical database as knowledge sources is described and the utilization of WSD is presented as an aid for TC.
Journal of Machine Learning ResearchIntroduction to special issue on machine learning approaches to shallow parsing
30 Citations2002HammertonJames, OsborneMiles +2 more
The problem of partial or shallow parsing (assigning partial syntactic structure to sentences) is introduced and why it is an important natural language processing (NLP) task is explained.
Query segmentation based on eigenspace similarity
17 Citations2009Chao Zhang, Nan Sun +3 more
A novel unsupervised learning approach to query segmentation based on principal eigenspace similarity of query-word-frequency matrix derived from web statistics is presented.
Using digest pages to increase user result space: Preliminary designs
15 Citations2008Shanu Sushmita, Mounia Lalmas +1 more
This paper presents preliminary designs regarding the construction, the presentation, and the ranking of digest pages and shows how digest pages can be used to capture the context of a query through the concept of an aggregated digest page, which is based on the aggregated search paradigm offered by some search engines.
