login

Semantic expansion using word embedding clustering and convolutional neural network for improving short text classification

NeurocomputingPublished 18 October 2015
Peng Wang, Bo Xu, Jiaming Xu, Guanhua Tian, Cheng‐Lin Liu, Hongwei Hao
Citations305
SJR quartileQ1
SJR score1.47
SNIP1.94

TL;DR

A unified framework to expand short texts based on word embedding clustering and convolutional neural network and semantic cliques via fast clustering is proposed, which validates the effectiveness of the proposed method on two open benchmarks.

Abstract

Text classification can help users to effectively handle and exploit useful information hidden in large-scale documents. However, the sparsity of data and the semantic sensitivity to context often hinder the classification performance of short texts. In order to overcome the weakness, we propose a unified framework to expand short texts based on word embedding clustering and convolutional neural network (CNN). Empirically, the semantically related words are usually close to each other in embedding spaces. Thus, we first discover semantic cliques via fast clustering. Then, by using additive composition over word embeddings from context with variable window width, the representations of multi-scale semantic units1 in short texts are computed. In embedding spaces, the restricted nearest word embeddings (NWEs)2 of the semantic units are chosen to constitute expanded matrices, where the semantic cliques are used as supervision information. Finally, for a short text, the projected matrix3 and expanded matrices are combined and fed into CNN in parallel. Experimental results on two open benchmarks validate the effectiveness of the proposed method.

Keywords

Computer Science