login

Improving Distributional Similarity with Lessons Learned from Word Embeddings

Transactions of the Association for Computational LinguisticsPublished 1 December 2015Open access
Omer Levy, Yoav Goldberg, Ido Dagan
Citations1,342
SJR quartileQ1
SJR score1.82
SNIP4.13
View PDF

TL;DR

It is revealed that much of the performance gains of word embeddings are due to certain system design choices and hyperparameter optimizations, rather than the embedding algorithms themselves, and these modifications can be transferred to traditional distributional models, yielding similar gains.

Abstract

Recent trends suggest that neural-network-inspired word embedding models outperform traditional count-based distributional models on word similarity and analogy detection tasks. We reveal that much of the performance gains of word embeddings are due to certain system design choices and hyperparameter optimizations, rather than the embedding algorithms themselves. Furthermore, we show that these modifications can be transferred to traditional distributional models, yielding similar gains. In contrast to prior reports, we observe mostly local or insignificant performance differences between the methods, with no global advantage to any single approach over the others.

Keywords

Computer Science