Bilingual Word Embeddings for Phrase-Based Machine Translation
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A method to learn bilingual embeddings from a large unlabeled corpus, while utilizing MT word alignments to constrain translational equivalence is proposed, which significantly out-perform baselines in word semantic similarity.
Abstract
We introduce bilingual word embeddings: se-mantic embeddings associated across two lan-guages in the context of neural language mod-els. We propose a method to learn bilingual embeddings from a large unlabeled corpus, while utilizing MT word alignments to con-strain translational equivalence. The new em-beddings significantly out-perform baselines in word semantic similarity. A single semantic similarity feature induced with bilingual em-beddings adds near half a BLEU point to the results of NIST08 Chinese-English machine translation task. 1
