Universal Neural Machine Translation for Extremely Low Resource Languages
Published 1 January 2018Open access
Jiatao Gu, Hany Hassan, Jacob Devlin, Victor O. K. Li
Citations241
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The proposed approach utilizing a transfer-learning approach to share lexical and sentence level representations across multiple source languages into one target language is able to achieve 23 BLEU on Romanian-English WMT2016 using a tiny parallel corpus of 6k sentences.
Abstract
Jiatao Gu, Hany Hassan, Jacob Devlin, Victor O.K. Li. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Keywords
Computer Science
Neural ComputationLong Short-Term Memory
98,079 Citations1997Sepp Hochreiter, Jürgen Schmidhuber
A novel, efficient, gradient based method called long short-term memory (LSTM) is introduced, which can learn to bridge minimal time lags in excess of 1000 discrete-time steps by enforcing constant error flow through constant error carousels within special units.
UvA-DARE (University of Amsterdam)Adam: A Method for Stochastic Optimization
84,783 Citations2014Diederik P. Kingma, Jimmy Ba
DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)MizAR 60 for Mizar 50
76,311 Citations2023Jakubův, Jan, Chvalovský, Karel +7 more
DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)
50,318 Citations2021Mandi, Jayanta, Canoy, Rocsildes +2 more
A simple numeric simulation of DNA-co-polymerized hydrogel shape change and a genetic algorithm that generates and selects large batches of material designs that compete with one another to evolve and converge on optimal objective-matching designs are constructed.
arXiv (Cornell University)Neural Machine Translation by Jointly Learning to Align and Translate
14,565 Citations2014Dzmitry Bahdanau
arXiv (Cornell University)Sequence to Sequence Learning with Neural Networks
13,362 Citations2014Ilya Sutskever, Oriol Vinyals +1 more
Transactions of the Association for Computational LinguisticsEnriching Word Vectors with Subword Information
9,784 Citations2017Piotr Bojanowski, Édouard Grave +2 more
A new approach based on the skipgram model, where each word is represented as a bag of character n-grams, with words being represented as the sum of these representations, which achieves state-of-the-art performance on word similarity and analogy tasks.
arXiv (Cornell University)Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
5,755 Citations2017Chelsea Finn, Pieter Abbeel +1 more
arXiv (Cornell University)Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
5,634 Citations2016Yonghui Wu, Mike Schuster +29 more
GNMT, Google's Neural Machine Translation system, is presented, which attempts to address many of the weaknesses of conventional phrase-based translation systems and provides a good balance between the flexibility of "character"-delimited models and the efficiency of "word"-delicited models.
arXiv (Cornell University)Convolutional Sequence to Sequence Learning
1,897 Citations2017Jonas Gehring, Michael Auli +3 more
This work introduces an architecture based entirely on convolutional neural networks, which outperform the accuracy of the deep LSTM setup of Wu et al. (2016) on both WMT'14 English-German and WMT-French translation at an order of magnitude faster speed, both on GPU and CPU.
Transactions of the Association for Computational LinguisticsGoogle’s Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation
1,742 Citations2017Melvin Johnson, Mike Schuster +10 more
This work proposes a simple solution to use a single Neural Machine Translation (NMT) model to translate between multiple languages using a shared wordpiece vocabulary, and introduces an artificial token at the beginning of the input sentence to specify the required target language.
International Conference on Machine LearningConvolutional Sequence to Sequence Learning
1,340 Citations2017Jonas Gehring, Michael Auli +3 more
Proceedings of the AAAI Conference on Artificial IntelligenceCharacter-Aware Neural Language Models
1,046 Citations2016Yoon Kim, Yacine Jernite +2 more
A simple neural language model that relies only on character-level inputs that is able to encode, from characters only, both semantic and orthographic information and suggests that on many languages, character inputs are sufficient for language modeling.
arXiv (Cornell University)Character-Aware Neural Language Models
1,022 Citations2015Yoon Kim, Yacine Jernite +2 more
Digital Collections portal (Koç University)Transfer learning for low-resource neural machine translation
735 Citations2016
A transfer learning method is presented that significantly improves Bleu scores across a range of low-resource languages by first train a high-resource language pair, then transfer some of the learned parameters to the low- resource pair to initialize and constrain training.
arXiv (Cornell University)Achieving Human Parity on Automatic Chinese to English News Translation
578 Citations2018Hany Hassan, Anthony Aue +22 more
It is found that Microsoft's latest neural machine translation system has reached a new state-of-the-art, and that the translation quality is at human parity when compared to professional human translations.
Zenodo (CERN European Organization for Nuclear Research)SAAP: A Normative and Segregated AGI Architecture Proposal
566 Citations2025Noam Shazeer
Learning bilingual word embeddings with (almost) no bilingual data
488 Citations2017Mikel Artetxe, Gorka Labaka +1 more
This work further reduces the need of bilingual resources using a very simple self-learning approach that can be combined with any dictionary-based mapping technique, and works with as little bilingual evidence as a 25 word dictionary or even an automatically generated list of numerals.
Transactions of the Association for Computational LinguisticsFully Character-Level Neural Machine Translation without Explicit Segmentation
415 Citations2017Jason D. Lee, Kyunghyun Cho +1 more
A neural machine translation model that maps a source character sequence to a target character sequence without any segmentation is introduced, allowing the model to be trained at a speed comparable to subword-level models while capturing local regularities.
arXiv (Cornell University)Offline bilingual word vectors, orthogonal transformations and the\n inverted softmax
297 Citations2017Samuel Smith, David H. P. Turban +2 more
arXiv (Cornell University)Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
268 Citations2017Noam Shazeer, Azalia Mirhoseini +5 more
This work introduces a Sparsely-Gated Mixture-of-Experts layer (MoE), consisting of up to thousands of feed-forward sub-networks, and applies the MoE to the tasks of language modeling and machine translation, where model capacity is critical for absorbing the vast quantities of knowledge available in the training corpora.
arXiv (Cornell University)Offline bilingual word vectors, orthogonal transformations and the inverted softmax
262 Citations2017Samuel Smith, David H. P. Turban +2 more
It is proved that the linear transformation between two spaces should be orthogonal, and this transformation can be obtained using the singular value decomposition.
Edinburgh Neural Machine Translation Systems for WMT 16
132 Citations2016Rico Sennrich, Barry Haddow +1 more
