Overcoming Catastrophic Forgetting During Domain Adaptation of Neural Machine Translation
Published 1 January 2019Open access
Brian J. Thompson, Jeremy Gwinnup, Huda Khayrallah, Kevin Duh, Philipp Koehn
Citations125
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This work adapts Elastic Weight Consolidation (EWC)—a machine learning method for learning a new task without forgetting previous tasks—to mitigate the drop in general-domain performance as catastrophic forgetting of general- domain knowledge.
Abstract
Brian Thompson, Jeremy Gwinnup, Huda Khayrallah, Kevin Duh, Philipp Koehn. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 2019.
Keywords
Computer Science
UvA-DARE (University of Amsterdam)Adam: A Method for Stochastic Optimization
84,783 Citations2014Diederik P. Kingma, Jimmy Ba
DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)
50,318 Citations2021Mandi, Jayanta, Canoy, Rocsildes +2 more
A simple numeric simulation of DNA-co-polymerized hydrogel shape change and a genetic algorithm that generates and selects large batches of material designs that compete with one another to evolve and converge on optimal objective-matching designs are constructed.
arXiv (Cornell University)Distilling the Knowledge in a Neural Network
13,965 Citations2015Geoffrey E. Hinton, Oriol Vinyals +1 more
This work shows that it can significantly improve the acoustic model of a heavily used commercial system by distilling the knowledge in an ensemble of models into a single model and introduces a new type of ensemble composed of one or more full models and many specialist models which learn to distinguish fine-grained classes that the full models confuse.
Proceedings of the National Academy of SciencesOvercoming catastrophic forgetting in neural networks
7,158 Citations2017James Kirkpatrick, Razvan Pascanu +12 more
It is shown that it is possible to overcome the limitation of connectionist models and train networks that can maintain expertise on tasks that they have not experienced for a long time and selectively slowing down learning on the weights important for previous tasks.
Moses
4,868 Citations2007Philipp Koehn, Richard Zens +12 more
An open-source toolkit for statistical machine translation whose novel contributions are support for linguistically motivated factors, confusion network decoding, and efficient data formats for translation models and language models.
Neural ComputationA Practical Bayesian Framework for Backpropagation Networks
2,934 Citations1992David Mackay
A quantitative and practical Bayesian framework is described for learning of mappings in feedforward networks that automatically embodies "Occam's razor," penalizing overflexible and overcomplex models.
RWTH Publications (RWTH Aachen)Moses: Open Source Toolkit for Statistical Machine Translation
1,451 Citations2007Philipp Koehn, Hieu Hoang +12 more
Sequence-Level Knowledge Distillation
775 Citations2016Yoon Kim, Alexander M. Rush
It is demonstrated that standard knowledge distillation applied to word-level prediction can be effective for NMT, and two novel sequence-level versions of knowledge distilling are introduced that further improve performance, and somewhat surprisingly, seem to eliminate the need for beam search.
arXiv (Cornell University)An Empirical Investigation of Catastrophic Forgetting in Gradient-Based Neural Networks
496 Citations2014Ian Goodfellow, Mehdi Mirza +3 more
It is found that it is always best to train using the dropout algorithm--the drop out algorithm is consistently best at adapting to the new task, remembering the old task, and has the best tradeoff curve between these two extremes.
arXiv (Cornell University)An Empirical Investigation of Catastrophic Forgetting in Gradient-Based Neural Networks
492 Citations2013Ian Goodfellow, Mehdi Mirza +3 more
Findings of the 2017 Conference on Machine Translation (WMT17)
418 Citations2017Ondřej Bojar, Rajen Chatterjee +14 more
The results of the WMT17 shared tasks, which included three machine translation (MT) tasks (news, biomedical, and multimodal), two evaluation tasks (metrics and run-time estimation of MT quality), an automatic post-editing task, a neural MT training task, and a bandit learning task are presented.
Journal of Machine Learning ResearchNew Insights and Perspectives on the Natural Gradient Method
256 Citations2020James Martens
This paper critically analyze this method and its properties, and shows how it can be viewed as a type of approximate 2nd-order optimization method, where the Fisher information matrix can be view as an approximation of the Hessian.
An Empirical Comparison of Domain Adaptation Methods for Neural Machine Translation
229 Citations2017Chenhui Chu, Raj Dabre +1 more
A novel domain adaptation method named “mixed fine tuning” for neural machine translation (NMT) is proposed which combines two existing approaches namely fine tuning and multi domain NMT.
arXiv (Cornell University)Sockeye: A Toolkit for Neural Machine Translation
194 Citations2017Felix Hieber, Tobias Domhan +5 more
This paper highlights Sockeye's features and benchmark it against other NMT toolkits on two language arcs from the 2017 Conference on Machine Translation (WMT): English-German and Latvian-English, and reports competitive BLEU scores across all three architectures.
arXiv (Cornell University)Fast Domain Adaptation for Neural Machine Translation
164 Citations2016Markus Freitag, Yaser Al-Onaizan
This paper proposes an approach for adapting a NMT system to a new domain with the main idea behind domain adaptation that the availability of large out-of-domain training data and a small in- domain training data.
Revisiting Natural Gradient for Deep Networks
139 Citations2014Razvan Pascanu, Yoshua Bengio
It is described how one can use unlabeled data to improve the generalization error obtained by natural gradient and empirically evaluate the robustness of the algorithm to the ordering of the training set compared to stochastic gradient descent.
PerspectivesResistance and accommodation: factors for the (non-) adoption of machine translation among professional translators
137 Citations2017Patrick Cadwell, Sharon O’Brien +1 more
arXiv (Cornell University)Revisiting Natural Gradient for Deep Networks
122 Citations2013Razvan Pascanu, Yoshua Bengio
arXiv (Cornell University)New insights and perspectives on the natural gradient method
94 Citations2014James Martens
Regularized Training Objective for Continued Training for Domain Adaptation in Neural Machine Translation
78 Citations2018Huda Khayrallah, Brian J. Thompson +2 more
This work adds an auxiliary term to the training objective during continued training that minimizes the cross entropy between the in-domain model’s output word distribution and that of the out-of- domain model to prevent the model‘s output from differing too much from the original out- of- domains model.
Freezing Subnetworks to Analyze Domain Adaptation in Neural Machine Translation
72 Citations2018Brian J. Thompson, Huda Khayrallah +8 more
It is found that freezing any single component during continued training has minimal impact on performance, and that performance is surprisingly good when a single component is adapted while holding the rest of the model fixed.
Learning Hidden Unit Contribution for Adapting Neural Machine Translation Models
35 Citations2018David Vilar
This paper explores the use of Learning Hidden Unit Contribution for the task of neural machine translation and shows that the proposed method achieves improvements of up to 2.6 BLEU points over a general system and up to 6 BLEu points if the initial system has only been trained on out-of-domain data.
COPPA V2. 0: Corpus Of Parallel Patent Applications Building Large Parallel Corpora with GNU Make
13 Citations2016Marcin Junczys-Dowmunt, Bruno Pouliquen +1 more
With the goal of supporting innovation in the Machine Translation field, WIPO offers the updated corpus under the same conditions as before, the product being notably free of charge for academic and private research institutions for research purposes only (without redistribution right).
Study on the use of machine translation and post-editing in Swiss-based language service providers
3 Citations2017Victoria Porro Rodriguez, Lucía Morado Vázquez +1 more
A survey among Swiss language service providers revealed that only 2 out of 16 companies were using a machine translation system combined with human post-editing in their translation workflow.
