Edinburgh Research Explorer (University of Edinburgh)Open access
Jerin Philip, Alexandre Bérard, Matthias Gallé, Laurent Besacier
Citations17
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
We propose a novel adapter layer formalism for adapting multilingual models. They are more parameter-efficient than existing adapter layers while obtaining as good or better performance. The layers are specific to one language (as opposed to bilingual adapters) allowing to compose them and generalize to unseen language-pairs. In this zero-shot setting, they obtain a median improvement of +2.77 BLEU points over a strong 20-language multilingual Transformer baseline trained on TED talks.
Keywords
Computer Science
UvA-DARE (University of Amsterdam)Adam: A Method for Stochastic Optimization
84,783 Citations2014Diederik P. Kingma, Jimmy Ba
DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)MizAR 60 for Mizar 50
76,311 Citations2023Jakubův, Jan, Chvalovský, Karel +7 more
fairseq: A Fast, Extensible Toolkit for Sequence Modeling
2,471 Citations2019Myle Ott, Sergey Edunov +6 more
Fairseq is an open-source sequence modeling toolkit that allows researchers and developers to train custom models for translation, summarization, language modeling, and other text generation tasks and supports distributed training across multiple GPUs and machines.
Transactions of the Association for Computational LinguisticsGoogle’s Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation
1,742 Citations2017Melvin Johnson, Mike Schuster +10 more
This work proposes a simple solution to use a single Neural Machine Translation (NMT) model to translate between multiple languages using a shared wordpiece vocabulary, and introduces an artificial token at the beginning of the input sentence to specify the required target language.
Oxford University Research Archive (ORA) (University of Oxford)Learning multiple visual domains with residual adapters
495 Citations2017Sylvestre-Alvise Rebuffi, Hakan Bilen +1 more
Massively Multilingual Neural Machine Translation
406 Citations2019Roee Aharoni, Melvin Johnson +1 more
It is shown that massively multilingual many-to-many models are effective in low resource settings, outperforming the previous state-of-the-art while supporting up to 59 languages in 116 translation directions in a single model.
When and Why Are Pre-Trained Word Embeddings Useful for Neural Machine Translation?
325 Citations2018Qi Ye, Devendra Singh Sachan +3 more
It is shown that pre-trained word embeddings can be surprisingly effective in NMT tasks – providing gains of up to 20 BLEU points in the most favorable setting.
Simple, Scalable Adaptation for Neural Machine Translation
307 Citations2019Ankur Bapna, Orhan Fırat
The proposed approach consists of injecting tiny task specific adapter layers into a pre-trained model, which adapt the model to multiple individual tasks simultaneously, paving the way towards universal machine translation.
arXiv (Cornell University)Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
296 Citations2019Naveen Arivazhagan, Ankur Bapna +11 more
This work sets a milestone by building a single massively multilingual NMT model handling 103 languages trained on over 25 billion examples, and demonstrates effective transfer learning ability, significantly improving translation quality of low-resource languages, while keeping high-resource language translation quality on-par with competitive bilingual baselines.
Improving Massively Multilingual Neural Machine Translation and Zero-Shot Translation
214 Citations2020Biao Zhang, Philip Williams +2 more
It is argued that multilingual NMT requires stronger modeling capacity to support language pairs with varying typological characteristics, and overcome this bottleneck via language-specific components and deepening NMT architectures.
Rapid Adaptation of Neural Machine Translation to New Languages
192 Citations2018Graham Neubig, Junjie Hu
This paper proposes methods based on starting with massively multilingual “seed models”, which can be trained ahead-of-time, and then continuing training on data related to the LRL, leading to a novel, simple, yet effective method of “similar-language regularization”.
Investigating Multilingual NMT Representations at Scale
101 Citations2019Sneha Kudugunta, Ankur Bapna +2 more
This work attempts to understand massively multilingual NMT representations using Singular Value Canonical Correlation Analysis (SVCCA), a representation similarity framework that allows us to compare representations across different languages, layers and models.
UPCommons institutional repository (Universitat Politècnica de Catalunya)Multilingual machine translation: Closing the gap between shared and language-specific encoder-decoders
32 Citations2021Carlos Escolano, Marta R. Costa‐jussà +2 more
arXiv (Cornell University)A Comprehensive Survey of Multilingual Neural Machine Translation
23 Citations2020Raj Dabre, Chenhui Chu +1 more
An in-depth survey of existing literature on multilingual neural machine translation is presented, which categorizes various approaches based on their central use-case and then further categorize them based on resource scenarios, underlying modeling principles, core-issues and challenges.
UDapter: Language Adaptation for Truly Universal Dependency Parsing
10 Citations2020Ahmet Üstün, Arianna Bisazza +2 more
A novel multilingual task adaptation approach based on recent work in parameter-efficient transfer learning, which allows for an easy but effective integration of existing linguistic typology features into the parsing network, and consistently outperforms strong monolingual and multilingual baselines on both high-resource and low-resource languages.
