mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer
Published 1 January 2021Open access
Linting Xue, Noah Constant, Adam P. Roberts, Mihir Kale, Rami Al‐Rfou, Aditya Siddhant
Citations1,572
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, Colin Raffel. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Keywords
Computer ScienceSocial Sciences
DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)MizAR 60 for Mizar 50
76,311 Citations2023Jakubův, Jan, Chvalovský, Karel +7 more
32,525 Citations2019Jacob Devlin, Ming‐Wei Chang +2 more
A new language representation model, BERT, designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers, which can be fine-tuned with just one additional output layer to create state-of-the-art models for a wide range of tasks.
DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)HISTORIAE, History of Socio-Cultural Transformation as Linguistic Data Science. A Humanities Use Case
17,334 Citations2019Yinhan Liu, Myle Ott +8 more
This work considers the task of building machine learning models to automatically select the best combination for a problem instance and contributes to the automatic learning of instance features directly from the high-level representation of a problem instance using a transformer encoder.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
6,328 Citations2016Pranav Rajpurkar, Jian Zhang +2 more
A strong logistic regression model is built, which achieves an F1 score of 51.0%, a significant improvement over a simple baseline (20%).
arXiv (Cornell University)Cross-lingual Language Model Pretraining
1,616 Citations2019Guillaume Lample, Alexis Conneau
This work proposes two methods to learn cross-lingual language models (XLMs): one unsupervised that only relies on monolingual data, and one supervised that leverages parallel data with a new cross-lingsual language model objective.
XNLI: Evaluating Cross-lingual Sentence Representations
934 Citations2018Alexis Conneau, Ruty Rinott +5 more
This work constructs an evaluation set for XLU by extending the development and test sets of the Multi-Genre Natural Language Inference Corpus to 14 languages, including low-resource languages such as Swahili and Urdu and finds that XNLI represents a practical and challenging evaluation suite and that directly translating the test data yields the best performance among available baselines.
CamemBERT: a Tasty French Language Model
710 Citations2020Louis Martin, Benjamin Müller +6 more
This paper investigates the feasibility of training monolingual Transformer-based language models for other languages, taking French as an example and evaluating their language models on part-of-speech tagging, dependency parsing, named entity recognition and natural language inference tasks.
arXiv (Cornell University)Multilingual Denoising Pre-training for Neural Machine Translation
607 Citations2020Yinhan Liu, Jiatao Gu +6 more
Transfer Learning in Natural Language Processing
579 Citations2019Sebastian Ruder, Matthew E. Peters +2 more
An overview of modern transfer learning methods in NLP, how models are pre-trained, what information the representations they learn capture, and review examples and case studies on how these models can be integrated and adapted in downstream NLP tasks are presented.
Document Ranking with a Pretrained Sequence-to-Sequence Model
414 Citations2020Rodrigo Nogueira, Zhiying Jiang +2 more
Surprisingly, it is found that the choice of target tokens impacts effectiveness, even for words that are closely related semantically, which sheds some light on why the sequence-to-sequence formulation for document ranking is effective.
arXiv (Cornell University)The Natural Language Decathlon: Multitask Learning as Question Answering
338 Citations2018Bryan McCann, Nitish Shirish Keskar +2 more
Presented on August 28, 2018 at 12:15 p.m. in the Pettit Microelectronics Research Center, Room 102 A/B.
Transactions of the Association for Computational LinguisticsT<scp>y</scp>D<scp>i</scp> QA: A Benchmark for Information-Seeking Question Answering in <i>Ty</i>pologically <i>Di</i>verse Languages
336 Citations2020Jonathan H. Clark, Eunsol Choi +5 more
A quantitative analysis of the data quality and example-level qualitative linguistic analyses of observed language phenomena that would not be found in English-only corpora are presented.
Cross-lingual Name Tagging and Linking for 282 Languages
301 Citations2017Xiaoman Pan, Boliang Zhang +4 more
This work develops a cross-lingual name tagging and linking framework for 282 languages that exist in Wikipedia that is able to identify name mentions, assign a coarse-grained or fine- grained type to each mention, and link it to an English Knowledge Base (KB) if it is linkable.
arXiv (Cornell University)XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization
299 Citations2020Junjie Hu, Sebastian Ruder +4 more
The Cross-lingual TRansfer Evaluation of Multilingual Encoders XTREME benchmark is introduced, a multi-task benchmark for evaluating the cross-lingually generalization capabilities of multilingual representations across 40 languages and 9 tasks.
arXiv (Cornell University)Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
296 Citations2019Naveen Arivazhagan, Ankur Bapna +11 more
This work sets a milestone by building a single massively multilingual NMT model handling 103 languages trained on over 25 billion examples, and demonstrates effective transfer learning ability, significantly improving translation quality of low-resource languages, while keeping high-resource language translation quality on-par with competitive bilingual baselines.
arXiv (Cornell University)CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data
246 Citations2019Guillaume Wenzek, Marie-Anne Lachaux +5 more
PAWS-X: A Cross-lingual Adversarial Dataset for Paraphrase Identification
228 Citations2019Yinfei Yang, Yuan Zhang +2 more
PAWS-X, a new dataset of 23,659 human translated PAWS evaluation pairs in six typologically distinct languages, shows the effectiveness of deep, multilingual pre-training while also leaving considerable headroom as a new challenge to drive multilingual research that better captures structure and contextual information.
arXiv (Cornell University)BERTje: A Dutch BERT Model
214 Citations2019Wietse de Vries, Andreas van Cranenburgh +4 more
The transformer-based pre-trained language model BERT has helped to improve state-of-the-art performance on many natural language processing (NLP) tasks, but a monolingual Dutch BERT model called BERTje is developed and evaluated, which consistently outperforms the equally-sized multilingual Bert model on downstream NLP tasks.
arXiv (Cornell University)WT5?! Training Text-to-Text Models to Explain their Predictions
103 Citations2020Sharan Narang, Colin Raffel +4 more
This paper uses the text-to-text framework proposed by Raffel et al. (2019) to train language models to output a natural text explanation alongside their prediction, and shows that this approach not only obtains state-of-the-art results on explainability benchmarks, but also permits learning from a limited set of labeled explanations and transferring rationalization abilities across datasets.
Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering
99 Citations2021Gautier Izacard, Édouard Grave
Interestingly, it is observed that the performance of this method significantly improves when increasing the number of retrieved passages, evidence that sequence-to-sequence models offers a flexible framework to efficiently aggregate and combine evidence from multiple passages.
arXiv (Cornell University)Pre-training via Paraphrasing
89 Citations2020Mike Lewis, Marjan Ghazvininejad +4 more
It is shown that fine-tuning gives strong performance on a range of discriminative and generative tasks in many languages, making MARGE the most generally applicable pre-training method to date.
arXiv (Cornell University)InfoXLM: An Information-Theoretic Framework for Cross-Lingual Language Model Pre-Training
77 Citations2020Zewen Chi, Li Dong +8 more
arXiv (Cornell University)Pre-training via Paraphrasing
56 Citations2020Mike Lewis, Marjan Ghazvininejad +4 more
Proceedings of the AAAI Conference on Artificial IntelligenceFILTER: An Enhanced Fusion Method for Cross-lingual Language Understanding
55 Citations2021Yuwei Fang, Shuohang Wang +3 more
FILTER is proposed, an enhanced fusion method that takes cross-lingual data as input for XLM finetuning and proposes an additional KL-divergence self-teaching loss for model training, based on auto-generated soft pseudo-labels for translated text in the target language.
arXiv (Cornell University)Playing with Words at the National Library of Sweden -- Making a Swedish\n BERT
50 Citations2020Martin Malmsten, Love Börjeson +1 more
arXiv (Cornell University)PTT5: Pretraining and validating the T5 model on Brazilian Portuguese\n data
32 Citations2020Diedre Carmo, Marcos Piau +3 more
RobBERT: a Dutch RoBERTa-based Language Model
31 Citations2020Pieter Delobelle, Thomas Winters +1 more
It is found that RobBERT improves state-of-the-art results for various tasks, and especially significantly outperforms other models when dealing with smaller datasets, indicating that it is a powerful pre-trained model for a large variety of Dutch language tasks.
arXiv (Cornell University)Playing with Words at the National Library of Sweden -- Making a Swedish BERT
31 Citations2020Martin Malmsten, Love Börjeson +1 more
The Swedish BERT ("KB-BERT") developed by the KBLab for data-driven research at the National Library of Sweden is introduced and it is demonstrated that KB-berT outperforms existing models in a range of NLP tasks from named entity recognition (NER) to part-of-speech tagging (POS).
arXiv (Cornell University)VECO: Variable Encoder-decoder Pre-training for Cross-lingual Understanding and Generation
31 Citations2021Fuli Luo, Wei Wang +6 more
A variable encoder-decoder (VECO) pre-training approach to unify the two mainstreams in both model architectures and pre- training tasks, which delivers new state-of-the-art results on various cross-lingual understanding tasks of the XTREME benchmark.
arXiv (Cornell University)GLU Variants Improve Transformer
20 Citations2020Noam Shazeer
Gated Linear Units (GLU) consist of the component-wise product of two linear projections, one of which is first passed through a sigmoid function, and it is found that some of them yield quality improvements over the typically-used ReLU or GELU activations.
arXiv (Cornell University)Exploring Fine-tuning Techniques for Pre-trained Cross-lingual Models via Continual Learning
17 Citations2020Zihan Liu, Genta Indra Winata +2 more
The method achieves better performance than other fine-tuning baselines on zero-shot cross-lingual part-of-speech tagging and named entity recognition tasks and preserves the original cross- Lingual ability of the pre-trained model when the authors fine-tune it to downstream cross-lingsual tasks.
arXiv (Cornell University)PTT5: Pretraining and validating the T5 model on Brazilian Portuguese data
10 Citations2020Diedre Carmo, Marcos Piau +3 more
This paper pretrain a T5 model on the BrWac corpus, an extensive collection of web pages in Portuguese, and evaluates its performance against other Portuguese pretrained models and multilingual models on the sentence similarity and sentence entailment tasks.
