UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning
Published 1 January 2021Open access
Wei Li, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu, Jiachen Liu
Citations234
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A UNIfied-MOdal pre-training architecture, namely UNIMO, which can effectively adapt to both single- modal and multi-modal understanding and generation tasks, and is able to learn more generalizable representations.
Abstract
Wei Li, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu, Jiachen Liu, Hua Wu, Haifeng Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Keywords
Computer Science
Deep Residual Learning for Image Recognition
222,082 Citations2016Kaiming He, Xiangyu Zhang +2 more
This work presents a residual learning framework to ease the training of networks that are substantially deeper than those used previously, and provides comprehensive empirical evidence showing that these residual networks are easier to optimize, and can gain accuracy from considerably increased depth.
arXiv (Cornell University)Very Deep Convolutional Networks for Large-Scale Image Recognition
75,540 Citations2014Karen Simonyan, Andrew Zisserman
This work investigates the effect of the convolutional network depth on its accuracy in the large-scale image recognition setting using an architecture with very small convolution filters, which shows that a significant improvement on the prior-art configurations can be achieved by pushing the depth to 16-19 weight layers.
IEEE Transactions on Pattern Analysis and Machine IntelligenceFaster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
54,186 Citations2016Shaoqing Ren, Kaiming He +2 more
This work introduces a Region Proposal Network (RPN) that shares full-image convolutional features with the detection network, thus enabling nearly cost-free region proposals and further merge RPN and Fast R-CNN into a single network by sharing their convolutionAL features.
Lecture notes in computer scienceMicrosoft COCO: Common Objects in Context
42,219 Citations2014Tsung-Yi Lin, Michael Maire +6 more
A new dataset with the goal of advancing the state-of-the-art in object recognition by placing the question of object recognition in the context of the broader question of scene understanding by gathering images of complex everyday scenes containing common objects in their natural context.
32,525 Citations2019Jacob Devlin, Ming‐Wei Chang +2 more
A new language representation model, BERT, designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers, which can be fine-tuned with just one additional output layer to create state-of-the-art models for a wide range of tasks.
DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)HISTORIAE, History of Socio-Cultural Transformation as Linguistic Data Science. A Humanities Use Case
17,334 Citations2019Yinhan Liu, Myle Ott +8 more
This work considers the task of building machine learning models to automatically select the best combination for a problem instance and contributes to the automatic learning of instance features directly from the high-level representation of a problem instance using a transformer encoder.
Momentum Contrast for Unsupervised Visual Representation Learning
12,039 Citations2020Kaiming He, Haoqi Fan +3 more
arXiv (Cornell University)A Simple Framework for Contrastive Learning of Visual Representations
7,335 Citations2020Ting Chen, Simon Kornblith +2 more
It is shown that composition of data augmentations plays a critical role in defining effective predictive tasks, and introducing a learnable nonlinear transformation between the representation and the contrastive loss substantially improves the quality of the learned representations, and contrastive learning benefits from larger batch sizes and more training steps compared to supervised learning.
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
6,774 Citations2013Richard Socher, Alex Perelygin +5 more
A Sentiment Treebank that includes fine grained sentiment labels for 215,154 phrases in the parse trees of 11,855 sentences and presents new challenges for sentiment compositionality, and introduces the Recursive Neural Tensor Network.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
6,328 Citations2016Pranav Rajpurkar, Jian Zhang +2 more
A strong logistic regression model is built, which achieves an F1 score of 51.0%, a significant improvement over a simple baseline (20%).
arXiv (Cornell University)Learning Transferable Visual Models From Natural Language Supervision
5,296 Citations2021Alec Radford, Jong Wook Kim +10 more
International Journal of Computer VisionVisual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
5,051 Citations2017Ranjay Krishna, Yuke Zhu +10 more
The Visual Genome dataset is presented, which contains over 108K images where each image has an average of 35 objects and contains dense annotations of objects, attributes, and relationships within each image to learn these models.
Transactions of the Association for Computational LinguisticsFrom image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
2,461 Citations2014Peter Young, Alice Lai +2 more
This work proposes to use the visual denotations of linguistic expressions to define novel denotational similarity metrics, which are shown to be at least as beneficial as distributional similarities for two tasks that require semantic inference.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
2,084 Citations2017Yash Goyal, Tejas Khot +3 more
Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books
2,035 Citations2015Yukun Zhu, Ryan Kiros +5 more
To align movies and books, a neural sentence embedding that is trained in an unsupervised way from a large corpus of books, as well as a video-text neural embedding for computing similarities between movie clips and sentences in the book are proposed.
Oxford University Research Archive (ORA) (University of Oxford)Teaching Machines to Read and Comprehend
1,936 Citations2015Karl Moritz Hermann, Tomáš Kočiský +5 more
Lecture notes in computer scienceUNITER: UNiversal Image-TExt Representation Learning
1,847 Citations2020Yen-Chun Chen, Linjie Li +6 more
UNITER, a UNiversal Image-TExt Representation, learned through large-scale pre-training over four image-text datasets is introduced, which can power heterogeneous downstream V+L tasks with joint multimodal embeddings.
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning
1,806 Citations2018Piyush Sharma, Nan Ding +2 more
arXiv (Cornell University)ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
1,673 Citations2019Jiasen Lu, Dhruv Batra +2 more
arXiv (Cornell University)Microsoft COCO Captions: Data Collection and Evaluation Server
1,627 Citations2015Xinlei Chen, Hao Fang +5 more
The Microsoft COCO Caption dataset and evaluation server are described and several popular metrics, including BLEU, METEOR, ROUGE and CIDEr are used to score candidate captions.
Lecture notes in computer scienceOscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
1,484 Citations2020Xiujun Li, Xi Yin +10 more
This paper proposes a new learning method Oscar (Object-Semantics Aligned Pre-training), which uses object tags detected in images as anchor points to significantly ease the learning of alignments.
arXiv (Cornell University)VisualBERT: A Simple and Performant Baseline for Vision and Language
1,229 Citations2019Liunian Harold Li, Mark Yatskar +3 more
Analysis demonstrates that VisualBERT can ground elements of language to image regions without any explicit supervision and is even sensitive to syntactic relationships, tracking, for example, associations between verbs and image regions corresponding to their arguments.
TIB Data ManagerA simple framework for contrastive learning of visual representations
1,203 Citations2024Ting Chen
arXiv (Cornell University)Unified language model pre-training for natural language understanding and generation
949 Citations2024Li Dong
A new Unified pre-trained Language Model (UniLM) that can be fine-tuned for both natural language understanding and generation tasks that compares favorably with BERT on the GLUE benchmark, and the SQuAD 2.0 and CoQA question answering tasks.
Im2Text: Describing Images Using 1 Million Captioned Photographs
944 Citations2011Vicente Ordóñez, Girish Kulkarni +1 more
A new objective performance measure for image captioning is introduced and methods incorporating many state of the art, but fairly noisy, estimates of image content are developed to produce even more pleasing results.
Neural Information Processing SystemsViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
906 Citations2019Jiasen Lu, Dhruv Batra +2 more
ViLBERT (short for Vision-and-Language BERT), a model for learning task-agnostic joint representations of image content and natural language, is presented, extending the popular BERT architecture to a multi-modal two-stream model, pro-cessing both visual and textual inputs in separate streams that interact through co-attentional transformer layers.
Proceedings of the AAAI Conference on Artificial IntelligenceUnified Vision-Language Pre-Training for Image Captioning and VQA
838 Citations2020Luowei Zhou, Hamid Palangi +4 more
VLP is the first reported model that achieves state-of-the-art results on both vision-language generation and understanding tasks, as disparate as image captioning and visual question answering, across three challenging benchmark datasets: COCO Captions, Flickr30k Captions and VQA 2.0.
arXiv (Cornell University)VL-BERT: Pre-training of Generic Visual-Linguistic Representations
782 Citations2019Weijie Su, Xizhou Zhu +5 more
A new pre-trainable generic representation for visual-linguistic tasks, called Visual-Linguistic BERT (VL-BERT), which adopts the simple yet powerful Transformer model as the backbone, and extends it to take both visual and linguistic embedded features as input.
Proceedings of the AAAI Conference on Artificial IntelligenceUnicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-Training
745 Citations2020Gen Li, Nan Duan +3 more
After pretraining on large-scale image-caption pairs, Unicoder-VL is transferred to caption-based image-text retrieval and visual commonsense reasoning, with just one additional output layer, and shows the powerful ability of the cross-modal pre-training.
arXiv (Cornell University)Unified Language Model Pre-training for Natural Language Understanding\n and Generation
539 Citations2019Li Dong, Nan Yang +7 more
arXiv (Cornell University)Large-Scale Adversarial Training for Vision-and-Language Representation Learning
287 Citations2020Zhe Gan, Yen-Chun Chen +4 more
To enable large-scale training, VILLA adopts the "free" adversarial training strategy, and combines it with KL-divergence-based regularization to promote higher invariance in the embedding space.
Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning
211 Citations2021Zhicheng Huang, Zhaoyang Zeng +4 more
This paper proposes SOHO to "Seeing Out of tHe bOx" that takes a whole image as input, and learns vision-language representation in an end-to-end manner, and does not require bounding box annotations which enables inference 10 times faster than region-based approaches.
arXiv (Cornell University)Visual Entailment: A Novel Task for Fine-Grained Image Understanding
162 Citations2019Ning Xie, Farley Lai +2 more
A new inference task, Visual Entailed (VE) - consisting of image-sentence pairs whereby a premise is defined by an image, rather than a natural language sentence as in traditional Textual Entailment tasks is introduced.
arXiv (Cornell University)ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph
118 Citations2020F. Richard Yu, Jiji Tang +5 more
ERNIE-GEN: An Enhanced Multi-Flow Pre-training and Fine-tuning Framework for Natural Language Generation
107 Citations2020Dongling Xiao, Han Zhang +5 more
An enhanced multi-flow sequence to sequence pre-training and fine-tuning framework named ERNIE-GEN, which bridges the discrepancy between training and inference with an infilling generation mechanism and a noise-aware generation method to make generation closer to human writing patterns.
arXiv (Cornell University)WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training
85 Citations2021Yuqi Huo, Manli Zhang +33 more
arXiv (Cornell University)InterBERT: Vision-and-Language Interaction for Multi-modal Pretraining
56 Citations2020Junyang Lin, Yang An +4 more
A novel model, namely InterBERT (BERT for Interaction), which owns strong capability of modeling interaction between the information flows of different modalities, is proposed and developed, which is the first Chinese multi-modal pretrained model.
Scene Graph Parsing as Dependency Parsing
54 Citations2018Yu-Siang Wang, Chenxi Liu +2 more
This paper introduces an alternative but equivalent edge-centric view of scene graphs that connect to dependency parses, and combines the two-stage pipeline used in prior work (generic dependency parsing followed by simple post-processing) into one, enabling end-to-end training.
eLifeNeuronal populations in the occipital cortex of the blind synchronize to the temporal dynamics of speech
50 Citations2018Markus J. van Ackeren, Francesca M. Barbero +3 more
These findings suggest that the occipital cortex of the blind adopts an architecture allowing the tracking of speech material, and therefore does not fully abstract from the reorganized sensory inputs it receives.
Information SciencesMulti-modal neural machine translation with deep semantic interactions
40 Citations2020Jinsong Su, Jinchang Chen +6 more
This model extends the conventional multi-modal NMT by introducing the following two attention neural networks: a bi-directional attention network for modeling text and image representations, where the semantic representations of text are learned by referring to the image representation, and vice versa.
arXiv (Cornell University)ERNIE-GEN: An Enhanced Multi-Flow Pre-training and Fine-tuning Framework for Natural Language Generation
28 Citations2020Dongling Xiao, Han Zhang +5 more
CoLA: The Corpus of Linguistic Acceptability (with added annotations)
8 Citations2019Alex Warstadt, Amanpreet Singh +1 more
bioRxiv (Cold Spring Harbor Laboratory)Neuronal populations in the occipital cortex of the blind synchronize to the temporal dynamics of speech
3 Citations2017Markus J. van Ackeren, Francesca M. Barbero +3 more
