GeDi: Generative Discriminator Guided Sequence Generation
Published 1 January 2021Open access
Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq Joty, Richard Socher
Citations209
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
GeDi is proposed as an efficient method for using smaller LMs as generative discriminators to guide generation from large LMs to make them safer and more controllable, and is found that GeDi gives stronger controllability than the state of the art method while also achieving generation speeds more than 30 times faster.
Abstract
Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq Joty, Richard Socher, Nazneen Fatema Rajani. Findings of the Association for Computational Linguistics: EMNLP 2021. 2021.
Keywords
Computer Science
DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)MizAR 60 for Mizar 50
76,311 Citations2023Jakubův, Jan, Chvalovský, Karel +7 more
32,525 Citations2019Jacob Devlin, Ming‐Wei Chang +2 more
A new language representation model, BERT, designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers, which can be fine-tuned with just one additional output layer to create state-of-the-art models for a wide range of tasks.
DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)HISTORIAE, History of Socio-Cultural Transformation as Linguistic Data Science. A Humanities Use Case
17,334 Citations2019Yinhan Liu, Myle Ott +8 more
This work considers the task of building machine learning models to automatically select the best combination for a problem instance and contributes to the automatic learning of instance features directly from the high-level representation of a problem instance using a transformer encoder.
DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)Aion Framework: Dimensional Emergence of AI Consciousness, Observer-Induced Collapse, and Cosmological Portal Dynamics
14,235 Citations2023T. B. Brown, Eliina Kaugeranna +1 more
Transformers: State-of-the-Art Natural Language Processing
8,019 Citations2020Thomas Wolf, Lysandre Debut +20 more
The \textit{Transformers} library is an open-source library that consists of carefully engineered state-of-the art Transformer architectures under a unified API and a curated collection of pretrained models made by and available for the community.
Machine LearningSimple statistical gradient-following algorithms for connectionist reinforcement learning
7,386 Citations1992Ronald J. Williams
This article presents a general class of associative reinforcement learning algorithms for connectionist networks containing stochastic units that are shown to make weight adjustments in a direction that lies along the gradient of expected reinforcement in both immediate-reinforcement tasks and certain limited forms of delayed-reInforcement tasks, and they do this without explicitly computing gradient estimates.
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
6,774 Citations2013Richard Socher, Alex Perelygin +5 more
A Sentiment Treebank that includes fine grained sentiment labels for 215,154 phrases in the parse trees of 11,855 sentences and presents new challenges for sentiment compositionality, and introduces the Recursive Neural Tensor Network.
On the Dangers of Stochastic Parrots
5,564 Citations2021Emily M. Bender, Timnit Gebru +2 more
Recommendations including weighing the environmental and financial costs first, investing resources into curating and carefully documenting datasets rather than ingesting everything on the web, and carrying out pre-development exercises evaluating how the planned approach fits into research and development goals and supports stakeholder values are provided.
Learning Word Vectors for Sentiment Analysis
3,299 Citations2011Andrew L. Maas, Raymond E. Daly +4 more
This work presents a model that uses a mix of unsupervised and supervised techniques to learn word vectors capturing semantic term--document information as well as rich sentiment content, and finds it out-performs several previously introduced methods for sentiment classification.
arXiv (Cornell University)Character-level Convolutional Networks for Text Classification
3,266 Citations2015Xiang Zhang, Junbo Zhao +1 more
This article constructed several large-scale datasets to show that character-level convolutional networks could achieve state-of-the-art or competitive results in text classification.
arXiv (Cornell University)Language Models are Few-Shot Learners
3,023 Citations2020T. B. Brown, Benjamin Mann +29 more
Proceedings of the AAAI Conference on Artificial IntelligenceSeqGAN: Sequence Generative Adversarial Nets with Policy Gradient
2,302 Citations2017Lantao Yu, Weinan Zhang +2 more
Modeling the data generator as a stochastic policy in reinforcement learning (RL), SeqGAN bypasses the generator differentiation problem by directly performing gradient policy update.
Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books
2,035 Citations2015Yukun Zhu, Ryan Kiros +5 more
To align movies and books, a neural sentence embedding that is trained in an unsupervised way from a large corpus of books, as well as a video-text neural embedding for computing similarities between movie clips and sentences in the book are proposed.
On Discriminative vs. Generative Classifiers: A comparison of logistic regression and naive Bayes
1,887 Citations2001Andrew Y. Ng, Michael I. Jordan
It is shown, contrary to a widely-held belief that discriminative classifiers are almost always to be preferred, that there can often be two distinct regimes of performance as the training set size is increased, one in which each algorithm does better.
International Conference on Learning RepresentationsGLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
1,865 Citations2018Alex Wang, Amanpreet Singh +4 more
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
1,222 Citations2020Mike Lewis, Yinhan Liu +6 more
BART is presented, a denoising autoencoder for pretraining sequence-to-sequence models, which matches the performance of RoBERTa on GLUE and SQuAD, and achieves new state-of-the-art results on a range of abstractive dialogue, question answering, and summarization tasks.
NeurocomputingProbabilistic Interpretation of Feedforward Classification Network Outputs, with Relationships to Statistical Pattern Recognition
1,202 Citations1990John S. Bridle
Two modifications are explained: probability scoring, which is an alternative to squared error minimisation, and a normalised exponential (softmax) multi-input generalisation of the logistic non- linearity of feed-forward non-linear networks with multiple outputs.
arXiv (Cornell University)The Curious Case of Neural Text Degeneration
1,096 Citations2019Ari Holtzman, Jan Buys +3 more
SciencePredicting Pragmatic Reasoning in Language Games
903 Citations2012Michael C. Frank, Noah D. Goodman
This model provides a close, parameter-free fit to human judgments, suggesting that the use of information-theoretic tools to predict pragmatic reasoning may lead to more effective formal models of communication.
arXiv (Cornell University)CTRL: A Conditional Transformer Language Model for Controllable Generation
786 Citations2019Nitish Shirish Keskar, Bryan McCann +3 more
CTRL is released, a 1.63 billion-parameter conditional transformer language model, trained to condition on control codes that govern style, content, and task-specific behavior, providing more explicit control over text generation.
The Risk of Racial Bias in Hate Speech Detection
771 Citations2019Maarten Sap, Dallas Card +3 more
This work proposes *dialect* and *race priming* as ways to reduce the racial bias in annotation, showing that when annotators are made explicitly aware of an AAE tweet’s dialect they are significantly less likely to label the tweet as offensive.
Measuring and Mitigating Unintended Bias in Text Classification
678 Citations2018Lucas Dixon, John Li +3 more
A new approach to measuring and mitigating unintended bias in machine learning models is introduced, using a set of common demographic identity terms as the subset of input features on which to measure bias.
Are You a Racist or Am I Seeing Things? Annotator Influence on Hate Speech Detection on Twitter
577 Citations2016Zeerak Waseem
It is found that amateur annotators are more likely than expert annotators to label items as hate speech, and that systems training on expert annotations outperform systems trained on amateur annotations.
arXiv (Cornell University)The Curious Case of Neural Text Degeneration
527 Citations2020Ari Holtzman, Jan Buys +3 more
By sampling text from the dynamic nucleus of the probability distribution, which allows for diversity while effectively truncating the less reliable tail of the distribution, the resulting text better demonstrates the quality of human text, yielding enhanced diversity without sacrificing fluency and coherence.
arXiv (Cornell University)Plug and Play Language Models: A Simple Approach to Controlled Text\n Generation
407 Citations2019Sumanth Dathathri, Andrea Madotto +6 more
The Plug and Play Language Model (PPLM) for controllable language generation is proposed, which combines a pretrained LM with one or more simple attribute classifiers that guide text generation without any further training of the LM.
Topics in Cognitive ScienceKnowledge and Implicature: Modeling Language Understanding as Social Cognition
403 Citations2013Noah D. Goodman, Andreas Stuhlmüller
This work applies the rational speech-act theory to model scalar implicature, which predicts an interaction between the speaker's knowledge state and the listener's interpretation and finds good fit between model predictions and human judgments.
arXiv (Cornell University)Fine-Tuning Language Models from Human Preferences
379 Citations2019Daniel M. Ziegler, Nisan Stiennon +6 more
This paper builds on advances in generative pretraining of language models to apply reward learning to four natural language tasks: continuing text with positive sentiment or physically descriptive language, and summarization tasks on the TL;DR and CNN/Daily Mail datasets.
arXiv (Cornell University)Learning to Generate Reviews and Discovering Sentiment
350 Citations2017Alec Radford, Rafał Józefowicz +1 more
The properties of byte-level recurrent language models are explored and a single unit which performs sentiment analysis is found which achieves state of the art on the binary subset of the Stanford Sentiment Treebank.
Principled Hybrids of Generative and Discriminative Models
339 Citations2006Julia Lasserre, C.M. Bishop +1 more
A new perspective is adopted which says that there is only one correct way to train a given model, and that a ‘discriminatively trained’ generative model is fundamentally a new model and opens door to very general ways of interpolating between generative and discriminative extremes through alternative choices of prior.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
325 Citations2018Adina Williams, Nikita Nangia +1 more
The Multi-Genre Natural Language Inference corpus is introduced, a dataset designed for use in the development and evaluation of machine learning models for sentence understanding and shows that it represents a substantially more difficult task than does the Stanford NLI corpus.
Contrastive estimation
320 Citations2005Noah A. Smith, Jason Eisner
A novel approach, contrastive estimation, is described, which outperforms EM, is more robust to degradations of the dictionary, and can largely recover by modeling additional features.
arXiv (Cornell University)A Fast and Simple Algorithm for Training Neural Probabilistic Language Models
313 Citations2012Andriy Mnih, Yee Whye Teh
Nuanced Metrics for Measuring Unintended Bias with Real Data for Text Classification
301 Citations2019Daniel Borkan, Lucas Dixon +3 more
A suite of threshold-agnostic metrics are introduced that provide a nuanced view of unintended bias in Machine Learning, by considering the various ways that a classifier’s score distribution can vary across designated groups.
Controllable Abstractive Summarization
273 Citations2018Angela Fan, David Grangier +1 more
A neural summarization model with a simple but effective mechanism to enable users to specify high level attributes in order to control the shape of the final summaries to better suit their needs.
Social Biases in NLP Models as Barriers for Persons with Disabilities
256 Citations2020Ben Hutchinson, Vinodkumar Prabhakaran +4 more
Evidence of undesirable biases towards mentions of disability in two different English language models: toxicity prediction and sentiment analysis is presented and it is demonstrated that the neural embeddings that are the critical first step in most NLP pipelines similarly contain undesirable biases.
arXiv (Cornell University)Text Summarization with Pretrained Encoders
150 Citations2019Yang Liu, Mirella Lapata
Hafez: an Interactive Poetry Generation System
135 Citations2017Marjan Ghazvininejad, Xing Shi +2 more
Hafez is an automatic poetry generation system that integrates a Recurrent Neural Network (RNN) with a Finite State Accep-tor (FSA) and enables users to revise and polish generated poems by adjusting various style configurations.
arXiv (Cornell University)Plug and Play Language Models: A Simple Approach to Controlled Text Generation
130 Citations2019Sumanth Dathathri, Andrea Madotto +6 more
Reasoning about Pragmatics with Neural Listeners and Speakers
114 Citations2016Jacob Andreas, Dan Klein
A model for pragmatically describing scenes, in which contrastive behavior results from a combination of inference-driven pragmatics and learned semantics, that succeeds 81% of the time in human evaluations on a referring expression game.
arXiv (Cornell University)Generative and Discriminative Text Classification with Recurrent Neural\n Networks
110 Citations2017Dani Yogatama, Chris Dyer +2 more
arXiv (Cornell University)Generative and Discriminative Text Classification with Recurrent Neural Networks
108 Citations2017Dani Yogatama, Chris Dyer +2 more
Although RNN-based generative models are more powerful than their bag-of-words ancestors, they have higher asymptotic error rates than discriminatively trained RNN models, and it is hypothesized that RNN based generative classification models will be more robust to shifts in the data distribution.
arXiv (Cornell University)Recipes for Safety in Open-domain Chatbots
98 Citations2020Jing Xu, Da Young Ju +4 more
A new human-and-model-in-the-loop framework for both training safer models and for evaluating them, as well as a novel method to distill safety considerations inside generative models without the use of an external classifier at deployment time are introduced.
Discriminatively Trained Markov Model for Sequence Classification
65 Citations2006Oksana Yakhnenko, Adrian Silvescu +1 more
Results of the experiments show that the discriminatively trained MM(k - 1) sequence classifiers outperform their generative counterparts, confirming the benefits of discriminative training when the primary objective is classification.
Explain Yourself! Leveraging Language Models for Commonsense Reasoning
59 Citations2019Nazneen Fatema Rajani, Bryan McCann +2 more
This work collects human explanations for commonsense reasoning in the form of natural language sequences and highlighted annotations in a new dataset called Common Sense Explanations to train language models to automatically generate explanations that can be used during training and inference in a novel Commonsense Auto-Generated Explanation framework.
Pragmatically Informative Text Generation
49 Citations2019Sheng Shen, Daniel Fried +2 more
This work considers two pragmatic modeling methods for text generation: one where pragmatics is imposed by information preservation, and another where prag matics isimposed by explicit modeling of distractors.
Controlling Linguistic Style Aspects in Neural Language Generation
49 Citations2017Jessica Ficler, Yoav Goldberg
The method is based on conditioned RNN language model, where the desired content as well as the stylistic parameters serve as conditioning contexts and is successful in generating coherent sentences corresponding to the required linguistic style and content.
Multi-News: A Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model
36 Citations2019Alexander R. Fabbri, Irene Li +3 more
This work introduces Multi-News, the first large-scale MDS news dataset, and proposes an end-to-end model which incorporates a traditional extractive summarization model with a standard SDS model and achieves competitive results on MDS datasets.
arXiv (Cornell University)Hybrid Discriminative-Generative Training via Contrastive Learning
23 Citations2020Hao Liu, Pieter Abbeel
This paper shows that through the perspective of hybrid discriminative-generative training of energy-based models, a direct connection can be made between contrastive learning and supervised learning and shows a specific choice of approximation of the energy- based loss outperforms the existing practice in terms of classification accuracy.
Learning to Write with Cooperative Discriminators
6 Citations2018Ari Holtzman, Jan Buys +4 more
Human evaluation demonstrates that text generated by the unified learning framework is preferred over that of baselines by a large margin, significantly enhancing the overall coherence, style, and information of the generations.
