Dialogue Learning with Human Teaching and Feedback in End-to-End Trainable Task-Oriented Dialogue Systems
Published 1 January 2018Open access
Bing Liu, Gokhan Tür, Dilek Hakkani‐Tür, Pararth Shah, Larry Heck
Citations151
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
Experimental results show that the end-to-end dialogue agent can learn effectively from the mistake it makes via imitation learning from user teaching, and applying reinforcement learning with user feedback after the imitation learning stage further improves the agent’s capability in successfully completing a task.
Abstract
Bing Liu, Gokhan Tür, Dilek Hakkani-Tür, Pararth Shah, Larry Heck. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Keywords
Computer Science
UvA-DARE (University of Amsterdam)Adam: A Method for Stochastic Optimization
84,783 Citations2014Diederik P. Kingma, Jimmy Ba
DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)
50,318 Citations2021Mandi, Jayanta, Canoy, Rocsildes +2 more
A simple numeric simulation of DNA-co-polymerized hydrogel shape change and a genetic algorithm that generates and selects large batches of material designs that compete with one another to evolve and converge on optimal objective-matching designs are constructed.
Machine LearningSimple statistical gradient-following algorithms for connectionist reinforcement learning
7,386 Citations1992Ronald J. Williams
This article presents a general class of associative reinforcement learning algorithms for connectionist networks containing stochastic units that are shown to make weight adjustments in a direction that lies along the gradient of expected reinforcement in both immediate-reinforcement tasks and certain limited forms of delayed-reInforcement tasks, and they do this without explicitly computing gradient estimates.
Building End-To-End Dialogue Systems Using Generative Hierarchical Neural Network Models
1,725 Citations2016Iulian Vlad Serban, Alessandro Sordoni +3 more
arXiv (Cornell University)A Reduction of Imitation Learning and Structured Prediction to No-Regret\n Online Learning
1,313 Citations2010Stéphane Ross, Geoffrey J. Gordon +1 more
A Persona-Based Neural Conversation Model
897 Citations2016Jiwei Li, Michel Galley +4 more
This work presents persona-based models for handling the issue of speaker consistency in neural response generation that yield qualitative performance improvements in both perplexity and BLEU scores over baseline sequence-to-sequence models.
Proceedings of the IEEEPOMDP-Based Statistical Spoken Dialog Systems: A Review
892 Citations2013Steve Young, Milica Gašić +2 more
This review article provides an overview of the current state of the art in the development of POMDP-based spoken dialog systems.
arXiv (Cornell University)A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning
846 Citations2010Stéphane Ross, Geoffrey J. Gordon +1 more
This paper proposes a new iterative algorithm, which trains a stationary deterministic policy, that can be seen as a no regret algorithm in an online learning setting and demonstrates that this new approach outperforms previous approaches on two challenging imitation learning problems and a benchmark sequence labeling problem.
A Network-based End-to-End Trainable Task-oriented Dialogue System
806 Citations2017Tsung-Hsien Wen, David Vandyke +6 more
This work introduces a neural network-based text-in, text-out end-to-end trainable goal-oriented dialogue system along with a new way of collecting dialogue data based on a novel pipe-lined Wizard-of-Oz framework that can converse with human subjects naturally whilst helping them to accomplish tasks in a restaurant search domain.
The Second Dialog State Tracking Challenge
580 Citations2014Matthew Henderson, Blaise Thomson +1 more
The results suggest that while large improvements on a competitive baseline are possible, trackers are still prone to degradation in mismatched conditions and ensemble learning demonstrates the most accurate tracking can be achieved by combining multiple trackers.
Multi-Domain Joint Semantic Frame Parsing Using Bi-Directional RNN-LSTM
472 Citations2016Dilek Hakkani‐Tür, Gökhan Tür +5 more
Experimental results show the power of a holistic multi-domain, multi-task modeling approach to estimate complete semantic frames for all user utterances addressed to a conversational system over alternative methods based on single domain/task deep learning.
IEEE/ACM Transactions on Audio Speech and Language ProcessingUsing Recurrent Neural Networks for Slot Filling in Spoken Language Understanding
465 Citations2014Grégoire Mesnil, Yann Dauphin +9 more
This paper implemented and compared several important RNN architectures, including Elman, Jordan, and hybrid variants, and implemented these networks with the publicly available Theano neural network toolkit and completed experiments on the well-known airline travel information system (ATIS) benchmark.
FigshareEfficient Reductions for Imitation Learning
403 Citations2018Stéphane Ross, Drew Bagnell
This work proposes two alternative algorithms for imitation learning where training occurs over several episodes of interaction and shows that this leads to stronger performance guarantees and improved performance on two challenging problems: training a learner to play a 3D racing game and Mario Bros.
Word-Based Dialog State Tracking with Recurrent Neural Networks
366 Citations2014Matthew Henderson, Blaise Thomson +1 more
A new wordbased tracking method which maps directly from the speech recognition results to the dialog state without using an explicit semantic decoder is presented, based on a recurrent neural network structure which is capable of generalising to unseen dialog state hypotheses, and which requires very little feature engineering.
Hybrid Code Networks: practical and efficient end-to-end dialog control with supervised and reinforcement learning
341 Citations2017J. D. Williams, Kavosh Asadi +1 more
This work introduces Hybrid Code Networks (HCNs), which combine an RNN with domain-specific knowledge encoded as software and system action templates, and considerably reduce the amount of training data required, while retaining the key benefit of inferring a latent representation of dialog state.
Agenda-based user simulation for bootstrapping a POMDP dialogue system
339 Citations2007Jost Schatzmann, Blaise Thomson +3 more
This paper investigates the problem of bootstrapping a statistical dialogue manager without access to training data and proposes a new probabilistic agenda-based method for simulating user behaviour and shows that the learned policy was highly competitive, with task completion rates above 90%.
The Dialog State Tracking Challenge
331 Citations2013J. D. Williams, Antoine Raux +2 more
The dialog state tracking challenge seeks to address this by providing a heterogeneous corpus of 15K human-computer dialogs in a standard format, along with a suite of 11 evaluation metrics, and shows that the suite of performance metrics cluster into 4 natural groups.
Towards End-to-End Reinforcement Learning of Dialogue Agents for Information Access
300 Citations2017Bhuwan Dhingra, Lihong Li +5 more
This paper proposes KB-InfoBot - a multi-turn dialogue agent which helps users search Knowledge Bases without composing complicated queries by replacing symbolic queries with an induced “soft” posterior distribution over the KB that indicates which entities the user is interested in.
Let's go public! taking a spoken dialog system to the real world
230 Citations2005Antoine Raux, Brian Langner +3 more
The changes necessary to make the Let’s Go Public spoken dialog system usable for the general public are described and analysis of the calls and strategies used to ensure high performance is presented.
Towards End-to-End Learning for Dialog State Tracking and Management using Deep Reinforcement Learning
188 Citations2016Tiancheng Zhao, Maxine Eskénazi
This paper presents an end-to-end framework for task-oriented dialog systems using a variant of Deep Recurrent Q-Networks (DRQN) that is able to interface with a relational database and jointly learn policies for both language understanding and dialog strategy.
Composite Task-Completion Dialogue Policy Learning via Hierarchical Deep Reinforcement Learning
153 Citations2017Baolin Peng, Xiujun Li +5 more
This paper addresses the travel planning task by formulating the task in the mathematical framework of options over Markov Decision Processes (MDPs), and proposing a hierarchical deep reinforcement learning approach to learning a dialogue manager that operates at different temporal scales.
IEEE/ACM Transactions on Audio Speech and Language ProcessingGaussian Processes for POMDP-Based Dialogue Manager Optimization
132 Citations2013Milica Gašić, Steve Young
It is shown that GP policy optimization can be implemented for a real world POMDP dialog manager, and it is demonstrated that designer effort can be substantially reduced by basing the policy directly on the full belief space thereby avoiding ad hoc feature space modeling.
Bootstrapping a Neural Conversational Agent with Dialogue Self-Play, Crowdsourcing and On-Line Reinforcement Learning
127 Citations2018Pararth Shah, Dilek Hakkani‐Tür +2 more
This paper discusses the advantages of this approach for industry applications of conversational agents, wherein an agent can be rapidly bootstrapped to deploy in front of users and further optimized via interactive learning from actual users of the system.
arXiv (Cornell University)End-to-end LSTM-based dialog control optimized with supervised and reinforcement learning
122 Citations2016J. D. Williams, Geoffrey Zweig
The main component of the model is a recurrent neural network (an LSTM), which maps from raw dialog history directly to a distribution over system actions, which relieves the system developer of much of the manual feature engineering of dialog state.
Computational LinguisticsHybrid Reinforcement/Supervised Learning of Dialogue Policies from Fixed Data Sets
121 Citations2008James Henderson, Oliver Lemon +1 more
This work proposes a hybrid model that combines reinforcement learning with supervised learning for dialogue management policies from a fixed data set, which outperforms a pure supervised learning model and a pure reinforcement learning model.
Sample-efficient Actor-Critic Reinforcement Learning with Supervised Data for Dialogue Management
118 Citations2017Pei-Hao Su, Paweł Budzianowski +3 more
A practical approach to learn deep RL-based dialogue policies and demonstrate their effectiveness in a task-oriented information seeking domain is demonstrated.
Gated End-to-End Memory Networks
100 Citations2017Fei Liu, Julien Pérez
A novel end-to-end memory access regulation mechanism inspired by the current progress on the connection short-cutting principle in the field of computer vision is introduced, which is the first of its kind in the world.
On-line Active Reward Learning for Policy Optimisation in Spoken Dialogue Systems
97 Citations2016Pei-Hao Su, Milica Gašić +6 more
An on-line learning framework whereby the dialogue policy is jointly trained alongside the reward model via active learning with a Gaussian process model is proposed.
Joint Online Spoken Language Understanding and Language Modeling With Recurrent Neural Networks
94 Citations2016Bing Liu, Ian Lane
A recurrent neural network model that jointly performs intent detection, slot filling, and language modeling is described that outperforms the independent task training model on SLU tasks and shows advantageous performance in the realistic ASR settings with noisy speech input.
2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)Iterative policy learning in end-to-end trainable task-oriented neural dialog models
94 Citations2017Bing Liu, Ian Lane
A deep reinforcement learning (RL) framework for iterative dialog policy optimization in end-to-end task-oriented dialog systems by jointly optimizing the dialog agent and the user simulator with deep RL by simulating dialogs between the two agents.
An End-to-End Trainable Neural Network Model with Belief Tracking for Task-Oriented Dialog
92 Citations2017Bing Liu, Ian Lane
This work presents a novel end-to-end trainable neural network model that is able to track dialog state, issue API calls to knowledge base (KB), and incorporate structured KB query results into system responses to successfully complete task-oriented dialogs.
arXiv (Cornell University)End-to-End Optimization of Task-Oriented Dialogue Model with Deep Reinforcement Learning
51 Citations2017Bing Liu, Gökhan Tür +3 more
A neural network based task-oriented dialogue system that can be optimized end-to-end with deep reinforcement learning (RL) and shows that deep RL based optimization leads to significant improvement on task success rate and reduction in dialogue length comparing to supervised training model.
Computer Speech & LanguageReinforcement learning for parameter estimation in statistical spoken dialogue systems
48 Citations2011Filip Jurčíček, Blaise Thomson +1 more
Two novel reinforcement algorithms for learning the parameters of a dialogue model are presented, designed to optimise the model parameters while the policy is kept fixed and jointly optimises both the model and the policy parameters.
arXiv (Cornell University)Query-Regression Networks for Machine Comprehension.
11 Citations2016Min Joon Seo, Hannaneh Hajishirzi +1 more
Query-Regression Network is a single recurrent unit with internal memory and local sigmoid attention that is suitable for end-to-end machine comprehension and is able to effectively handle long-term dependencies and is highly parallelizable.
