Efficient Learning and Planning Within the Dyna Framework
Adaptive BehaviorPublished 1 March 1993
Jing Peng, Ronald J. Williams
Citations205
SJR quartileQ1
SJR score0.41
SNIP0.66
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
Sutton's Dyna framework provides a novel and computationally appealing way to integrate learning, planning, and reacting in autonomous agents. Examined here is a class of strategies designed to enhance the learning and planning power of Dyna systems by increasing their computational efficiency. The benefit of using these strategies is demonstrated on some simple abstract learning tasks.
Keywords
Computer Science
Machine LearningQ-learning
8,970 Citations1992Christopher J. Watkins, Peter Dayan
This paper presents and proves in detail a convergence theorem forQ-learning based on that outlined in Watkins (1989), showing that Q-learning converges to the optimum action-values with probability 1 so long as all actions are repeatedly sampled in all states and the action- values are represented discretely.
IBM Journal of Research and DevelopmentSome Studies in Machine Learning Using the Game of Checkers
4,355 Citations1959Arthur L. Samuel
A new signature-table technique is described together with an improved book-learning procedure which is thought to be much superior to the linear polynomial method and to permit the program to look ahead to a much greater depth than it otherwise could do.
Machine LearningLearning to Predict by the Methods of Temporal Differences
3,894 Citations1988Richard S. Sutton
This article introduces a class of incremental learning procedures specialized for prediction-that is, for using past experience with an incompletely known system to predict its future behavior, and proves their convergence and optimality for special cases and relate them to supervised-learning methods.
Principles of Artificial Intelligence
3,336 Citations1982Nils J. Nilsson
This classic introduction to artificial intelligence describes fundamental AI ideas that underlie applications such as natural language processing, automatic programming, robotics, machine vision, automatic theorem proving, and intelligent data retrieval.
Machine LearningLearning to predict by the methods of temporal differences
2,758 Citations1988Richard S. Sutton
Elsevier eBooksIntegrated Architectures for Learning, Planning, and Reacting Based on Approximating Dynamic Programming
1,351 Citations1990Richard S. Sutton
Results are shown for a simple Dyna-PI system that simultaneously learns by trial and error, learns a world model, and plans optimal routes using the evolving world model and it is shown that Dyna-Q architectures are easy to adapt for use in changing environments.
Computers and Thought
921 Citations1963Edward A. Feigenbaum, Julian Feldman
IEEE Control SystemsReinforcement learning is direct adaptive optimal control
513 Citations1992Richard S. Sutton, Andrew G. Barto +1 more
Reinforcement learning methods are presented as a computationally simple, direct approach to the adaptive optimal control of nonlinear systems.
Elsevier eBooksVariable Resolution Dynamic Programming: Efficiently Learning Action Maps in Multivariate Real-valued State-spaces
166 Citations1991Andrew Moore
How such an approach to create an autonomous reactive controller can be realized in real valued multivariate state spaces in which straightforward discretization falls prey to the curse of dimensionality is discussed.
Elsevier eBooksPlanning by Incremental Dynamic Programming
145 Citations1991Richard S. Sutton
The basic results and ideas of dynamic programming as they relate most directly to the concerns of planning in AI are presented, which form the theoretical basis for the incremental planning methods used in the integrated architecture Dyna.
Practical Issues in Temporal Difference Learning
116 Citations1992Gerald Tesauro
Neural Information Processing SystemsMemory-Based Reinforcement Learning: Efficient Computation with Prioritized Sweeping
19 Citations1992Andrew Moore, Christopher G. Atkeson
This work presents a new algorithm, Prioritized Sweeping, for efficient prediction and control of stochastic Markov systems, which successfully solves large state-space real time problems with which other methods have difficulty.
