login

Asynchronous stochastic approximation and Q-learning

Machine LearningPublished 1 September 1994Open access
John N. Tsitsiklis
Citations455
SJR quartileQ1
SJR score1.15
SNIP2.14
View PDF

TL;DR

The Q-learning algorithm, a reinforcement learning method for solving Markov decision problems, is studied to establish its convergence under conditions more general than previously available.

Abstract

We provide some general results on the convergence of a class of stochastic approximation algorithms and their parallel and asynchronous variants. We then use these results to study the Q-learning algorithm, a reinforcement learning method for solving Markov decision problems, and establish its convergence under conditions more general than previously available.

Keywords

Computer ScienceDecision Sciences