login

Bandit Based Monte-Carlo Planning

Lecture notes in computer sciencePublished 1 January 2006
Levente Kocsis, Csaba Szepesvári
Citations2,802
SJR quartileQ2
SJR score0.35
SNIP0.55

TL;DR

A new algorithm is introduced, UCT, that applies bandit ideas to guide Monte-Carlo planning and is shown to be consistent and finite sample bounds are derived on the estimation error due to sampling.

Abstract

For large state-space Markovian Decision Problems Monte-Carlo planning is one of the few viable approaches to find near-optimal solutions. In this paper we introduce a new algorithm, UCT, that applies bandit ideas to guide Monte-Carlo planning. In finite-horizon or discounted MDPs the algorithm is shown to be consistent and finite sample bounds are derived on the estimation error due to sampling. Experimental results show that in several domains, UCT is significantly more efficient than its alternatives.

Keywords

Computer ScienceDecision Sciences