Discounted Dynamic Programming
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
This chapter discusses the discounted dynamic programming. A policy is any rule for choosing actions. Thus, the action chosen by a policy depends on the history of the process up to that point. An important subclass of the class of all policies is the class of stationary policies. Here, a policy is said to be stationary if it is nonrandomized and the action it chooses at time t depends on the state of the process at t. To determine policies that are optimal, an optimality criterion first needs to be decided. The chapter uses the total expected discounted return as the criterion. The use of a discount factor is economically motivated by the fact that a reward to be earned in the future is less valuable than one earned today. Therefore, a policy is α-optimal if its expected α-discounted return is maximal for every initial state. The chapter also discusses the method of successive approximations.
