> For the complete documentation index, see [llms.txt](https://theshank.gitbook.io/ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://theshank.gitbook.io/ai/reinforcement-learning/deep-reinforcement-learning/value-function-approximation.md).

# Value Function Approximation

## Value Function Approximation

SO far we have represented value function using look-up table:\
\- Every state s has an entry V(s)\
\- Every state-action pair have an entry Q(s,a)

Now it is impossible to have V(s) entry if we have large number of states. Hence, to solve this we use **Value Function Approximator.** Which are basically functoins which can map state to its value i.e $$v: s \rightarrow a$$ .

## Stochastic Gradient Descent

$$
J(w) = E\_\pi\[(v\_\pi(s)-\hat{v}(s,w))^2]
$$

Here, $$v\_\pi$$ is the original taget and $$\hat{v}$$ is the value approximation function with parameters w. Now we will update the $$w$$ using gradient descent in order to minimize the the cost $$J(w)$$ . Hence

$$
\triangle w = \alpha (v\_\pi(s)-\hat{v}(s,w))\triangledown\_w\hat{v}
(s,w)
$$

But is the problem that we don't know the target value funtion $$v\_\pi(s)$$ in reinforcement learning before hand . We will see below how to handle this problem.

## Different Value approximation

### Linear Function Approximation

**Feature Vector:** We represent state using a vector as below.&#x20;

![](https://1877261540-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LFDuA0A2VRqmT31Blrq%2F-Lbn9KEfr19fEnrFORhv%2F-LbrEC6lu04kzws1wEay%2Fimage.png?alt=media\&token=dc830bf8-03d2-4268-a0ef-23f4577b1f38)

Using this state feature vector and weights. Linear approximation is as follows:

$$
\hat{v}(s,w) = x(s)^Tw = \sum x\_j(s)w\_j
$$

Hence in this case update become:

$$
\triangle w = \alpha (v\_\pi(s)-\hat{v}(s,w))x(s)
$$

**Table Lookup Features:** It is special case for linear function approximation. Here the state vector have all the states in itself. see below:<br>

![](https://1877261540-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LFDuA0A2VRqmT31Blrq%2F-Lbn9KEfr19fEnrFORhv%2F-LbrFOvp0wRTTiOAXd41%2Fimage.png?alt=media\&token=44d7e32f-2a86-461c-8551-c613b46ab16c)

## Incremental Prediction Algorithms

Here we talk about the ways in which we can substitute the value of $$v\_\pi(s)$$ from the the weights.&#x20;

### Monte-Carlo with Value Funtoni Approx.

Use $$G\_t$$ calucuated from the episodes here in place of $$v\_\pi(s)$$. We will have \
&#x20;$$<(S\_1, G\_1), (S\_2, G\_2),...(S\_T, G\_T)>$$ , which will be ou training data.&#x20;

Hence using linear Monte-carlo policy evaluation, algorithm will be:

$$
\triangle w = \alpha (G\_t-\hat{v}(s\_t,w))x(s\_t)
$$

**Monte-Carlo evaluation will converge to a local optimum.  Even with non-linear value funtion approximation.**&#x20;

### **TD Learning with value funtion approx.**

The TD-target $$R\_{t+1} + γ v̂ (S\_{t+1} , w)$$ is a biased sample of true value $$v\_\pi(s\_t)$$**.**&#x20;

Supervised learning can be applies using traingin data: $$<(S\_1, R\_2 + γ v̂ (S\_2 , w)), (S\_2, R\_3 + γ v̂ (S\_3 , w)),...(S\_{T-1}, R\_T)>$$

$$
\triangle w = \alpha (R\_{t+1} + γ v̂ (S\_{t+1} , w)-\hat{v}(s\_t,w))x(s\_t)
$$

Linear TD(0) coverges(close) to global optimum.&#x20;

### TD() with value funtion approximation.&#x20;

The λ-return $$G\_t^λ$$ is also a biased sample of true value $$v\_\pi(s\_t)$$

Apply Supervised learning with following data: $$<(S\_1, G\_1^\lambda), (S\_2, G\_2^\lambda),...(S\_T, G\_T^\lambda)>$$

Forward View linear TD():

$$
\triangle w = \alpha (G\_t^\lambda-\hat{v}(s\_t,w))x(s\_t)
$$

Backward view linear TD():

$$
\delta\_t = R\_{t+1} +  γ v̂ (S\_{t+1} , w)-\hat{v}(s\_t,w)\\
E\_t = \gamma \lambda E\_{t-1} + x(S\_t)\\
\triangle w = \alpha \delta\_t E\_t
$$

Forward and backward view llinear TD() are equivalent.&#x20;

## Incremental Control Algorithms

Policy Evalutaion: Approximate policy evaluation: $$\hat{q}(.,.,w)  \approx q\_\pi$$ \
Policy Improvement : $$\epsilon-greedy$$ policy improvement

Here

$$
J(w) = E\_\pi\[(q\_\pi(S,A,w)-\hat{q}(S,A,w))^2]
$$

$$
\triangle w = \alpha (q\_\pi(S,A,w)-\hat{q}(S,A,w))\triangle\_w \hat{q}(S,A,w)
$$

### Linear approximation:

![](https://1877261540-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LFDuA0A2VRqmT31Blrq%2F-Lbn9KEfr19fEnrFORhv%2F-LbrPsE2lifvD1RxViag%2Fimage.png?alt=media\&token=3168c2d0-5e96-403e-bbda-cf8e904d73e7)

$$
\hat{q}(s,a,w) = x(s,a)^Tw = \sum x\_j(s,a)w\_j
$$

### Algorithms

Now we have to substitute a target for $$q\_\pi(S,A)$$&#x20;

**MC:** The taget wiill be return $$G\_t$$&#x20;

$$
\triangle w = \alpha (G\_t-\hat{q}(S,A,w))\triangle\_w \hat{q}(S,A,w)
$$

**TD(0):** The target is TD target $$R\_{t+1} + γ \hat{q} (S\_{t+1}, A\_{t+1} , w)$$

$$
\triangle w = \alpha (R\_{t+1} + γ \hat{q} (S\_{t+1}, A\_{t+1} , w)-\hat{q}(S,A,w))\triangle\_w \hat{q}(S,A,w)
$$

**Same goes with forward and backward view TD(lambda).**

![Sarsa Control with action-value funtion approximation](https://1877261540-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LFDuA0A2VRqmT31Blrq%2F-Lbn9KEfr19fEnrFORhv%2F-LbrVt6tEoRhYo7kuSfm%2Fimage.png?alt=media\&token=256cfe91-8d11-43f5-916b-2f864de43eb3)

### Gradient Temporal-Difference Learning

## Batch Methods

### Least Square Prediction

Given value function $$\hat{v}(s,w)$$ approximation amd experience D consisting of \<state, value> pairs. $$D = {\<s\_1, v\_1^\pi>,\<s\_2, v\_2^\pi>,...\<s\_t, v\_t^\pi>}$$&#x20;

Learning paramters w for the best fitting value funtion$$\hat{v}(s,w)$$.

Least-square algorithm: find $$w$$ minimizing the least square error between apprx funtion and target values.&#x20;

$$
LS(w) = \sum\_{t=1}^T(v\_t^\pi - \hat{v}(s\_t,w))^2 \\
\=E\_D\[(v^\pi-\hat{v}(s,w))^2]
$$

#### SGD with experience replay

* Sample state, value from experience\
  &#x20;$$\<s,v^\pi> \sim D$$
* Apply SGD update\
  &#x20;$$\triangle w = \alpha (v^\pi-\hat{v}(s,w))\triangledown\_w\hat{v}(s,w)$$

This converges to least square solution

$$
w^\pi = arg \min\_w LS(w)
$$

#### Experience Replay in Deep Q-Networks

![Deep Q network using experience replay](https://1877261540-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LFDuA0A2VRqmT31Blrq%2F-Lbn9KEfr19fEnrFORhv%2F-LbrdQFWKpLahuxlvhm-%2Fimage.png?alt=media\&token=26c34b26-05a2-44a6-83dc-7a8a01f56573)

### Least Square Control

#### Least square Action-Value Function Approximation

Approximate action-value funtion.&#x20;

Minimize least square error between $$\hat{q}(s,a,w)$$ and $${q}\_\pi(s,a,w)$$ from experiences generated using policy $$\pi$$ consisting of <(state, actoin), value> pairs $$D = {<(s\_1,a\_1), v\_1^\pi>,<(s\_2,a\_2), v\_2^\pi>,...<(s\_t,a\_t), v\_t^\pi>}$$

![](https://1877261540-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LFDuA0A2VRqmT31Blrq%2F-Lbn9KEfr19fEnrFORhv%2F-LbrrpRsk9i3VqUrQJLG%2Fimage.png?alt=media\&token=b37f5dd3-4a59-46e5-bcca-aa25ad22c0c1)

![](https://1877261540-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LFDuA0A2VRqmT31Blrq%2F-Lbn9KEfr19fEnrFORhv%2F-LbrruDYnpjE7RWxbqUe%2Fimage.png?alt=media\&token=fb35abc7-126f-4114-b53b-f7b9e4a7f2aa)

![](https://1877261540-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LFDuA0A2VRqmT31Blrq%2F-Lbn9KEfr19fEnrFORhv%2F-LbrrxgfTNgrgi4aTdvH%2Fimage.png?alt=media\&token=166a0a6a-19b6-4a17-b7b1-f5155041bf6d)
