> For the complete documentation index, see [llms.txt](https://theshank.gitbook.io/ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://theshank.gitbook.io/ai/reinforcement-learning/deep-reinforcement-learning.md).

# Deep Reinforcement Learning

## Goal of Deep RL

As deep RL have parameters $$\theta$$ , hence our objective is to get $$\theta$$ such that **reward** of each state in an episode is maximized.&#x20;

![](https://1877261540-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LFDuA0A2VRqmT31Blrq%2F-LiDTIjPcP3AVzaKqN36%2F-Li8d2VaQFWgXv187NLx%2Fimage.png?alt=media\&token=7addda0d-d75e-4730-aee6-e7eafc101894)

**Note: Here we are not considering discounted reward**

## Finite & Infinite Horizon Objective Function

![](https://1877261540-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LFDuA0A2VRqmT31Blrq%2F-LiDTIjPcP3AVzaKqN36%2F-Li8kGJpaO4e69CRuCIf%2Fimage.png?alt=media\&token=231c30fe-d73e-4ea8-acd4-0ce7e247bfe8)

![](https://1877261540-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LFDuA0A2VRqmT31Blrq%2F-LiDTIjPcP3AVzaKqN36%2F-Li8kUp0CiIuz5wc58qq%2Fimage.png?alt=media\&token=ff9aaa3e-9dec-4e17-a584-b2fda788c165)

### Stationary Distribution: Markov Chains

A **stationary distribution** of a [Markov chain](https://brilliant.org/wiki/markov-chains/) is a probability distribution that remains unchanged in the Markov chain as time progresses. Typically, it is represented as a row vector π whose entries are probabilities summing to 1, and given [transition matrix](https://brilliant.org/wiki/markov-chains/#transition-matrices) $$\textbf{P}$$ , it satisfies

$$
\pi = \pi \textbf{P}
$$

Once you are in a stationary distribution, you will remain in a stationary distribution. Stationery distribution is a vector conataining probabilities of being in state for each state in corresponding element.\
Seeing the equation above, eigen vector of a transition matrix P is always an stationary distribution.&#x20;

## Expectation makes objective function smooth

![](https://1877261540-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LFDuA0A2VRqmT31Blrq%2F-LiDTIjPcP3AVzaKqN36%2F-Li98qdSzGUUOwexYgwb%2Fimage.png?alt=media\&token=23c6952e-3710-46cd-9876-aa5fbcb8c8f4)

Taking expectation of a function makes it smoother, allowing it differential wrt to the parameter. Look at example above. The reward function above is non-differentiable, but expectation makes it differentiable, making gradient based learning feasible for RL.&#x20;

## Types of RL Algos

![](https://1877261540-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LFDuA0A2VRqmT31Blrq%2F-LiDTIjPcP3AVzaKqN36%2F-LiC1wcltFfxhVn3UiQZ%2Fimage.png?alt=media\&token=63a5cd54-2a7e-4877-bf1a-ab8058462991)

### Model-Based RL Algo

![](https://1877261540-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LFDuA0A2VRqmT31Blrq%2F-LiDTIjPcP3AVzaKqN36%2F-LiC8XZWdcUVNNRs0AJN%2Fimage.png?alt=media\&token=4a829c23-6c7a-49fc-a186-e6a2461b2d57)

### Value-Based RL Algo

![](https://1877261540-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LFDuA0A2VRqmT31Blrq%2F-LiDTIjPcP3AVzaKqN36%2F-LiC9KvFzAz546OxjZ7v%2Fimage.png?alt=media\&token=3a8344a4-107e-4338-acaf-b0b00900ee78)

### Direct Policy Gradients

![](https://1877261540-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LFDuA0A2VRqmT31Blrq%2F-LiDTIjPcP3AVzaKqN36%2F-LiCFjKOMZjrpjFS0ppA%2Fimage.png?alt=media\&token=715037a2-2911-4e53-b0b3-6275d6db9e3c)

### Actor-Critic

![](https://1877261540-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LFDuA0A2VRqmT31Blrq%2F-LiDTIjPcP3AVzaKqN36%2F-LiCGMbhqDK56JteVy1d%2Fimage.png?alt=media\&token=c944af18-07b9-448f-a2d3-498cecc4a2df)

## Trade-Offs

### Sample Efficiency

![](https://1877261540-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LFDuA0A2VRqmT31Blrq%2F-LiDTIjPcP3AVzaKqN36%2F-LiCMC9mPlJX6W5aDt4r%2Fimage.png?alt=media\&token=cbc46692-fbaf-4b35-b943-2acb171bd53d)

**Off-Policy:** Able to improve the policy without generating new samples from that policy\
**On-policy:** each time the policy is changed, even a little bit, we need to generate new samples

* **Conventional Policy Gradient Methods are on-Policy. New samples are generated each time with a updated policy.**&#x20;
* **Actor-Critic method can either be on-policy or off-policy depending upon the details**
* **Model based are more efficient as it is intuitive itself, because having a model can reduce the need of samples.**

### Convergence and Stability

![](https://1877261540-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LFDuA0A2VRqmT31Blrq%2F-LiDTIjPcP3AVzaKqN36%2F-LiCWUIoqREhSI08NzZ3%2Fimage.png?alt=media\&token=21af5f9c-6dee-4c57-a6a8-08b209b90f75)

* Model-based RL method are gradient descent to get the best model. i.e. we are fit the for the model but nowhere maximizing the reward.&#x20;
