You're reading from Mastering Reinforcement Learning with Python

Product typeBook

Published inDec 2020

Reading LevelBeginner

PublisherPackt

ISBN-139781838644147

Edition1st Edition

Languages

Python

Tools

PyTorch

Concepts

Reinforcement Learning

Author (1)

Enes Bilgin

Chapter 11: Achieving Generalization and Overcoming Partial Observability

Deep reinforcement learning (RL) has achieved what was impossible with the earlier AI methods, such as beating world champions in games like Go, Dota 2, and StarCraft II. Yet, applying RL to real-world problems is still challenging. Two important obstacles to this end are generalization of trained policies to a broad set of environment conditions and developing policies that can handle partial observability. As we will see in the chapter, these are closely related challenges, for which we will present solution approaches.

Here is what we will cover in this chapter:

Focusing on generalization in reinforcement learning
Enriching agent experience via domain randomization
Using memory to overcome partial observability
Quantifying generalization via CoinRun

These topics are critical to understand for a successful implementation of RL in real-world settings. So, let's dive right in...

Focusing on generalization in reinforcement learning

The core goal in most machine learning projects is to obtain models that will work beyond training, and under a broad set of conditions during test time. Yet, when you start learning about RL, efforts to prevent overfitting and achieve generalization are not always at the forefront of the discussion, as opposed to how it is with supervised learning. In this section, we discuss what leads to this discrepancy, describe how generalization is closely related to partial observability in RL, and present a general recipe to handle these challenges.

Generalization and overfitting in supervised learning

When we train an image recognition or forecasting model, what we really want to achieve is high accuracy on unseen data. After all, we already know the labels for the data at hand. We use various methods to this end:

We use separate training, dev, and test sets, for model training, hyperparameter selection, and model performance...

Enriching agent experience via domain randomization

DR is simply about randomizing the parameters defining (part of) the environment during training to enrich the training data. It is a useful technique to obtain policies that are robust and generalizable, both in fully and partially observable environments. In this section, we first present a classification of such parameters, in other words, different dimensions of randomization. Then, we discuss two curriculum learning approaches to guide RL training along those dimensions.

Dimensions of randomization

Borrowed from (Rivlin, 2019), a useful categorization of how two environments belonging to the same problem class (e.g., autonomous driving) can differ is as follows.

Different observations for the same/similar states

In this case, two environments emit different observations although the underlying state and transition functions are the same or very similar. An example to this is the same Atari game scene but with different...

Using memory to overcome partial observability

A memory is nothing but a way of processing a sequence of observations as the input to the agent policy. If you worked with other types of sequence data with neural networks, such as in time series prediction or natural language processing (NLP), you can adopt similar approaches to use observation memory as the input your RL model.

Let's go into more details of how this can be done.

Stacking observations

A simple way of passing an observation sequence to the model is to stitch them together and treat this stack as a single observation. Denoting the observation at time as , we can form a new observation to be passed to the model as follows:

where is the length of the memory. Of course, for , we need to somehow initialize the earlier parts of the memory, such as using vectors of zeros that are the same dimension as .

In fact, simply stacking observations is how the original DQN work handled...

Quantifying generalization via CoinRun

There are various ways of testing whether certain algorithms/approaches generalize to unseen environment conditions better than others, such as:

Creating validation and test environments with separate sets of environment parameters,
Assessing policy performance in real-life deployment.

Real-life deployment may not necessarily be an option, so the latter is not always practical. The challenge with the former is to have consistency and to ensure that validation/test data are indeed not used in training. Also, it is possible to overfit to the validation environment when too many models are tried based on validation performance. One approach to overcome these challenges is to use procedurally generated environments. To this end, OpenAI has created the CoinRun environment to benchmark algorithms on their generalization capabilities. Let's look into it in more detail.

CoinRun environment

In the CoinRun environment, we have...

Summary

In this chapter, we have covered an important topic in RL: Generalization and partial observability, which are key for real-world applications. Note that this is an active research area: Keep our discussion here as directional suggestions and the first methods to try for your problem. New approaches come out periodically, so watch out for them. The important thing is you should always keep an eye on the generalization and partial observability for a successful RL implementation outside of video games. In the next section, we will take our expedition to yet a next advanced level with meta-learning. So, stay tuned!

References

Cobbe, K., Klimov, O., Hesse, C., Kim, T., & Schulman, J. (2018). Quantifying Generalization in Reinforcement Learning. Retrieved from ArXiv: https://arxiv.org/abs/1812.02341
Lee, K., Lee, K., Shin, J., & Lee, H. (2020). {Network Randomization: A Simple Technique for Generalization in Deep Reinforcement Learning. Retrieved from ArXiv: https://arxiv.org/abs/1910.05396
Rivlin, O. (2019, Nov 21). Generalization in Deep Reinforcement Learning. Retrieved from Towards Data Science: https://towardsdatascience.com/generalization-in-deep-reinforcement-learning-a14a240b155b
Cobbe, K., Klimov, O., Hesse, C., Kim, T., & Schulman, J. (2018). Quantifying Generalization in Reinforcement Learning: https://arxiv.org/abs/1812.0234
Lee, K., Lee, K., Shin, J., & Lee, H. (2020). "Network Randomization: A Simple Technique for Generalization in Deep Reinforcement Learning.": https://arxiv.org/abs/1910.0539
Parisotto, Emilio, et al. (2019) "...

The rest of the chapter is locked

You have been reading a chapter from

Mastering Reinforcement Learning with Python

Published in: Dec 2020Publisher: PacktISBN-13: 9781838644147

A free Packt account unlocks extra newsletters, articles, discounted offers, and much more. Start advancing your knowledge today.

undefined

Unlock this book and the full library FREE for 7 days

Get unlimited access to 7000+ expert-authored eBooks and videos courses covering every tech area you can think of

Start free trial

Renews at $15.99/month. Cancel anytime

Author (1)

Enes Bilgin

Enes Bilgin works as a senior AI engineer and a tech lead in Microsoft's Autonomous Systems division. He is a machine learning and operations research practitioner and researcher with experience in building production systems and models for top tech companies using Python, TensorFlow, and Ray/RLlib. He holds an M.S. and a Ph.D. in systems engineering from Boston University and a B.S. in industrial engineering from Bilkent University. In the past, he has worked as a research scientist at Amazon and as an operations research scientist at AMD. He also held adjunct faculty positions at the McCombs School of Business at the University of Texas at Austin and at the Ingram School of Engineering at Texas State University.
Read more about Enes Bilgin

Other recommended products

Related to this chapter

TensorFlow Reinforcement Learning Quick Start Guide

This book is an essential guide for anyone interested in Reinforcement Learning. The book provides an actionable reference for Reinforcement Learning algorithms and their applications using TensorFlow and Python. It will help readers leverage the power of algorithms such as Deep Q-Network (DQN), Deep Deterministic Policy Gradients (DDPG), and Proximal Policy Optimization (PPO) to solve challenging control and decision-making problems.

BookMar 2019184 pages

TensorFlow 2 Reinforcement Learning Cookbook

This cookbook will help you to gain a solid understanding of deep reinforcement learning (RL) algorithms with the help of concise, easy-to-follow implementations from scratch. You'll learn how to implement these algorithms with minimal code and develop AI applications to solve real-world and business problems using RL.

BookJan 2021472 pages

PyTorch 1.x Reinforcement Learning Cookbook

This book presents practical solutions to the most common reinforcement learning problems. The recipes in this book will help you understand the fundamental concepts to develop popular RL algorithms. You will gain practical experience in the RL domain using the modern offerings of the PyTorch 1.x library.

BookOct 2019340 pages

Deep Reinforcement Learning with Python

Deep Reinforcement Learning with Python - Second Edition will help you learn reinforcement learning algorithms, techniques and architectures – including deep reinforcement learning – from scratch. This new edition is an extensive update of the original, reflecting the state-of-the-art latest thinking in reinforcement learning.

BookSep 2020760 pages

Hands-On Reinforcement Learning with Python

Reinforcement learning is a self-evolving type of machine learning that takes us closer to achieving true artificial intelligence. This easy-to-follow guide explains everything from scratch using rich examples written in Python.

BookJun 2018318 pages

Reinforcement Learning Algorithms with Python

With this book, you will understand the core concepts and techniques of reinforcement learning. You will take a look into each RL algorithm and will develop your own self-learning algorithms and models. You will optimize the algorithms for better precision, use high-speed actions and lower the risk of anomalies in your applications.

BookOct 2019366 pages

Hands-On Intelligent Agents with OpenAI Gym

Walks through the hands-on process of building intelligent agents from the basics and all the way up to solving complex problems including playing Atari games and driving a car autonomously in the CARLA simulator. Discusses various learning environments and how to transform real-world problems into learning environments and solve using the agents.

BookJul 2018254 pages

Reinforcement Learning with TensorFlow

Reinforcement learning allows you to develop intelligent, self-learning systems. This book shows you how to put the concepts of Reinforcement Learning to train efficient models.You will use popular reinforcement learning algorithms to implement use-cases in image processing and NLP, by combining the power of TensorFlow and OpenAI Gym.

BookApr 2018334 pages

Hands-On Q-Learning with Python

Q-learning is the reinforcement learning approach behind Deep-Q-Learning and is a values-based learning algorithm in RL. This book will help you get comfortable with developing the effective agents for Q learning and also make you learn to effectively develop and deploy Deep Q networks for complex AI applications.

BookApr 2019212 pages

The Reinforcement Learning Workshop

With the help of practical examples and engaging activities, The Reinforcement Learning Workshop takes you through reinforcement learning’s core techniques and frameworks. Following a hands-on approach, it allows you to learn reinforcement learning at your own pace to develop your own intelligent applications with ease.

BookAug 2020822 pages

Hands-On Markov Models with Python

This book will help you become familiar with HMMs and different inference algorithms by working on real-world problems. You will start with an introduction to the basic concepts of Markov chains, Markov processes and then delve deeper into understanding hidden Markov models and its types using practical examples.

BookSep 2018178 pages

Practical Reinforcement Learning

Reinforcement learning (RL) is becoming a popular tool for constructing autonomous systems that improve themselves with experience. We will break the RL framework into its core building blocks, and provide you with details of each element. This book is divided into three parts. The first part defines Reinforcement Learning and describes the basics and the Python and Java frameworks, which we are going to use later in the book. The second part discusses learning techniques with basic algorithms such as Temporal Difference, Monte Carlo, and Policy Gradient—all with practical examples. Lastly, in the third part we apply Reinforcement Learning with the most recent and widely used algorithms, via practical applications.

BookOct 2017336 pages

Personalised recommendations for you

Based on your interests and search pattern

Et al.

Ever wonder why speech recognition systems don't understand the Scottish accent, or what would happen if an astronaut only ate mac 'n' cheese, or other spurious reflections you'd have at a bar? We did, then collated those deliberations into absurd research articles with fake figures and methodologies inspired by even more fictionally absurd studies.

BookAug 2023230 pages5

Generative AI with LangChain

This book is a comprehensive introduction to LLMs and LangChain, demystifying the basic mechanics of LangChain, its functionalities, and the myriad of applications it can be integrated into.

BookDec 2023360 pages4

Generative AI with LangChain

This book is a comprehensive introduction to LLMs and LangChain, demystifying the basic mechanics of LangChain, its functionalities, and the myriad of applications it can be integrated into.

BookDec 2023360 pages5

Generative AI with LangChain

This book is a comprehensive introduction to LLMs and LangChain, demystifying the basic mechanics of LangChain, its functionalities, and the myriad of applications it can be integrated into.

BookDec 2023360 pages1

Generative AI with LangChain

This book is a comprehensive introduction to LLMs and LangChain, demystifying the basic mechanics of LangChain, its functionalities, and the myriad of applications it can be integrated into.

BookDec 2023360 pages5

Mastering Tableau 2023

This book is a comprehensive resource to mastering your Tableau skills and becoming a BI expert. As you progress, you will learn how to build advanced dashboards and improve your storytelling to derive key business insight, as well as make you well-versed with advanced functionalities of Tableau in the business intelligence domain.

BookAug 2023684 pages

Building AI Applications with ChatGPT APIs

This guide covers all ChatGPT API features for effortless creation of robust AI powered apps. With its help, you’ll be able to leverage ChatGPT’s cutting-edge NLP models to take your app development skills to the next level. You’ll also work on ten exciting projects that will give you the practical know-how that you can apply to your existing applications.

BookSep 2023258 pages5

Building AI Applications with ChatGPT APIs

This guide covers all ChatGPT API features for effortless creation of robust AI powered apps. With its help, you’ll be able to leverage ChatGPT’s cutting-edge NLP models to take your app development skills to the next level. You’ll also work on ten exciting projects that will give you the practical know-how that you can apply to your existing applications.

BookSep 2023258 pages2

Data Engineering with AWS

Embark on a journey to master data engineering pipelines on AWS! Our book offers a hands-on experience of AWS services for ingesting, transforming, and consuming data. Whether you're an absolute beginner or someone with basic data engineering experience, this guide is an indispensable resource.

BookOct 2023636 pages5

Modern Data Architecture on AWS

Every organization wants an agile, performant, and cost-effective data platform that meets all their current and future business needs. Purpose-built AWS analytics services and their features play a big part in building such a modern data platform. This book brings to you all the design and architectural patterns that’ll help you achieve this goal.

BookAug 2023420 pages5

Practical Guide to Applied Conformal Prediction in Python

Discover the power of Conformal Prediction with the "Practical Guide to Applied Conformal Prediction in Python." Master the latest techniques to quantify uncertainty in machine learning and computer vision models, and seamlessly apply them to your industry applications.

BookDec 2023240 pages

TinyML Cookbook

With over 70 project-based recipes, the TinyML Cookbook is a practical guide that will help you to get the most out of your microcontrollers. It provides a comprehensive understanding of the theoretical foundations while giving you hands-on experience training ML models for deployment on Arduino Nano 33 BLE Sense, Raspberry Pi Pico, and SparkFun RedBoard Artemis Nano microcontrollers.

BookNov 2023664 pages