Reinforcement Learning Q-learning 算法学习-3

//Q-learning 源码分析。 
import java.util.Random;

public class QLearning1

{

    private static final int Q_SIZE = 6;

    private static final double GAMMA = 0.8;

    private static final int ITERATIONS = 10;

    private static final int INITIAL_STATES[] = new int[] {1, 3, 5, 2, 4, 0};

    private static final int R[][] = new int[][] {{-1, -1, -1, -1, 0, -1},

                                                  {-1, -1, -1, 0, -1, 100},

                                                  {-1, -1, -1, 0, -1, -1},

                                                  {-1, 0, 0, -1, 0, -1},

                                                  {0, -1, -1, 0, -1, 100},

                                                  {-1, 0, -1, -1, 0, 100}};

    private static int q[][] = new int[Q_SIZE][Q_SIZE];

    private static int currentState = 0;

    private static void train()

    {

        initialize();

        // Perform training, starting at all initial states.

        for(int j = 0; j < ITERATIONS; j++)

        {

            for(int i = 0; i < Q_SIZE; i++)

            {

                episode(INITIAL_STATES[i]);

            } // i

        } // j

        System.out.println("Q Matrix values:");

        for(int i = 0; i < Q_SIZE; i++)

        {

            for(int j = 0; j < Q_SIZE; j++)

            {

                System.out.print(q[i][j] + ",\t");

            } // j

            System.out.print("\n");

        } // i

        System.out.print("\n");

        return;

    }

    private static void test()

    {

        // Perform tests, starting at all initial states.

        System.out.println("Shortest routes from initial states:");

        for(int i = 0; i < Q_SIZE; i++)

        {

            currentState = INITIAL_STATES[i];

            int newState = 0;

            do

            {

                newState = maximum(currentState, true);

                System.out.print(currentState + ", ");

                currentState = newState;

            }while(currentState < 5);

            System.out.print("5\n");

        }

        return;

    }

    private static void episode(final int initialState)

    {

        currentState = initialState;

        // Travel from state to state until goal state is reached.

        do

        {

            chooseAnAction();

        }while(currentState == 5);

        // When currentState = 5, Run through the set once more for convergence.

        for(int i = 0; i < Q_SIZE; i++)

        {

            chooseAnAction();

        }

        return;

    }

    private static void chooseAnAction()

    {

        int possibleAction = 0;

        // Randomly choose a possible action connected to the current state.

        possibleAction = getRandomAction(Q_SIZE);

        if(R[currentState][possibleAction] >= 0){

            q[currentState][possibleAction] = reward(possibleAction);

            currentState = possibleAction;

        }

        return;

    }

    private static int getRandomAction(final int upperBound)

    {

        int action = 0;

        boolean choiceIsValid = false;

        // Randomly choose a possible action connected to the current state.

        while(choiceIsValid == false)

        {

            // Get a random value between 0(inclusive) and 6(exclusive).

            action = new Random().nextInt(upperBound);

            if(R[currentState][action] > -1){

                choiceIsValid = true;

            }

        }

        return action;

    }

    private static void initialize()

    {

        for(int i = 0; i < Q_SIZE; i++)

        {

            for(int j = 0; j < Q_SIZE; j++)

            {

                q[i][j] = 0;

            } // j

        } // i

        return;

    }

    private static int maximum(final int State, final boolean ReturnIndexOnly)

    {

        // If ReturnIndexOnly = True, the Q matrix index is returned.

        // If ReturnIndexOnly = False, the Q matrix value is returned.

        int winner = 0;

        boolean foundNewWinner = false;

        boolean done = false;

        while(!done)

        {

            foundNewWinner = false;

            for(int i = 0; i < Q_SIZE; i++)

            {

                if(i != winner){             // Avoid self-comparison.

                    if(q[State][i] > q[State][winner]){

                        winner = i;

                        foundNewWinner = true;

                    }

                }

            }

            if(foundNewWinner == false){

                done = true;

            }

        }

        if(ReturnIndexOnly == true){

            return winner;

        }else{

            return q[State][winner];

        }

    }

    private static int reward(final int Action)

    {

        return (int)(R[currentState][Action] + (GAMMA * maximum(Action, false)));

    }

    public static void main(String[] args)

    {

        train();

        test();

        return;

    }

}

Reinforcement Learning Q-learning 算法学习-3的更多相关文章

Reinforcement Learning Q-learning 算法学习-2
在阅读了Q-learning 算法学习-1文章之后. 我分析了这个算法的本质. 算法本质个人分析. 1.算法的初始状态是随机的,所以每个初始状态都是随机的,所以每个初始状态出现的概率都一样的.如果训练 ...
增强学习（五）----- 时间差分学习(Q learning, Sarsa learning)
接下来我们回顾一下动态规划算法(DP)和蒙特卡罗方法(MC)的特点,对于动态规划算法有如下特性: 需要环境模型,即状态转移概率\(P_{sa}\) 状态值函数的估计是自举的(bootstrapping ...
强化学习9-Deep Q Learning
之前讲到Sarsa和Q Learning都不太适合解决大规模问题,为什么呢? 因为传统的强化学习都有一张Q表,这张Q表记录了每个状态下,每个动作的q值,但是现实问题往往极其复杂,其状态非常多,甚至是连 ...
机器学习实战（Machine Learning in Action）学习笔记————08.使用FPgrowth算法来高效发现频繁项集
机器学习实战(Machine Learning in Action)学习笔记————08.使用FPgrowth算法来高效发现频繁项集关键字:FPgrowth.频繁项集.条件FP树.非监督学习作者:米 ...
机器学习实战（Machine Learning in Action）学习笔记————07.使用Apriori算法进行关联分析
机器学习实战(Machine Learning in Action)学习笔记————07.使用Apriori算法进行关联分析关键字:Apriori.关联规则挖掘.频繁项集作者:米仓山下时间:2018 ...
机器学习实战（Machine Learning in Action）学习笔记————06.k-均值聚类算法（kMeans）学习笔记
机器学习实战(Machine Learning in Action)学习笔记————06.k-均值聚类算法(kMeans)学习笔记关键字:k-均值.kMeans.聚类.非监督学习作者:米仓山下时间: ...
机器学习实战（Machine Learning in Action）学习笔记————02.k-邻近算法（KNN）
机器学习实战(Machine Learning in Action)学习笔记————02.k-邻近算法(KNN) 关键字:邻近算法(kNN: k Nearest Neighbors).python.源 ...
强化学习_Deep Q Learning(DQN)_代码解析
Deep Q Learning 使用gym的CartPole作为环境,使用QDN解决离散动作空间的问题. 一.导入需要的包和定义超参数 import tensorflow as tf import n ...
如何用简单例子讲解 Q - learning 的具体过程？
作者:牛阿链接:https://www.zhihu.com/question/26408259/answer/123230350来源:知乎著作权归作者所有.商业转载请联系作者获得授权,非商业转载请注明 ...
机器学习实战（Machine Learning in Action）学习笔记————09.利用PCA简化数据
机器学习实战(Machine Learning in Action)学习笔记————09.利用PCA简化数据关键字:PCA.主成分分析.降维作者:米仓山下时间:2018-11-15机器学习实战(Ma ...

随机推荐

LeetCode：二叉树的前序遍历【144】
LeetCode:二叉树的前序遍历[144] 题目描述给定一个二叉树,返回它的前序遍历. 示例: 输入: [1,null,2,3] 1 \ 2 / 3 输出: [1,2,3] 题目分析如果用递 ...
60. Permutation Sequence（求全排列的第k个排列）
The set [1,2,3,…,n] contains a total of n! unique permutations. By listing and labeling all of the p ...
centos6.5系统python2.6升级到python3.6
1.安装必备的工具 wget:yum install wget gcc:yum install gcc zlib zlib-devel: yum install zlib zlib-devel -y ...
hadoop19---动态代理
Action调用service里面的方法,动态代理:改变方法的实现在方法前后加逻辑不是加新方法. 在学习Spring的时候,我们知道Spring主要有两大思想,一个是IoC,另一个就是AOP,对于Io ...
Tomcat 启动内存修改
内存修改文件 Windows 文件 /bin/catalina.bat Linux 文件 /bin/catalina.sh 方法一 # 设置参数 JAVA_OPTS='-Xms[初始化内存大小] -X ...
机器学习中的numpy库
日常学习中总是遇到数据需要处理等问题,这时候我们就可以借助numpy这个工具来做一些有意思的事. 1.生成随机数的几种方式 x=np.random.random(12) ###生成12 ...
jQuery单选框跟复选框美化
在线演示本地下载
Node.Js安装教程
Node.Js安装教程介绍下我的环境环境值操作系统 win10 64bit Node.Js 8.9.4 emmmm 表格中毒了,为什么出不来效果一.下载及安装这个可以去Node.Js官网上 ...
MyEclipse 为xml添加本地的dtd文件
在使用Eclipse或MyEclipse编辑XML文件的时候经常会碰到编辑器不提示的现象,这常常是因为其xml文件需要参考的DTD文件找不到,还有因为网络的问题不能及时提示而产生的.Eclipse/M ...
包嗅探和包回放 —tcpdump、tcpreplay--重放攻击
攻击方式:tcpdump 进行嗅探,获取报文消息:然后用tcpreplay回放攻击 arp欺骗可以使用 arpspoof kali linux有这三个工具转载地址https://www.cnblog ...

Reinforcement Learning Q-learning 算法学习-3

Reinforcement Learning Q-learning 算法学习-3的更多相关文章

随机推荐

热门专题