Multi-shot Pedestrian Re-identification via Sequential Decision Making

Multi-shot Pedestrian Re-identification via Sequential Decision Making

2019-07-31 20:33:37

Paper: http://openaccess.thecvf.com/content_cvpr_2018/papers/Zhang_Multi-Shot_Pedestrian_Re-Identification_CVPR_2018_paper.pdf

Code: https://github.com/TuSimple/rl-multishot-reid

1. Background and Motivation:

本文引入 DRL 到 person re-ID 任务，通过序列决策来完成难易样本的识别问题。主要动机如下图所示：

2. The Proposed Method：

2.1 Image-level feature extraction:

作者对图像特征提取，采用了多个组合损失函数的形式，即：classification loss, pairwise verification loss, and triplet verification loss。用了两种经典的骨干网络，即：Inception-BN 和 AlexNet。作者将一个序列中所有图像的 feature 进行聚合，得到 l2-normalized features，即：

并且根据 l (*, *) 进行 identities 的排序，即：

2.2 Sequence Level Feature Aggregation :

作者将该问题看做是 Markov Decision Processes (MDP), 表达为 (S, A, T, R)。在每一个步骤中，agent 将会从两个输入序列中得到一个选择的图像对，来观察 state，然后选择一个动作，接下来该 agent 将会得到一个奖励 r。在此之后，如果序列没有结束，该智能体将会接收下一个 image pair，然后得到一个新的 state。

Actions and Transitions:

首先随机的从两个序列中，选择两个图像，构成 image pair。然后将该样本对输入到 agent 中，agent 会输出三个动作：same, different, and unsure。前两个动作将会停止当前的 episode，然后即可输出当前的结果。作者认为当智能体收集到了足够的信息，并且足够自信来进行决策的时候，就可以及时停止以避免不必要的计算代价。如果智能体选择的 action 是 unsure，那么我们将会选择其他的 image pair 来进行判别。

Rewards：我们定义如下的奖励情况：

如果 agent 给定的结果和 gt 一致，那么给定 +1 的奖励；

如果 at 与 gt 不同，奖励将是 -1；要么当 t = $t_{max}$ 时，at 仍然是 unsure 的时候；

当 t < $t_{max}$，$a_t$ 是 unsure 的时候，奖励是 $r_p$ ；

这里的 rp 可能是 + 也可能是 -，具体看情况：If rp is negative, it will be penalized for requesting more pairs; on the other hand, if rp is positive, we encourage the agent to gather more pairs, and stop gathering when it has collected $t_{max}$ pairs to avoid a penalty of -1. 这个值，将会极大地影响最终 agent 的行为。

States and Deep Q-learning:

我们使用 deep Q-learning 来找到最优的策略。对于每一个 state and action $(s_t, a_t)$, $Q(s_t, a_t)$ 代表了当前状态和动作下的折扣的累积奖励。在训练阶段，我们可以迭代的更新 Q-function：

在时刻 t，状态 st 由如下的三个部分构成：

1). the first part is the observation $o_t$，即图像的特征；

2). the second part is a weighted average of the difference between historical image features of two sequences; 权重计算方法如下：

3). we also augment the image features with hand-crafted features for better discrimination.

3. Experimental Results：

Multi-shot Pedestrian Re-identification via Sequential Decision Making的更多相关文章

Parallel Gradient Boosting Decision Trees
本文转载自:链接 Highlights Three different methods for parallel gradient boosting decision trees. My algori ...
ICCV 2017论文分析（文本分析）标题词频分析这算不算大数据第一步：数据清洗（删除作者和无用的页码）
IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017. IEE ...
ICLR 2013 International Conference on Learning Representations深度学习论文papers
ICLR 2013 International Conference on Learning Representations May 02 - 04, 2013, Scottsdale, Arizon ...
metasploit-post模块信息
Name Disclosure Date Rank Description ---- ...
Andrew Ng机器学习公开课笔记–Reinforcement Learning and Control
网易公开课,第16课 notes,12 前面的supervised learning,对于一个指定的x可以明确告诉你,正确的y是什么但某些sequential decision making问题,比 ...
Learning Structured Representation for Text Classification via Reinforcement Learning 学习笔记
Representation learning : 表征学习,端到端的学习 pre-specified 预先指定的 demonstrate 论证;证明,证实;显示,展示;演示,说明 attempt ...
David Silver强化学习Lecture1：强化学习简介
课件:Lecture 1: Introduction to Reinforcement Learning 视频:David Silver深度强化学习第1课 - 简介 (中文字幕) 强化学习的特征作为 ...
（转）Applications of Reinforcement Learning in Real World
Applications of Reinforcement Learning in Real World 2018-08-05 18:58:04 This blog is copied from: h ...
论文笔记之：SeqGAN: Sequence generative adversarial nets with policy gradient
SeqGAN: Sequence generative adversarial nets with policy gradient AAAI-2017 Introduction : 产生序列模拟数 ...

随机推荐

Flink源码分析 - 剖析一个简单的Flink程序
本篇文章首发于头条号Flink程序是如何执行的?通过源码来剖析一个简单的Flink程序,欢迎关注头条号和微信公众号"大数据技术和人工智能"(微信搜索bigdata_ai_tech) ...
Python爬虫系列：四、Cookie的使用
Cookie,指某些网站为了辨别用户身份.进行session跟踪而储存在用户本地终端上的数据(通常经过加密) 比如说有些网站需要登录后才能访问某个页面,在登录之前,你想抓取某个页面内容是不允许的.那么 ...
SpringBoot2.x服务器端主动推送技术
一.服务端推送常用技术介绍服务端主流推送技术:websocket.SSE等 1.客户端轮询:ajax定时拉取后台数据 js setInterval定时函数 + ajax异步加载定时向服务 ...
redis发布订阅实现各类定时业务（优惠券过期，商品不支付自动撤单，自动收货等）
修改redis配置文件找到机器上redis配置文件conf/redis.conf,新增一行 notify-keyspace-events Ex 最后的Ex代表监听失效的键值修改后效果如下图代码 ...
Fuel
1. fuel简介 fuel是Mirantis公司提供的一款开源的自动化安装部署OpenStack的工具.为OpenStack相关的社区项目和插件的部署和管理提供了一种直观的GUI驱动体验. Fuel ...
小程序~列表渲染~key
如果列表中项目的位置会动态改变或者有新的项目添加到列表中,并且希望列表中的项目保持自己的特征和状态(如 <input/> 中的输入内容, <switch/> 的选中状态),需要 ...
mysql 查询今天，昨天，上个月sql语句
今天 select * from 表名 where to_days(时间字段名) = to_days(now()); 昨天Select * FROM 表名 Where TO_DAYS( NOW( ) ...
Mybatis框架-联表查询显示问题解决
需求:查询结果要求显示用户名,用户密码,用户的角色因为在用户表中只有用户角色码值,没有对应的名称,角色名称是在码表smbms_role表中,这时我们就需要联表查询了. 这里需要在User实体类中添加 ...
LeetCode 1048. Longest String Chain
原题链接在这里:https://leetcode.com/problems/longest-string-chain/ 题目: Given a list of words, each word con ...
Windbg妙用
计算器当你在调试,需要做一些从十六进制到十进制的简单转换,一些整数计算你不需要切换到calc.exe,你可以只使用windbg的表达式计算器.假设你得到了一个十六进制的大小,比如说2e903000, ...

Multi-shot Pedestrian Re-identification via Sequential Decision Making

Multi-shot Pedestrian Re-identification via Sequential Decision Making的更多相关文章

随机推荐

热门专题