开篇第一篇就写一个paper reading吧,用markdown+vim写东西切换中英文挺麻烦的,有些就偷懒都用英文写了。

Stereo DSO: Large-Scale Direct Sparse Visual Odometry with Stereo Cameras

Abstract

Optimization objectives:

  1. intrinsic/extrinsic parameters of all keyframes
  2. all selected pixels' depth

Integrate constraints from static stereo (左右两个相机的立体视觉约束是静态的) into the bundle adjustment pipeline of temporal multi-view stereo.
Fixed-baseline stereo resolves scale drift.

? It also reduces the sensitivities to large optical flow and to rolling shutter effect which are known shortcomings of direct image alignment methods.

1. Introduction

stem from: working in an effective way
heuristically: 启发式的
hallucinate: 出现幻觉
strip down: reduced to its simplest form

Strasdat et al. proposed to expand the concept of keyframes to integrate scale and proposed a double window optimization (Figure out what is it)

Direct methods aim at computing geometry and motion directly from the images thereby skipping the intermediate keypoint selection step.

The key idea of LSD SLAM is to incrementally track the camera and simultaneously perform a pose graph optimization in order to keep the entire camera trajectory globally consistent. 作者认为这种方式没有减少累计误差,只是把它扩散到整个轨迹中( So the meaning of pose graph is? )。

Three drawbacks of DSO:

  1. The mentioned performance was gained on a photometrically calibrated dataset, in its absense, the performance would degrade.
  2. Scale drift
  3. DSO is quite sensitive to geometric distortion as those induces by fast motion and rolling shutter. While techniques for calibrating rolling shutter exist for direct SLAM algorithm, these are often quite involved and far from real-time capable.

Contribution:

  1. A stereo version of DSO. detail the proposed combination of temporal multi-view stereo and static stereo.
  2. Stereo DSO is good.

2. Direct Sparse VO with Stereo Camera

  • Absolute scale can be directly calculated from static stereo from the known baseline of the stereo camera
  • Static stereo can provide initial depth estimation for multi-view stereo
  • Static Stereo can only accurately triangulate 3D points within a limited depth range while this limit is resolved by temporal multi-view stereo.

New stereo frames are first tracked with respect to their reference keyframe in a coarse-to-fine mannar.

A joint optimization of their poses, affine brightness (两个参数:a和b) parameters, as well as the depts of all the observed 3D points and camera intrinsics, is performed.

2.1 Notation

Nothing important.

2.2 Direct Image Alignment Formulation

\[
E_{ij}=\sum_{p\in P_i}\omega_p \left\| I_j[p'] - I_i[p] \right\|_\gamma
\]

where \(\omega_p\) is the weight which is shown as follows.(梯度越大权重越小,不知道为啥)
\[
\omega_p = \frac{c^2}{c^2+\left\| \nabla I_i(p) \right\| ^2_2}
\]

光度误差对突然的光照变化非常敏感。

2.3 Tracking

All the potins inside the active window are projected into the new frame. Then the pose of the new frame is optimized by minimizing the energy function.
在之前的单目DSO中,用随机深度值来初始化,所以都会需要一个确定模式的移动来初始化。在本文中,因为这时候stereo image pair的affine brightness transfer factor是位置的,所以用NCC在水平极限上的3*5的领域中搜索。

2.4 Frame Management

The basic idea is to check if the scene or the illumination has sufficiently changed.

  • scene change: 用mean square optical flow和 mean squared optical flow without rotation between the current frame and the last keyframe来衡量。
  • illumination change: 用relative brightness factor \(|a_j - a_i|\) 来衡量。

-> 一个点如果是梯度大于一个阈值并且是一个block里最大的点,那么他会被选择。

-> Before a candidate point is activated and optimized in the windowed optimization, its inverse depth is constantly refined by the following non-keyframes. (找出来怎么做的)

-> 旧去新来:在边缘化点的时候把候选点加入到联合优化中。

-> The constraints from static stereo introduce scale information into the system, and they also provide good geometric priors to temporal multi-view stereo.

2.5 Windowed Optimization

-> Temporal Multi-View Stereo: 就一般的不同时刻的图片之间的立体视觉
-> Static Stereo:
-> Stereo Coupling: 为了平衡上两种约束的权重,我们引入了\(\lambda\)参数。
-> Margninalization: 在边缘化一个关键帧之前,我们首先会边缘化所有没有被过去两个关键帧看到所有active window中的点。

3. Evaluation

暂且略过不表

4. Conclusion

未来可以做的两件事:

  • Loop closuring and a database for map maintenance (LDSO半闲居士做过了)
  • Dynamic object handling to further boost the VO accuracy and robustness. (用深度学习做动态物体检测然后动的点不要了?)
    虽然自己在SLAM领域还有很多可以学习的,但是这样感觉直接法的东西也做完了?悲伤。。

Paper Reading: Stereo DSO的更多相关文章

  1. [Paper Reading]--Exploiting Relevance Feedback in Knowledge Graph

    <Exploiting Relevance Feedback in Knowledge Graph> Publication: KDD 2015 Authors: Yu Su, Sheng ...

  2. Paper Reading: Perceptual Generative Adversarial Networks for Small Object Detection

    Perceptual Generative Adversarial Networks for Small Object Detection 2017-07-11  19:47:46   CVPR 20 ...

  3. Paper Reading: In Defense of the Triplet Loss for Person Re-Identification

    In Defense of the Triplet Loss for Person Re-Identification  2017-07-02  14:04:20   This blog comes ...

  4. Paper Reading - Attention Is All You Need ( NIPS 2017 ) ★

    Link of the Paper: https://arxiv.org/abs/1706.03762 Motivation: The inherently sequential nature of ...

  5. Paper Reading - Convolutional Sequence to Sequence Learning ( CoRR 2017 ) ★

    Link of the Paper: https://arxiv.org/abs/1705.03122 Motivation: Compared to recurrent layers, convol ...

  6. Paper Reading - Deep Captioning with Multimodal Recurrent Neural Networks ( m-RNN ) ( ICLR 2015 ) ★

    Link of the Paper: https://arxiv.org/pdf/1412.6632.pdf Main Points: The authors propose a multimodal ...

  7. Paper Reading - Deep Visual-Semantic Alignments for Generating Image Descriptions ( CVPR 2015 )

    Link of the Paper: https://arxiv.org/abs/1412.2306 Main Points: An Alignment Model: Convolutional Ne ...

  8. Paper Reading - Mind’s Eye: A Recurrent Visual Representation for Image Caption Generation ( CVPR 2015 )

    Link of the Paper: https://ieeexplore.ieee.org/document/7298856/ A Correlative Paper: Learning a Rec ...

  9. Paper Reading - Show and Tell: A Neural Image Caption Generator ( CVPR 2015 )

    Link of the Paper: https://arxiv.org/abs/1411.4555 Main Points: A generative model ( NIC, GoogLeNet ...

随机推荐

  1. 【源码】HashMap源码及线程非安全分析

    最近工作不是太忙,准备再读读一些源码,想来想去,还是先从JDK的源码读起吧,毕竟很久不去读了,很多东西都生疏了.当然,还是先从炙手可热的HashMap,每次读都会有一些收获.当然,JDK8对HashM ...

  2. ElasticSearch(十二)删除数据插件delete-by-query

    在ElasticSearch2.0之后的版本中没有默认的delete-by-query,想使用此命令需要安装这个插件. 首先需要进入ES的目录 [root@node122 elasticsearch] ...

  3. day 04

    今天学些的内容 流程控制 1.在python一般代码执行顺序都从上到下依次解释执行称为顺序结构. 2.然而遇到需要一些条件判断的时候需要选择不同的执行路线去解释执行 这种流程控制称为 分支结构 也可叫 ...

  4. 机器学习总结(二)bagging与随机森林

    一:Bagging与随机森林 与Boosting族算法不同的是,Bagging和随机森林的个体学习器之间不存在强的依赖关系,可同时生成并行化的方法. Bagging算法 bagging的算法过程如下: ...

  5. 第一次远程ubuntu用c写Hello Word出现的问题

    2019-03-26 21:48:48 之前已经学过c语言,但是一直是在win10上用VC++6.0来写代码的,现在想尝试在linux下用c语言 其实主要目的是来学习linux 1.问题: 在ubun ...

  6. pyhton抛出自定义的异常

    用raise语句来引发一个异常.异常/错误对象必须有一个名字,且它们应是Error或Exception类的子类 下面是一个引发异常的例子: class ShortInputException(Exce ...

  7. Go 参数传递是传值还是传引用

    什么是传值(值传递)? 传值的意思是:函数传递的总是原来这个东西的一个副本.一个副拷贝.比如我们传递一个 int 类型的参数,传递 的其实这个参数的一个副本:传递一个指针类型的参数,其实传递的是这个指 ...

  8. C语言-第5次作业

    1.本章学习总结 1.1思维导图 1.2 本章学习体会及代码量学习体会 1.2.1学习体会 感受:和数组一样,这又是一个非常陌生的知识点--指针,刚刚开始学习的时候,被陌生的各种赋值方式搞得眼花缭乱, ...

  9. Lombok之使用详解

    前言 在Java中,封装是一个非常好的机制,最常见的封装莫过于get,set方法了,无论是Intellij idea 还是Eclipse,都提供了快速生成get,set方法的快捷键,使用起来很是方便, ...

  10. C# Selenium 破解腾讯滑动验证

    什么是Selenium? WebDriver是主流Web应用自动化测试框架,具有清晰面向对象 API,能以最佳的方式与浏览器进行交互. 支持的浏览器: Mozilla Firefox Google C ...