voc-fcn-alexnet网络结构理解

一、写在前面

fcn是首次使用cnn来实现语义分割的，论文地址：fully convolutional networks for semantic segmentation

实现代码地址：https://github.com/shelhamer/fcn.berkeleyvision.org

全卷积神经网络主要使用了三种技术：

1. 卷积化（Convolutional）

2. 上采样（Upsample）

3. 跳跃结构（Skip Layer）

为了便于理解，我拿最简单的结构voc-fcn-alexnet进行说明，该网络结构主要用到了前面两个技术，不包含跳跃结构。

二、voc-fcn-alexnet 的train.prototxt文件

layer {

  name: "data"

  type: "Python"

  top: "data"

  top: "label"

  python_param {

    module: "voc_layers"

    layer: "SBDDSegDataLayer"

    param_str: "{\'sbdd_dir\': \'../data/sbdd/dataset\', \'seed\': 1337, \'split\': \'train\', \'mean\': (104.00699, 116.66877, 122.67892)}"

  }

}

layer {

  name: "conv1"

  type: "Convolution"

  bottom: "data"

  top: "conv1"

  convolution_param {

    num_output:

    pad:

    kernel_size:

    group:

    stride:

  }

}

layer {

  name: "relu1"

  type: "ReLU"

  bottom: "conv1"

  top: "conv1"

}

layer {

  name: "pool1"

  type: "Pooling"

  bottom: "conv1"

  top: "pool1"

  pooling_param {

    pool: MAX

    kernel_size:

    stride:

  }

}

layer {

  name: "norm1"

  type: "LRN"

  bottom: "pool1"

  top: "norm1"

  lrn_param {

    local_size:

    alpha: 0.0001

    beta: 0.75

  }

}

layer {

  name: "conv2"

  type: "Convolution"

  bottom: "norm1"

  top: "conv2"

  convolution_param {

    num_output:

    pad:

    kernel_size:

    group:

    stride:

  }

}

layer {

  name: "relu2"

  type: "ReLU"

  bottom: "conv2"

  top: "conv2"

}

layer {

  name: "pool2"

  type: "Pooling"

  bottom: "conv2"

  top: "pool2"

  pooling_param {

    pool: MAX

    kernel_size:

    stride:

  }

}

layer {

  name: "norm2"

  type: "LRN"

  bottom: "pool2"

  top: "norm2"

  lrn_param {

    local_size:

    alpha: 0.0001

    beta: 0.75

  }

}

layer {

  name: "conv3"

  type: "Convolution"

  bottom: "norm2"

  top: "conv3"

  convolution_param {

    num_output:

    pad:

    kernel_size:

    group:

    stride:

  }

}

layer {

  name: "relu3"

  type: "ReLU"

  bottom: "conv3"

  top: "conv3"

}

layer {

  name: "conv4"

  type: "Convolution"

  bottom: "conv3"

  top: "conv4"

  convolution_param {

    num_output:

    pad:

    kernel_size:

    group:

    stride:

  }

}

layer {

  name: "relu4"

  type: "ReLU"

  bottom: "conv4"

  top: "conv4"

}

layer {

  name: "conv5"

  type: "Convolution"

  bottom: "conv4"

  top: "conv5"

  convolution_param {

    num_output:

    pad:

    kernel_size:

    group:

    stride:

  }

}

layer {

  name: "relu5"

  type: "ReLU"

  bottom: "conv5"

  top: "conv5"

}

layer {

  name: "pool5"

  type: "Pooling"

  bottom: "conv5"

  top: "pool5"

  pooling_param {

    pool: MAX

    kernel_size:

    stride:

  }

}

layer {

  name: "fc6"

  type: "Convolution"

  bottom: "pool5"

  top: "fc6"

  convolution_param {

    num_output:

    pad:

    kernel_size:

    group:

    stride:

  }

}

layer {

  name: "relu6"

  type: "ReLU"

  bottom: "fc6"

  top: "fc6"

}

layer {

  name: "drop6"

  type: "Dropout"

  bottom: "fc6"

  top: "fc6"

  dropout_param {

    dropout_ratio: 0.5

  }

}

layer {

  name: "fc7"

  type: "Convolution"

  bottom: "fc6"

  top: "fc7"

  convolution_param {

    num_output:

    pad:

    kernel_size:

    group:

    stride:

  }

}

layer {

  name: "relu7"

  type: "ReLU"

  bottom: "fc7"

  top: "fc7"

}

layer {

  name: "drop7"

  type: "Dropout"

  bottom: "fc7"

  top: "fc7"

  dropout_param {

    dropout_ratio: 0.5

  }

}

layer {

  name: "score_fr"

  type: "Convolution"

  bottom: "fc7"

  top: "score_fr"

  param {

    lr_mult:

    decay_mult:

  }

  param {

    lr_mult:

    decay_mult:

  }

  convolution_param {

    num_output:

    pad:

    kernel_size:

  }

}

layer {

  name: "upscore"

  type: "Deconvolution"

  bottom: "score_fr"

  top: "upscore"

  param {

    lr_mult:

  }

  convolution_param {

    num_output:

    bias_term: false

    kernel_size:

    stride:

  }

}

layer {

  name: "score"

  type: "Crop"

  bottom: "upscore"

  bottom: "data"

  top: "score"

  crop_param {

    axis:

    offset:

  }

}

layer {

  name: "loss"

  type: "SoftmaxWithLoss"

  bottom: "score"

  bottom: "label"

  top: "loss"

  loss_param {

    ignore_label:

    normalize: true

  }

}

三、网络结构

假设输入的图片为500x500，

根据train.prototxt文件，可以得到上图的网络结构，该网络结构除了前五层的卷积层，也把后面的三层改为了卷积层，score_fr是卷积层的最后一层，也叫heatmap热图，热图就是我们最重要的高维特诊图，得到高维特征的heatmap之后，就是最重要的一步也是最后的一步，对原图像进行upsampling（即反卷积），把图像进行放大，得到原图像的大小。

四、损失函数

该网络的损失函数为SoftmaxWithLoss。首先进行softmax求解，求出每个像素点属于不同类别的概率，因为总共是分为21类，所以每个像素点对应21个概率值（输出通道数为21）。然后求解每个像素点所属实际类别概率的log值之和的平均，再取负数，可得到损失函数，参考如下：

end

voc-fcn-alexnet网络结构理解的更多相关文章

pascalcontext-fcn全卷积网络结构理解
一.说明 fcn的开源代码:https://github.com/shelhamer/fcn.berkeleyvision.org 论文地址:fully convolutional networks ...
Alexnet网络结构
最近试一下kaggle的文字检测的题目,目前方向有两个ssd和cptn.直接看看不太懂,看到Alexnet是基础,今天手写一下网络,记录一下啊. 先理解下Alexnet中使用的原件和作用: 激活函数使 ...
Xception网络结构理解
Xception网络是由inception结构加上depthwise separable convlution,再加上残差网络结构改进而来/ 常规卷积是直接通过一个卷积核把空间信息和通道信息直接提取出 ...
深入理解AlexNet网络
原文地址:https://blog.csdn.net/luoluonuoyasuolong/article/details/81750190 AlexNet论文:<ImageNet Classi ...
LeNet, AlexNet, VGGNet, GoogleNet, ResNet的网络结构
1. LeNet 2. AlexNet 3. 参考文献: 1. 经典卷积神经网络结构——LeNet-5.AlexNet.VGG-16 2. 初探Alexnet网络结构 3.
深度学习与CV教程(14) | 图像分割 (FCN,SegNet,U-Net,PSPNet,DeepLab,RefineNet)
作者:韩信子@ShowMeAI 教程地址:http://www.showmeai.tech/tutorials/37 本文地址:http://www.showmeai.tech/article-det ...
【深度学习系列】用PaddlePaddle和Tensorflow实现AlexNet
上周我们用PaddlePaddle和Tensorflow实现了图像分类,分别用自己手写的一个简单的CNN网络simple_cnn和LeNet-5的CNN网络识别cifar-10数据集.在上周的实验表现 ...
【深度学习系列】用PaddlePaddle和Tensorflow实现经典CNN网络AlexNet
上周我们用PaddlePaddle和Tensorflow实现了图像分类,分别用自己手写的一个简单的CNN网络simple_cnn和LeNet-5的CNN网络识别cifar-10数据集.在上周的实验表现 ...
tensorflow学习笔记——AlexNet
1,AlexNet网络的创新点 AlexNet将LeNet的思想发扬光大,把CNN的基本原理应用到了很深很宽的网络中.AlexNet主要使用到的新技术点如下: (1)成功使用ReLU作为CNN的激活函 ...

随机推荐

Unity ECS 初探
1.安装安装两个包 2.初探实例化注:实例化的实体并不会在Hierarchy视图里面显示,可在EntityDebugger窗口里面显示,因此需要显示的话需要添加Rendermeshcompone ...
关于几天来研究使用css3动画的一点总结
1. 为避免合成线程频繁计算导致性能降低, 使用transform属性,尽量避免使用width / height / padding / left 等. 2. 侧边栏下阴影遮罩层动画, 如果用back ...
url_encode和base64
在用一个某开源插件做封装,想要传一些参数进去. 多数字段都是普通字符串参数,但是有一个字段传的是json,结果发现这个插件一看到大括号和双引号就识别错误了. 不想改这个插件的源码,考虑自己传进去的时候 ...
Python 语法2
文档字符串,这个只能是用于紧跟函数第二行的内容. 输出文档说明的部分,代码格式是函数名称.(dot)__(双下划线)doc__(双下划线) ///////////////////////////// ...
task打印执行结果
使用debug输出task执行的register: - name: check extract session # script: /app/ansiblecfg/XXX/roles/test/tas ...
[转载]Fiddler 解析！抓包抓得好真的可以为所欲为 [一]
说起抓包,很多人以为就是用个工具,简简单单地抓一下就可以了.昨天在面试一个安卓逆向,直接告诉我[抓包没有技术含量].在这里,我必须发一个教程,解析一下抓包神器——Fiddler.Fiddler仅仅是一 ...
转载 * jQuery实现动态分割div—通过拖动分隔栏实现上下、左右动态改变左右、上下两个相邻div的大小
由jQuery实现上下.左右动态改变左右.上下两个div的大小,需要自己引入jquery1.8.0.min.js包可用于页面布局. //============================ind ...
vue+koa实现简单的图书小程序（2）
记录一下实现我们图书的扫码功能: https://developers.weixin.qq.com/miniprogram/dev/api/scancode.html要多读文档 scanBook () ...
"Loading a plug-in failed The plug-in or one of its prerequisite plug-ins may be missing or damaged and may need to be reinstalled"
The Unarchiver 虽好,但存在问题比我们在mac上zip打包一个软件xcode, 然后copy to another mac, 这时用The Unarchiver解压缩出来的xcode包不 ...
数组中array==null和array.length==0的区别
//代码public class Test1 { public static void main(String[] args) { int[] a1 = new int[0]; int[] a2 = ...

voc-fcn-alexnet网络结构理解

voc-fcn-alexnet网络结构理解的更多相关文章

随机推荐

热门专题