主角torch.nn.LSTM()

初始化时要传入的参数

 |  Args:

 |      input_size: The number of expected features in the input `x`

 |      hidden_size: The number of features in the hidden state `h`

 |      num_layers: Number of recurrent layers. E.g., setting ``num_layers=2``

 |          would mean stacking two LSTMs together to form a `stacked LSTM`,

 |          with the second LSTM taking in outputs of the first LSTM and

 |          computing the final results. Default: 1

 |      bias: If ``False``, then the layer does not use bias weights `b_ih` and `b_hh`.

 |          Default: ``True``

 |      batch_first: If ``True``, then the input and output tensors are provided

 |          as `(batch, seq, feature)` instead of `(seq, batch, feature)`.

 |          Note that this does not apply to hidden or cell states. See the

 |          Inputs/Outputs sections below for details.  Default: ``False``

 |      dropout: If non-zero, introduces a `Dropout` layer on the outputs of each

 |          LSTM layer except the last layer, with dropout probability equal to

 |          :attr:`dropout`. Default: 0

 |      bidirectional: If ``True``, becomes a bidirectional LSTM. Default: ``False``

 |      proj_size: If ``> 0``, will use LSTM with projections of corresponding size. Default: 0

input_size：一般是词嵌入的大小

hidden_size：隐含层的维度

num_layers：默认是1，单层LSTM

bias：是否使用bias

batch_first：默认为False，如果设置为True，则表示第一个维度表示的是batch_size

dropout：直接看英文吧

bidirectional：默认为False，表示单向LSTM，当设置为True，表示为双向LSTM，一般和num_layers配合使用（需要注意的是当该项设置为True时，将num_layers设置为1，表示由1个双向LSTM构成）

模型输入输出-单向LSTM

import torch

import torch.nn as nn

import numpy as np

inputs_numpy = np.random.random((64,32,300))

inputs = torch.from_numpy(inputs_numpy).to(torch.float32)

inputs.shape

torch.Size([64, 32, 300])：表示[batchsize, max_length, embedding_size]

hidden_size = 128

lstm = nn.LSTM(300, 128, batch_first=True, num_layers=1)

output, (hn, cn) = lstm(inputs)

print(output.shape)

print(hn.shape)

print(cn.shape)

torch.Size([64, 32, 128])

torch.Size([1, 64, 128])

torch.Size([1, 64, 128])

说明：

output：保存了每个时间步的输出，如果想要获取最后一个时间步的输出，则可以这么获取：output_last = output[:,-1,:]

h_n：包含的是句子的最后一个单词的隐藏状态，与句子的长度seq_length无关

c_n：包含的是句子的最后一个单词的细胞状态，与句子的长度seq_length无关

另外：最后一个时间步的输出等于最后一个隐含层的输出

output_last = output[:,-1,:]

hn_last = hn[-1]

print(output_last.eq(hn_last))

模型输入输出-双向LSTM

首先我们要明确：

output ：（seq_len, batch, num_directions * hidden_size）

h_n：(num_layers * num_directions, batch, hidden_size)

c_n ：（num_layers * num_directions, batch, hidden_size）

其中num_layers表示层数，这里是1，num_directions表示方向数，由于是双向的，这里是2，也是，我们就有下面的结果：

import torch

import torch.nn as nn

import numpy as np

inputs_numpy = np.random.random((64,32,300))

inputs = torch.from_numpy(inputs_numpy).to(torch.float32)

inputs.shape

hidden_size = 128

lstm = nn.LSTM(300, 128, batch_first=True, num_layers=1, bidirectional=True)

output, (hn, cn) = lstm(inputs)

print(output.shape)

print(hn.shape)

print(cn.shape)

torch.Size([64, 32, 256])

torch.Size([2, 64, 128])

torch.Size([2, 64, 128])

这里面的hn包含两个元素，一个是正向的隐含层输出，一个是方向的隐含层输出。

#获取反向的最后一个output

output_last_backward = output[:,0,-hidden_size:]

#获反向最后一层的hn

hn_last_backward = hn[-1]

#反向最后的output等于最后一层的hn

print(output_last_backward.eq(hn_last_backward))

#获取正向的最后一个output

output_last_forward = output[:,-1,:hidden_size]

#获取正向最后一层的hn

hn_last_forward = hn[-2]

# 反向最后的output等于最后一层的hn

print(output_last_forward.eq(hn_last_forward))

https://www.cnblogs.com/LiuXinyu12378/p/12322993.html

https://blog.csdn.net/m0_45478865/article/details/104455978

https://blog.csdn.net/foneone/article/details/104002372

关于torch.nn.LSTM()的输入和输出的更多相关文章

torch.nn.LSTM()函数维度详解
123456789101112lstm=nn.LSTM(input_size, hidden_size, num_la ...
PyTorch官方中文文档：torch.nn
torch.nn Parameters class torch.nn.Parameter() 艾伯特(http://www.aibbt.com/)国内第一家人工智能门户,微信公众号:aibbtcom ...
pytorch nn.LSTM()参数详解
输入数据格式:input(seq_len, batch, input_size)h0(num_layers * num_directions, batch, hidden_size)c0(num_la ...
pytorch中文文档-torch.nn.init常用函数-待添加
参考:https://pytorch.org/docs/stable/nn.html torch.nn.init.constant_(tensor, val) 使用参数val的值填满输入tensor ...
pytorch中文文档-torch.nn常用函数-待添加-明天继续
https://pytorch.org/docs/stable/nn.html 1)卷积层 class torch.nn.Conv2d(in_channels, out_channels, kerne ...
torch.nn.functional中softmax的作用及其参数说明
参考:https://pytorch-cn.readthedocs.io/zh/latest/package_references/functional/#_1 class torch.nn.Soft ...
torch.nn.Embedding理解
Pytorch官网的解释是:一个保存了固定字典和大小的简单查找表.这个模块常用来保存词嵌入和用下标检索它们.模块的输入是一个下标的列表,输出是对应的词嵌入. torch.nn.Embedding(nu ...
pytorch torch.nn.functional实现插值和上采样
interpolate torch.nn.functional.interpolate(input, size=None, scale_factor=None, mode='nearest', ali ...
pytorch torch.nn 实现上采样——nn.Upsample
Vision layers 1)Upsample CLASS torch.nn.Upsample(size=None, scale_factor=None, mode='nearest', align ...

随机推荐

python之读取excel实例演示
1.基础知识点击这里 import openpyxl def read_excel(workbook,sheetname=None): wd=openpyxl.load_workbook(workbo ...
并发王者课-铂金1：探本溯源-为何说Lock接口是Java中锁的基础
欢迎来到<并发王者课>,本文是该系列文章中的第14篇. 在黄金系列中,我们介绍了并发中一些问题,比如死锁.活锁.线程饥饿等问题.在并发编程中,这些问题无疑都是需要解决的.所以,在铂金系列文 ...
H5简介（转）
H5究竟是什么? "HTML5(WEB前端)技术由HTML(结构).CSS(样式).JavaScript(行为)组成.HTML5是WEB的未来,HTML5不仅在PC端,更是在移动端上也有广泛 ...
【NX二次开发】NX内部函数，libuifw.dll文件中的内部函数
本文分为两部分:"带参数的函数"和 "带修饰的函数". 浏览这篇博客前请先阅读: [NX二次开发]NX内部函数,查找内部函数的方法带参数的函数: void U ...
【C++】vector容器的用法
检测vector容器是否为空: 1 #include <iostream> 2 #include <string> 3 #include <vector> 4 us ...
遇到禁止复制该怎么办？幸好我会Python...
相信大家都有遇到这种情况(无法复制): 或者是这种情况以上这种情况都是网页无法复制文本的情况.不过这些对于Python来说都不是问题.今天辰哥就叫你们用Python去解决. 思路:利用pdfkit库 ...
docker创建和使用mysql
container和image是两种不同的概念,image即指存在的镜像,container指docker运行起来后image的实例. 当使用docker kill 把某个正在运行的实例kill掉之后 ...
[202103] Interview Summary
整理 2021 March「偷」到的算法题. 题目: 阿里:https://codeforces.com/contest/1465/problem/C 字节:输出 LCS Jump Trading:给 ...
一个线上 Maven 诡异问题排查过程
å. 前言现在的大部分 Java 应用基本都是通过 Maven 进行组织的,不论是分布式应用还是单体集群应用往往都会通过一个父 POM 加若干子 POM 完成项目的组织.然而这种多应用多模块的拆分 ...
Golang十六进制字符串和byte数组互转
Golang十六进制字符串和byte数组互转需求 Golang十六进制字符串和byte数组互相转换,使用"encoding/hex"包实现Demo package main i ...