NLP（十九）双向LSTM情感分类模型

原文链接：http://www.one2know.cn/nlp19/

使用IMDB情绪数据来比较CNN和RNN两种方法，预处理与上节相同

from __future__ import print_function

import numpy as np

import pandas as pd

from keras.preprocessing import sequence

from keras.models import Sequential

from keras.layers import Dense,Dropout,Embedding,LSTM,Bidirectional

from keras.datasets import imdb

from sklearn.metrics import accuracy_score,classification_report

# 限制最大的特征数

max_features = 15000

max_len = 300

batch_size = 64

# 加载数据

(x_train,y_train),(x_test,y_test) = imdb.load_data(num_words=max_features)

print(len(x_train),'train observations')

print(len(x_test),'test observations')

输出：

Using TensorFlow backend.

25000 train observations

25000 test observations

如何实现

1.预处理

2.LSTM模型的构建和验证

3.模型评估
代码

from __future__ import print_function

import numpy as np

import pandas as pd

from keras.preprocessing import sequence

from keras.models import Sequential

from keras.layers import Dense,Dropout,Embedding,LSTM,Bidirectional

from keras.datasets import imdb

from sklearn.metrics import accuracy_score,classification_report

# 限制最大的特征数

max_features = 15000

max_len = 300

batch_size = 64

# 加载数据

(x_train,y_train),(x_test,y_test) = imdb.load_data(num_words=max_features)

# print(len(x_train),'train observations')

# print(len(x_test),'test observations')

# 通过序列填充将所有的数据整合为一个固定维度，提高运行效率

x_train_2 = sequence.pad_sequences(x_train,maxlen=max_len)

x_test_2 = sequence.pad_sequences(x_test,maxlen=max_len)

print('x_train_2.shape:',x_train_2.shape)

print('x_test_2.shape:',x_test_2.shape)

y_train = np.array(y_train)

y_test = np.array(y_test)

# keras框架 => 双向LSTM模型

# 双向LSTM网络有前向和后向连接，使句子中的单词可以同时与左右词汇产生连接

model = Sequential()

model.add(Embedding(max_features,128,input_length=max_len)) # 嵌入层将维数降到128

model.add(Bidirectional(LSTM(64))) # 双向LSTM层

model.add(Dropout(0.5)) # 随机失活

model.add(Dense(1,activation='sigmoid')) # 稠密层 将情感分类0或1

model.compile('adam','binary_crossentropy',metrics=['accuracy']) # 二元交叉熵

print(model.summary())

model.fit(x_train_2,y_train,batch_size=batch_size,epochs=4,validation_split=0.2)

# 预测及结果

y_train_predclass = model.predict_classes(x_train_2,batch_size=1000)

y_test_predclass = model.predict_classes(x_test_2,batch_size=1000)

y_train_predclass.shape = y_train.shape

y_test_predclass.shape = y_test.shape

print('\n\nLSTM Bidirectional Sentiment Classification - Train accuracy:',

      round(accuracy_score(y_train,y_train_predclass),3))

print('\nLSTM Bidirectional Sentiment Classification of Training data\n',

      classification_report(y_train,y_train_predclass))

print('\nLSTM Bidirectional Sentiment Classification - Train Confusion Matrix\n\n',

      pd.crosstab(y_train,y_train_predclass,rownames=['Actuall'],colnames=['Predicted']))

print('\nLSTM Bidirectional Sentiment Classification - Test accuracy:',

      round(accuracy_score(y_test,y_test_predclass),3))

print('\nLSTM Bidirectional Sentiment Classification of Test data\n',

      classification_report(y_test,y_test_predclass))

print('\nLSTM Bidirectional Sentiment Classification - Test Confusion Matrix\n\n',

      pd.crosstab(y_test,y_test_predclass,rownames=['Actuall'],colnames=['Predicted']))

输出：

Using TensorFlow backend.

x_train_2.shape: (25000, 300)

x_test_2.shape: (25000, 300)

WARNING:tensorflow:From D:\Python37\Lib\site-packages\tensorflow\python\framework\op_def_library.py:263: colocate_with (from tensorflow.python.framework.ops) is deprecated and will be removed in a future version.

Instructions for updating:

Colocations handled automatically by placer.

WARNING:tensorflow:From D:\Anaconda3\lib\site-packages\keras\backend\tensorflow_backend.py:3445: calling dropout (from tensorflow.python.ops.nn_ops) with keep_prob is deprecated and will be removed in a future version.

Instructions for updating:

Please use `rate` instead of `keep_prob`. Rate should be set to `rate = 1 - keep_prob`.

_________________________________________________________________

Layer (type)                 Output Shape              Param #

=================================================================

embedding_1 (Embedding)      (None, 300, 128)          1920000

_________________________________________________________________

bidirectional_1 (Bidirection (None, 128)               98816

_________________________________________________________________

dropout_1 (Dropout)          (None, 128)               0

_________________________________________________________________

dense_1 (Dense)              (None, 1)                 129

=================================================================

Total params: 2,018,945

Trainable params: 2,018,945

Non-trainable params: 0

_________________________________________________________________

None

WARNING:tensorflow:From D:\Python37\Lib\site-packages\tensorflow\python\ops\math_ops.py:3066: to_int32 (from tensorflow.python.ops.math_ops) is deprecated and will be removed in a future version.

Instructions for updating:

Use tf.cast instead.

Train on 20000 samples, validate on 5000 samples

Epoch 1/4

2019-07-07 20:03:45.649853: I tensorflow/core/platform/cpu_feature_guard.cc:141] Your CPU supports instructions that this TensorFlow binary was not compiled to use: AVX2

   64/20000 [..............................] - ETA: 18:21 - loss: 0.6915 - acc: 0.5781

  128/20000 [..............................] - ETA: 13:04 - loss: 0.6918 - acc: 0.5938

  192/20000 [..............................] - ETA: 11:14 - loss: 0.6915 - acc: 0.5729

  256/20000 [..............................] - ETA: 10:19 - loss: 0.6917 - acc: 0.5469

  320/20000 [..............................] - ETA: 9:45 - loss: 0.6915 - acc: 0.5469

  此处省略一堆epoch的一堆操作

LSTM Bidirectional Sentiment Classification - Train accuracy: 0.955

LSTM Bidirectional Sentiment Classification of Training data

               precision    recall  f1-score   support

           0       0.96      0.95      0.95     12500

           1       0.95      0.96      0.95     12500

    accuracy                           0.95     25000

   macro avg       0.95      0.95      0.95     25000

weighted avg       0.95      0.95      0.95     25000

LSTM Bidirectional Sentiment Classification - Train Confusion Matrix

 Predicted      0      1

Actuall

0          11928    572

1            561  11939

LSTM Bidirectional Sentiment Classification - Test accuracy: 0.859

LSTM Bidirectional Sentiment Classification of Test data

               precision    recall  f1-score   support

           0       0.86      0.86      0.86     12500

           1       0.86      0.85      0.86     12500

    accuracy                           0.86     25000

   macro avg       0.86      0.86      0.86     25000

weighted avg       0.86      0.86      0.86     25000

LSTM Bidirectional Sentiment Classification - Test Confusion Matrix

 Predicted      0      1

Actuall

0          10809   1691

1           1829  10671

time============== 2080.618681907654

NLP（十九）双向LSTM情感分类模型的更多相关文章

NLP学习（2）----文本分类模型
实战:https://github.com/jiangxinyang227/NLP-Project 一.简介: 1.传统的文本分类方法:[人工特征工程+浅层分类模型] (1)文本预处理: ①(中文) ...
pytorch LSTM情感分类全部代码
先运行main.py进行文本序列化,再train.py模型训练 dataset.py from torch.utils.data import DataLoader,Dataset import to ...
tensorflow学习笔记(三十九):双向rnn
tensorflow 双向 rnn 如何在tensorflow中实现双向rnn 单层双向rnn 单层双向rnn (cs224d) tensorflow中已经提供了双向rnn的接口,它就是tf.nn.b ...
Python之路【第二十九篇】:django ORM模型层
ORM简介 MVC或者MVC框架中包括一个重要的部分,就是ORM,它实现了数据模型与数据库的解耦,即数据模型的设计不需要依赖于特定的数据库,通过简单的配置就可以轻松更换数据库,这极大的减轻了开发人员的 ...
PaddlePaddle︱开发文档中学习情感分类（CNN、LSTM、双向LSTM）、语义角色标注
PaddlePaddle出教程啦,教程一部分写的很详细,值得学习. 一期涉及新手入门.识别数字.图像分类.词向量.情感分析.语义角色标注.机器翻译.个性化推荐. 二期会有更多的图像内容. 随便,帮国产 ...
[DeeplearningAI笔记]序列模型2.9情感分类
5.2自然语言处理觉得有用的话,欢迎一起讨论相互学习~Follow Me 2.9 Sentiment classification 情感分类情感分类任务简单来说是看一段文本,然后分辨这个人是否喜欢 ...
kaggle——Bag of Words Meets Bags of Popcorn（IMDB电影评论情感分类实践）
kaggle链接:https://www.kaggle.com/c/word2vec-nlp-tutorial/overview 简介:给出 50,000 IMDB movie reviews,进行0 ...
基于双向LSTM和迁移学习的seq2seq核心实体识别
http://spaces.ac.cn/archives/3942/ 暑假期间做了一下百度和西安交大联合举办的核心实体识别竞赛,最终的结果还不错,遂记录一下.模型的效果不是最好的,但是胜在“端到端”, ...
NLP文本情感分类传统模型+深度学习（demo）
文本情感分类: 文本情感分类(一):传统模型摘自:http://spaces.ac.cn/index.php/archives/3360/ 测试句子:工信处女干事每月经过下属科室都要亲口交代24口交 ...

随机推荐

勘误：EOS资源抵押退还
关键字:勘误,delegatebw,undelegatebw,listbw,资源管理,抵押,解抵押,返还资源 EOS中,资源抵押与解抵押是通过一对命令完成的:delegatebw,undelegate ...
Vue事件修饰符详解
整体学习Vue时看到Vue文档中有事件修饰符的描述,但是看了之后并没有理解是什么意思,于是查阅了资料,现在记录下来与大家分享先给大家画一个示意图理解一下冒泡和捕获 (1) .stop修饰符请看如下 ...
chapter01作业
1. 请用命令查出ifconfig命令程序的绝对路径 [root@localhost chen]# which ifconfig /usr/sbin/ifconfig 2.请用命令展示以下命令哪些是内 ...
多线程编程（Linux C）
多线程编程可以说每个程序员的基本功,同时也是开发中的难点之一,本文以Linux C为例,讲述了线程的创建及常用的几种线程同步的方式,最后对多线程编程进行了总结与思考并给出代码示例. 一.创建线程多线 ...
【React踩坑记四】React项目中引入并使用js-xlsx上传插件(结合antdesign的上传组件)
最近有一个前端上传并解析excel/csv表格数据的需求. 于是在github上找到一个14K star的前端解析插件 github传送门官方也有,奈何实在太过于浅薄.于是做了以下整理,避免道友们少 ...
ieda控制台缓冲区限制问题
一.现象控制台输出数据若超过默认值时,将从后向前取默认值大小数据(1024) 二.解决方案 1.配置文件(idea安装目录/bin/idea.properties) 2.找到该栏:idea.cycl ...
Springmvc的运行原理 SpringMvc的优点
SpringMVC框架运行原理 1:客户端发送请求到前端控制器(DispatcherServlet),前端控制器根据请求信息(url),查询一个或多个HandlerMapping, 前端控制器,来决定 ...
如何获取app中的toast
前言 Toast是什么呢?在这个手机飞速发展的时代,app的种类也越来越多,那们在日常生活使用中,经常会发现,当你在某个app的输入框输入非法字符或者非法执行某个流程时,经常看到系统会给你弹出一个黑色 ...
带你剖析WebGis的世界奥秘----点和线的世界
前言昨天写了好久的博文我没保存,今天在来想继续写居然没了,气死人啊这种情况你们见到过没,所以今天重新写,我还是切换到了HTML格式的书写上.废话不多说了,我们现在就进入主题,上周我仔细研究了WebG ...
JMM浅析
背景学习Java并发编程,JMM是绕不过的槛.在Java规范里面指出了JMM是一个比较开拓性的尝试,是一种试图定义一个一致的.跨平台的内存模型.JMM的最初目的,就是为了能够支多线程程序设计的,每个 ...

NLP（十九） 双向LSTM情感分类模型

NLP（十九） 双向LSTM情感分类模型的更多相关文章

随机推荐

热门专题

NLP（十九）双向LSTM情感分类模型

NLP（十九）双向LSTM情感分类模型的更多相关文章