Preprocessing

Tokenizer

source code：https://github.com/keras-team/keras-preprocessing/blob/master/keras_preprocessing/text.py#L490-L519

some important functions and variables

init
def fit_on_texts(self, texts) #texts can be a string or a list of strings or a list of list of strings
self.word_index # the type of variance is dictonary, which contain a specific word subject to a unique index
self.index_word #r eserve the key and value of the word_index

sample

  import tensorflow as tf

  from tensorflow import keras

  # the package which can tokenizer

  from tensorflow.keras.preprocessing.text import Tokenizer

  '''

    transform the word into number

  '''

  sentences= ['i love my dog', 'i love my cat','you love my dog!']

  tokenizer = Tokenizer(num_words = 100)

  tokenizer.fit_on_texts(sentences)

  word_index = tokenizer.word_index

  print(word_index)

  # get the result {'love': 1, 'my': 2, 'i': 3, 'dog': 4, 'cat': 5, 'you': 6}

Serialization

texts_to_sequences(self,texts) # transforms each text in texts to a sequence of integers.
tf.keras.preprocessing.sequence.pad_sequences( sequences, maxlen=None, dtype='int32',padding='pre', truncating='pre', value=0.) # make the sentences with same length.
- sorce code https://github.com/tensorflow/tensorflow/blob/v2.5.0/tensorflow/python/keras/preprocessing/sequence.py#L88-L154

sample

sentences= ['i love my dog', 'i love my cat','you love my dog!','do you think my dog is amazing']

sequences = tokenizer.texts_to_sequences(sentences)

print(sequences)

 '''

   result is [[3, 1, 2, 4], [3, 1, 2, 5], [6, 1, 2, 4], [6, 2, 4]]

   which is not encoding for amazing, because it's not appear in fit texts

 '''

To solve this problem，we can set a oov in tokenizer to encode a word which not appear before.

tokenizer = Tokenizer(num_words = 100, oov_token = "<OOV>")

'''

    restart the code,we can get the result

    [[4, 2, 3, 5], [4, 2, 3, 6], [7, 2, 3, 5], [1, 7, 1, 3, 5, 1, 1]]

'''

but each sequences has the different length of the series, it's difficult for train a neuro network,so we need make the sequnces has the same length.

from tensorflow.keras.preprocessing.sequence import pad_sequences

padded_sequences = pad_sequences(sequences,

                                 padding = 'post',   # right padding

                                 maxlen = 5,         # max len of senquence

                                 truncating = 'post') # right cut

padded_sequences

'''

then we can get the result

array([[5, 3, 2, 4, 0],

       [5, 3, 2, 7, 0],

       [6, 3, 2, 4, 0],

       [8, 6, 9, 2, 4]])

'''

word processing in nlp with tensorflow的更多相关文章

论文阅读 | Text Processing Like Humans Do: Visually Attacking and Shielding NLP Systems
[code&data] [pdf] 主要工作文章首先证明了对抗攻击对NLP系统的影响力,然后提出了三种屏蔽方法: visual character embeddings adversaria ...
自然语言处理资源NLP
转自:https://github.com/andrewt3000/DL4NLP Deep Learning for NLP resources State of the art resources ...
NLP与深度学习（一）NLP任务流程
1. 自然语言处理简介根据工业界的估计,仅有21% 的数据是以结构化的形式展现的[1].在日常生活中,大量的数据是以文本.语音的方式产生(例如短信.微博.录音.聊天记录等等),这种方式是高度无结构化 ...
NLP新手入门指南|北大-TANGENT
开源的学习资源:<NLP 新手入门指南>,项目作者为北京大学 TANGENT 实验室成员. 该指南主要提供了 NLP 学习入门引导.常见任务的开发实现.各大技术教程与文献的相关推荐等内容, ...
分词（Tokenization） - NLP学习（1）
自从开始使用Python做深度学习的相关项目时,大部分时候或者说基本都是在研究图像处理与分析方面,但是找工作反而碰到了很多关于自然语言处理(natural language processing: N ...
TensorFlow系列专题（十一）：RNN的应用及注意力模型
磐创智能-专注机器学习深度学习的教程网站 http://panchuang.net/ 磐创AI-智能客服,聊天机器人,推荐系统 http://panchuangai.com/ 目录: 循环神经网络的应 ...
TensorFlow开发者证书中文手册
经过一个月的准备,终于通过了TensorFlow的开发者认证,由于官方的中文文档较少,为了方便大家了解这个考试,同时分享自己的备考经验,让大家少踩坑,我整理并制作了这个中文手册,请大家多多指正,有任何 ...
TensorFlow 在android上的Demo（1）
转载时请注明出处: 修雨轩陈系统环境说明: ------------------------------------ 操作系统 : ubunt 14.03 _ x86_64 操作系统内存: 8GB ...
初学者如何查阅自然语言处理（NLP）领域学术资料
1. 国际学术组织.学术会议与学术论文自然语言处理(natural language processing,NLP)在很大程度上与计算语言学(computational linguistics,CL ...

随机推荐

Blazor WebAssembly 渐进式 Web 应用程序 (PWA) 使用 LocalStorage 离线处理数据
原文链接:https://www.cnblogs.com/densen2014/p/16133343.html Window.localStorage 只读的localStorage 属性允许你访问一 ...
『忘了再学』Shell基础 — 11、变量定义的规则和分类
目录 1.定义变量的规则 2.变量的分类 1.定义变量的规则在定义变量时,有一些规则需要遵守变量名称可以由字母.数字和下划线组成,但是不能以数字开头.如果变量名是2name则是错误的. 在Bash ...
AcWing 【算法提高课】笔记02——搜索
搜索进阶 22.4.14 (PS:还有字串变换 A*两题生日蛋糕回转游戏没做) 感觉暂时用不上 BFS 1. Flood Fill 在线性时间复杂度内,找到某个点所在的连通块思路统计连通块 ...
从.net开发做到云原生运维(八)——DevOps实践
1. DevOps的一些介绍 DevOps(Development和Operations的组合词)是一组过程.方法与系统的统称,用于促进开发(应用程序/软件工程).技术运营和质量保障(QA)部门之间的 ...
树莓派开发笔记（十二）：入手研华ADVANTECH工控树莓派UNO-220套件（一）：介绍和运行系统
前言树莓派也可以做商业应用,工业控制,其稳定性和可靠性已经得到了验证,故而工业控制,一些停车场等场景也有采用树莓派作为主控的,本片介绍了研华ADVANTECH的树莓派套件组UNO-220-P4N ...
STS快捷键
在类或者方法上方加注释:shift+alt+J
ghostnet论文解析：ghost
创建日期: 2020-03-02 17:02:54 简介: GhostNet是2020CVPR录用的一篇对卷积操作进行改进的论文.文章的核心内容是Ghost模块(Ghost Module),可以用来替 ...
Java高并发-无锁
一.无锁类的原理 1.1 CAS CAS算法的过程是这样:它包含3个参数CAS(V,E,N).V表示要更新的变量,E表示预期值,N表示新值.仅当V值等于E值时,才会将V的值设为N,如果V值和E值不同, ...
使用虚拟机在3台centos7系统安装docker和k8s集群
一.安装docker 环境:准备3台centos7系统,都安装上docker环境,具体安装步骤和流程如下参考: https://docs.docker.com/install/linux/docke ...
BUUCTF刷题记录（更新中...）
极客大挑战 2019]EasySQL-1 直接通过输入万能密码:' or 1=1#实现注入: 思考:服务端sql语句应该为:select * from users where username='xx ...

word processing in nlp with tensorflow

Preprocessing

Tokenizer

some important functions and variables

sample

Serialization

sample

word processing in nlp with tensorflow的更多相关文章

随机推荐

热门专题