Coursera, Deep Learning 5, Sequence Models, week2, Natural Language Processing & Word Embeddings

mashuai_191 2024-08-25 15:21:36 原文

Word embeding

给word 加feature，用来区分word 之间的不同，或者识别word之间的相似性.

　　

用于学习 Embeding matrix E 的数据集非常大，比如 1B - 100B 的word corpos. 所以即使你输入的是没见过的 durian cutivator 也知道和 orange farmer 很相近. 这是transfter learning 的一个case.

　　

　　

　　

　　

因为t-SNE 做了non-liner 的转化，所以在原来的300维空间的平行的向量在转化过后的2D空间里基本上不会再平行.

　　

看两个向量的相似性，可以用cosine similarity.

　　

　　

说了这么久 word embeding E 有这么大的用处，那E是怎么得到的呢？下面讲怎么learn embeding.

Learning embedings:Word2Vec&GloVe

就是怎么得到 E。

下图提到的在附近找一个词来做context 预测 target 词，就是Skip Gram 算法.

　　

研究人员发现如果你真想要 build language model，很自然的选择最近的一些words 作为context；如果你是为了learn word embeding, 你可以选择任意下面的context都能得到很好的结果.

　　

Word2Vec

前面讲了使用前4个word来学习参数 E 进而预测target word, 这里讲用1个word 来学习E并用来预测target word. 下图就是Word2Vec 的 Skip-grams model，还有一个叫 CBow.

　　

Word2Vec 最大的问题就是计算p(t|c) 时候计算量很大.下图提到了hierachial softmax 的概念是一个解决的思路.

　　

接下里有两种算法来解决上面讲到的softmax计算量大的问题

Negagtive Sampling

　　

这个negative sampling model 把一次性计算softmax (在我们的例子里softmax 输出的维度是10k), 改成了多（10k）次(1个positive example, 4个随机选取的negative example)计算sigmoid.

　　

到底怎么随机选取negative sampling 呢？去绝对随机，还是有什么方法？有两种极端，一种是根据经验取常见的word, 还有一种是在vacab 里绝对的随机选取，都不好，下面给出了一个经验公式.但是我没搞懂那个p(wi) 怎么用到实际的选取中.

　　

GloVe

　　

　　

　　

Sentiment Classification

情感分析一个大问题是没有足够的label dataset.

Coursera, Deep Learning 5, Sequence Models, week2, Natural Language Processing & Word Embeddings的更多相关文章

Coursera, Deep Learning 5, Sequence Models, week3, Sequence models & Attention mechanism
Sequence to Sequence models basic sequence-to-sequence model: basic image-to-sequence or called imag ...
课程五(Sequence Models)，第二周（Natural Language Processing & Word Embeddings） —— 1.Programming assignments：Operations on word vectors - Debiasing
Operations on word vectors Welcome to your first assignment of this week! Because word embeddings ar ...
Coursera Deep Learning笔记序列模型（二）NLP & Word Embeddings(自然语言处理与词嵌入)
参考 1. Word Representation 之前介绍用词汇表表示单词,使用one-hot 向量表示词,缺点:它使每个词孤立起来,使得算法对相关词的泛化能力不强. 从上图可以看出相似的单词分布距 ...
Coursera, Deep Learning 5, Sequence Models, week1 Recurrent Neural Networks
有哪些sequence model Notation: RNN - Recurrent Neural Network 传统NN 在解决sequence input 时有什么问题? RNN就没有上面的问 ...
课程五(Sequence Models)，第二周（Natural Language Processing & Word Embeddings） —— 2.Programming assignments：Emojify
Emojify! Welcome to the second assignment of Week 2. You are going to use word vector representation ...
课程五(Sequence Models)，第二周（Natural Language Processing & Word Embeddings） —— 0.Practice questions：Natural Language Processing & Word Embeddings
[解释] The dimension of word vectors is usually smaller than the size of the vocabulary. Most common s ...
Predicting effects of noncoding variants with deep learning–based sequence model | 基于深度学习的序列模型预测非编码区变异的影响
Predicting effects of noncoding variants with deep learning–based sequence model PDF Interpreting no ...
吴恩达《深度学习》-课后测验-第五门课序列模型(Sequence Models)-Week 2: Natural Language Processing and Word Embeddings (第二周测验：自然语言处理与词嵌入)
Week 2 Quiz: Natural Language Processing and Word Embeddings (第二周测验:自然语言处理与词嵌入) 1.Suppose you learn ...
[C5W2] Sequence Models - Natural Language Processing and Word Embeddings
第二周自然语言处理与词嵌入(Natural Language Processing and Word Embeddings) 词汇表征(Word Representation) 上周我们学习了 RN ...

随机推荐

Manjaro下带供电的USB Hub提示error -71
问题描述这款USB Hub是绿联出的1转7带供电的白色款. 在lsusb中显示为 Bus 004 Device 023: ID 05e3:0616 Genesys Logic, Inc. hub B ...
MySQL数据库简单查询
--黑马程序员 DQL数据查询语言数据库执行DQL语句不会对数据进行改变,而是让数据库发送结果集给客户端.查询返回的结果集是一张虚拟表. 查询关键字:SELECT 语法: SELECT 列名 FRO ...
Hibernate4
内容简介:1.使用log4j的日志存储,2.一对一关系,3.二级缓存 1 整合log4j(了解) l slf4j 核心jar : slf4j-api-1.6.1.jar .slf4j是 ...
交叉编译jpeglib遇到的问题
由于要在开发板中加载libjpeg,不能使用gcc编译的库文件给以使用,需要自己配置使用另外的编译器编译该库文件. /usr/bin/ld: .libs/jaricom.o: Relocations ...
toString()和toLocaleString()有什么区别
偶然之间用到这两个方法然后在数字转换成字符串的时候,并没有感觉这两个方法有什么区别,如下: 1 2 3 4 5 6 7 8 var e=123 e.toString() "123& ...
JS判断一个数是否为质数
function isPrime(number) { if (typeof number !== 'number' || number<2) { // 不是数字或者数字小于2 return fa ...
HTML学习笔记Day6
一.元素类型 1.元素类型分类依据和元素类型分类根据css显示分类,XHTML元素被分为三种类型:块状元素.内联元素.行内块元素.可变元素 2.块状元素 1)块状元素在网页中就是以块的形式显示,所谓 ...
bzoj1027 状压dp
https://www.lydsy.com/JudgeOnline/problem.php?id=1072 题意给一个数字串s和正整数d, 统计s有多少种不同的排列能被d整除试了一下发现暴力可过 ...
Oracle_异常
问题1 描述:plsql客户端列值中的中文都成了问号分析:客户端和服务端编码不一致所致解决:1.查询服务端数据库编码 SQL> select userenv('language') from ...
Linux网卡调优篇-禁用ipv6与优化socket缓冲区大小
Linux网卡调优篇-禁用ipv6与优化socket缓冲区大小作者:尹正杰版权声明:原创作品,谢绝转载!否则将追究法律责任. 一般在内网环境中,我们几乎是用不到IPV6,因此我们没有必要把多不 ...