#论文阅读# Universial language model fine-tuing for text classification

论文链接：https://aclweb.org/anthology/P18-1031

对文章内容的总结

文章研究了一些在general corous上pretrain LM，然后把得到的model transfer到text classiffication上整个过程的训练技巧。

这些技巧的切入点是learning rate. 主要是三个：

（1）discriminative fine-tuning (其中的discriminative 指 fine-tune each layer with different learning rate LR)

（2）slanted triangular learning rate （在训练过程中先增加LR，增到预设的最大值后减小(减小速度＜增加速度，所以LR随训练步数的曲线看起来是slanted triangle)）

（3）在训练text classiffication model时， perform gradual unfreezing. (即先锁住所有层的参数，训练过程中从最后一层开始，每训练一个epoch向前放开一层)

以下是ABSTACT和INTRODUCTION主要内容的翻译：

Abstract：

Inductive transfer learning 已经在很大程度上影响了CV，但在NLP领域仍然需要task-specific微调或需要从头开始训练。这篇文章提出了一个Universial language model Fine-tuning （ULMFiT）,这是一种可以应用于任何NLP任务的迁移学习方法。此外，文章还介绍了几种主要的fine-tuning language model 的技术。实验证明：ULMFiT outperforms the state-of-the-art on 三种文本分类任务（共计6个数据集）。

实验结果：

Introduction:

p1: Inductive transfer learning 对CV应用于CV时，rarely需要从头开始训练，只要fine-tuning from models that have been pretrained on ImageNet, MS-COCO等。

P2: Text classification属于NLP tasks with rea-world applications such as spam, fraud, and bot detection, emergency response and commercial document classification,such as for legal discovery.

P3: DL models 在很多NLP tasks 上取得了state-of-the-art，这些models是从头开始训练，requiring large datasets 和days to converge. NLP领域中的transfer learning 研究大多是 transductive transfer. 而inductive transfer，如fine-tuning pre-trained word embeddings 这种只是针对第一层的transfer technique，已经在实际中有了large impact，并且也被应用到了很多state-of-the-art models。但是recent approaches 在使用embeddings时只是把它们作为fixed parameters从头开始训练main task model，这样做limit了这些embedding的作用。

P4: 按照pretraining的思路，我们可以 do better than randomly initializing 模型的其他参数。However，有文献说inductive transfer via fine-tuning has been unsuccessful.

P5: 本文并不是想强调LM fine-tuning这个想法，而是要指出对模型进行有效训练的技术的缺乏才是阻碍transfer learning应用的关键所在。 LMs overfit to small datasets and suffered catastrophic forgetting when fine-tuned with a classifier. 跟CV相比，NLP models 非常shallow，它们需要不同的fine-tuning methods.

P6: 本文提出ULMFiT来解决上述issuses, 并且在any NLP task上得到了robust inductive transfer. ULMFiT 的architecture是 3-layer LSTM。各层使用相同的超参数，除了 tuned dropout hyperparameters 没有任何额外东西。ULMFiT 能够 outperforms highly engineered models.

Contributions:

(1) 提出ULMFiT，一种可以在any NLP task上achieve CV-like transfer learning的方法。

(2) 提出用于retain previous knowledge进而avoid catastrophic forgetting的novel techniques: discriminative fine-tuning, slanted triangular learning rate, and gradual unfreezing.

(3) siginificantly outperform the state-of-the-art on six representative text classification datasets, with an error reduction of 18-24% on the majority of datasets.

(4) 本文方法能够实现非常sample-efficient 的transfer learning，并且做了extensive ablation analysis.

(5) 作者们预训练了模型并且可用于wider adoption.

#论文阅读# Universial language model fine-tuing for text classification的更多相关文章

论文笔记 - Noisy Channel Language Model Prompting for Few-Shot Text Classification
Direct && Noise Channel 进一步把语言模型推理的模式分为了: 直推模式(Direct): 噪声通道模式(Noise channel). 直观来看: Direct ...
论文列表——text classification
https://blog.csdn.net/BitCs_zt/article/details/82938086 列出自己阅读的text classification论文的列表,以后有时间再整理相应的笔 ...
论文分享|《Universal Language Model Fine-tuning for Text Classificatio》
https://www.sohu.com/a/233269391_395209 本周我们要分享的论文是<Universal Language Model Fine-tuning for Text ...
【论文翻译】KLMo: Knowledge Graph Enhanced Pretrained Language Model with Fine-Grained Relationships
KLMo:建模细粒度关系的知识图增强预训练语言模型 (KLMo: Knowledge Graph Enhanced Pretrained Language Model with Fine-Graine ...
论文阅读笔记 Word Embeddings A Survey
论文阅读笔记 Word Embeddings A Survey 收获 Word Embedding 的定义 dense, distributed, fixed-length word vectors, ...
论文阅读笔记 Improved Word Representation Learning with Sememes
论文阅读笔记 Improved Word Representation Learning with Sememes 一句话概括本文工作使用词汇资源--知网--来提升词嵌入的表征能力,并提出了三种基于 ...
NLP问题特征表达基础 - 语言模型（Language Model）发展演化历程讨论
1. NLP问题简介 0x1:NLP问题都包括哪些内涵人们对真实世界的感知被成为感知世界,而人们用语言表达出自己的感知视为文本数据.那么反过来,NLP,或者更精确地表达为文本挖掘,则是从文本数据出发 ...
YOLO 论文阅读
YOLO(You Only Look Once)是一个流行的目标检测方法,和Faster RCNN等state of the art方法比起来,主打检测速度快.截止到目前为止(2017年2月初),YO ...
BERT 论文阅读笔记
BERT 论文阅读 BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding 由 @快刀切草莓君 ...

随机推荐

java里getPath、 getAbsolutePath、getCanonicalPath的区别
本文链接:https://blog.csdn.net/wh_19910525/article/details/9314675 File的这三个方法在api中都有说明,仅以程序为例说明. package ...
MyBatisPLus入门项目实战各教程目录汇总
https://blog.csdn.net/BADAO_LIUMANG_QIZHI/column/info/37194 http://www.imooc.com/article/details/id/ ...
Docker安装redis3.2
1.拉取redis3.2镜像 2.使用docker images查看拉去下来的镜像 3.运行容器,命令如下 docker run -p : -v $PWD/data:/data -d redis:3. ...
Qt for Android （二） Qt打开android的相册
以下有一个可以用的Demo https://files.cnblogs.com/files/wzxNote/Qt-Android-Gallery-master.zip 网盘地址: https://pa ...
Winform 工程反编译后窗体如何显示
Winform反编译后,如果想要让它象正常的工程一样,可以在窗体编辑器中,编辑,需要做一些工作. 1. 转换.resources 为 .resx 利用resgen工具.这个工具是vs自带的. 在启动 ...
layer快速点击会触发多次回调
场景还原测试同学反馈点击了一次操作,为什么会有两条操作记录? 我:???? 排查思路查看日志,看一下是不是发了两次请求,果不其然啊: 并发了,同一时间发送了两次请求,出现了脏写. 原因系统的co ...
go协程理解
一.Golang 线程和协程的区别备注:需要区分进程.线程(内核级线程).协程(用户级线程)三个概念. 进程.线程和协程之间概念的区别对于进程.线程,都是有内核进行调度,有 CPU 时间片 ...
linux简单命令6---挂载
logback 和 log4j对比，及相关配置
Logback 一.logback的介绍 Logback是由log4j创始人设计的又一个开源日志组件.logback当前分成三个模块:logback-core,logback- classic和log ...
prometheus数据上报方式-pushgateway
pushgateway 客户端使用push的方式上报监控数据到pushgateway,prometheus会定期从pushgateway拉取数据.使用它的原因主要是: Prometheus 采用 pu ...

#论文阅读# Universial language model fine-tuing for text classification

#论文阅读# Universial language model fine-tuing for text classification的更多相关文章

随机推荐

热门专题