Twenty Newsgroups Classification任务之二seq2sparse

seq2sparse对应于mahout中的org.apache.mahout.vectorizer.SparseVectorsFromSequenceFiles，从昨天跑的算法中的任务监控界面可以看到这一步包含了7个Job信息，分别是：（1）DocumentTokenizer（2）WordCount（3）MakePartialVectors（4）MergePartialVectors（5）VectorTfIdf Document Frequency Count（6）MakePartialVectors（7）MergePartialVectors。打印SparseVectorsFromSequenceFiles的参数帮助信息可以看到如下的信息：

Usage:

 [--minSupport <minSupport> --analyzerName <analyzerName> --chunkSize

<chunkSize> --output <output> --input <input> --minDF <minDF> --maxDFSigma

<maxDFSigma> --maxDFPercent <maxDFPercent> --weight <weight> --norm <norm>

--minLLR <minLLR> --numReducers <numReducers> --maxNGramSize <ngramSize>

--overwrite --help --sequentialAccessVector --namedVector --logNormalize]

Options

  --minSupport (-s) minSupport        (Optional) Minimum Support. Default

                                      Value: 2

  --analyzerName (-a) analyzerName    The class name of the analyzer

  --chunkSize (-chunk) chunkSize      The chunkSize in MegaBytes. 100-10000 MB

  --output (-o) output                The directory pathname for output.

  --input (-i) input                  Path to job input directory.

  --minDF (-md) minDF                 The minimum document frequency.  Default

                                      is 1

  --maxDFSigma (-xs) maxDFSigma       What portion of the tf (tf-idf) vectors

                                      to be used, expressed in times the

                                      standard deviation (sigma) of the

                                      document frequencies of these vectors.

                                      Can be used to remove really high

                                      frequency terms. Expressed as a double

                                      value. Good value to be specified is 3.0.

                                      In case the value is less then 0 no

                                      vectors will be filtered out. Default is

                                      -1.0.  Overrides maxDFPercent

  --maxDFPercent (-x) maxDFPercent    The max percentage of docs for the DF.

                                      Can be used to remove really high

                                      frequency terms. Expressed as an integer

                                      between 0 and 100. Default is 99.  If

                                      maxDFSigma is also set, it will override

                                      this value.

  --weight (-wt) weight               The kind of weight to use. Currently TF

                                      or TFIDF

  --norm (-n) norm                    The norm to use, expressed as either a

                                      float or "INF" if you want to use the

                                      Infinite norm.  Must be greater or equal

                                      to 0.  The default is not to normalize

  --minLLR (-ml) minLLR               (Optional)The minimum Log Likelihood

                                      Ratio(Float)  Default is 1.0

  --numReducers (-nr) numReducers     (Optional) Number of reduce tasks.

                                      Default Value: 1

  --maxNGramSize (-ng) ngramSize      (Optional) The maximum size of ngrams to

                                      create (2 = bigrams, 3 = trigrams, etc)

                                      Default Value:1

  --overwrite (-ow)                   If set, overwrite the output directory

  --help (-h)                         Print out help

  --sequentialAccessVector (-seq)     (Optional) Whether output vectors should

                                      be SequentialAccessVectors. If set true

                                      else false

  --namedVector (-nv)                 (Optional) Whether output vectors should

                                      be NamedVectors. If set true else false

  --logNormalize (-lnorm)             (Optional) Whether output vectors should

                                      be logNormalize. If set true else false

在昨天算法的终端信息中该步骤的调用命令如下：

./bin/mahout seq2sparse -i /home/mahout/mahout-work-mahout/20news-seq -o /home/mahout/mahout-work-mahout/20news-vectors -lnorm -nv -wt tfidf

我们只看对应的参数，首先是-lnorm 对应的解释为输出向量是否要使用log函数进行归一化（设置则为true），-nv解释为输出向量被设置为named 向量，这里的named是啥意思？（暂时不清楚），-wt tfidf解释为使用权重的算法，具体参考 http://zh.wikipedia.org/wiki/TF-IDF 。

第（1）步在SparseVectorsFromSequenceFiles的253行的：

DocumentProcessor.tokenizeDocuments(inputDir, analyzerClass, tokenizedPath, conf);

这里进入可以看到使用的Mapper是：SequenceFileTokenizerMapper，没有使用Reducer。Mapper的代码如下：

protected void map(Text key, Text value, Context context) throws IOException, InterruptedException {

    TokenStream stream = analyzer.reusableTokenStream(key.toString(), new StringReader(value.toString()));

    CharTermAttribute termAtt = stream.addAttribute(CharTermAttribute.class);

    StringTuple document = new StringTuple();

    stream.reset();

    while (stream.incrementToken()) {

      if (termAtt.length() > 0) {

        document.add(new String(termAtt.buffer(), 0, termAtt.length()));

      }

    }

    context.write(key, document);

  }

该Mapper的setup函数主要设置Analyzer的，关于Analyzer的api参考： http://lucene.apache.org/core/3_0_3/api/core/org/apache/lucene/analysis/Analyzer.html ，其中在map中用到的函数为 reusableTokenStream( String fieldName, Reader reader) ：Creates a TokenStream that is allowed to be re-used from the previous time that the same thread called this method.
编写下面的测试程序：

package mahout.fansy.test.bayes;

import java.io.IOException;

import java.io.StringReader;

import org.apache.hadoop.conf.Configuration;

import org.apache.hadoop.io.Text;

import org.apache.lucene.analysis.Analyzer;

import org.apache.lucene.analysis.TokenStream;

import org.apache.lucene.analysis.tokenattributes.CharTermAttribute;

import org.apache.mahout.common.ClassUtils;

import org.apache.mahout.common.StringTuple;

import org.apache.mahout.vectorizer.DefaultAnalyzer;

import org.apache.mahout.vectorizer.DocumentProcessor;

public class TestSequenceFileTokenizerMapper {

	/**

	 * @param args

	 */

	private static Analyzer analyzer = ClassUtils.instantiateAs("org.apache.mahout.vectorizer.DefaultAnalyzer",

Analyzer.class);

	public static void main(String[] args) throws IOException {

		testMap();

	}

	public static void testMap() throws IOException{

		Text key=new Text("4096");

		Text value=new Text("today is also late.what about tomorrow?");

		TokenStream stream = analyzer.reusableTokenStream(key.toString(), new StringReader(value.toString()));

	    CharTermAttribute termAtt = stream.addAttribute(CharTermAttribute.class);

	    StringTuple document = new StringTuple();

	    stream.reset();

	    while (stream.incrementToken()) {

	      if (termAtt.length() > 0) {

	        document.add(new String(termAtt.buffer(), 0, termAtt.length()));

	      }

	    }

	    System.out.println("key:"+key.toString()+",document"+document);

	}

}

得出的结果如下：

key:4096,document[today, also, late.what, about, tomorrow]

其中，TokenStream有一个stopwords属性，值为：[but, be, with, such, then, for, no, will, not, are, and, their, if, this, on, into, a, or, there, in, that, they, was, is, it, an, the, as, at, these, by, to, of]，所以当遇到这些单词的时候就不进行计算了。

额，又太晚了。哎，早困了，刷个牙线。。。

分享，快乐，成长

转载请注明出处：http://blog.csdn.net/fansy1990

Twenty Newsgroups Classification任务之二seq2sparse的更多相关文章

Twenty Newsgroups Classification任务之二seq2sparse（5）
接上篇blog,继续分析.接下来要调用代码如下: // Should document frequency features be processed if (shouldPrune || proce ...
Twenty Newsgroups Classification任务之二seq2sparse（3）
接上篇,如果想对上篇的问题进行测试其实可以简单的编写下面的代码: package mahout.fansy.test.bayes.write; import java.io.IOException; ...
Twenty Newsgroups Classification任务之二seq2sparse（2）
接上篇,SequenceFileTokenizerMapper的输出文件在/home/mahout/mahout-work-mahout0/20news-vectors/tokenized-docum ...
mahout 运行Twenty Newsgroups Classification实例
按照mahout官网https://cwiki.apache.org/confluence/display/MAHOUT/Twenty+Newsgroups的说法,我只用运行一条命令就可以完成这个算法 ...
Twenty Newsgroups Classification实例任务之TrainNaiveBayesJob(一)
接着上篇blog,继续看log里面的信息如下: + echo 'Training Naive Bayes model' Training Naive Bayes model + ./bin/mahou ...
项目笔记《DeepLung:Deep 3D Dual Path Nets for Automated Pulmonary Nodule Detection and Classification》（二）（上）模型设计
我只讲讲检测部分的模型,后面两样性分类的试验我没有做,这篇论文采用了很多肺结节检测论文都采用的u-net结构,准确地说是具有DPN结构的3D版本的u-net,直接上图. DPN是颜水成老师团队的成果, ...
深度学习数据集Deep Learning Datasets
Datasets These datasets can be used for benchmarking deep learning algorithms: Symbolic Music Datase ...
Open Data for Deep Learning
Open Data for Deep Learning Here you’ll find an organized list of interesting, high-quality datasets ...
深度学习课程笔记（二）Classification： Probility Generative Model
深度学习课程笔记(二)Classification: Probility Generative Model 2017.10.05 相关材料来自:http://speech.ee.ntu.edu.tw ...

随机推荐

Ruby学习：类的定义和实例变量
ruby是完全面向对象的,所有的数据都是对象,没有独立在类外的方法,所有的方法都在类中定义的. 一.类的定义语法类的定义以 class 关键字开头,后面跟类名,以 end标识符结尾. 类中的方法以 ...
Poco::TCPServer框架解析
Poco::TCPServer框架解析 POCO C++ Libraries提供一套 C++ 的类库用以开发基于网络的可移植的应用程序,功能涉及线程.文件.流,网络协议包括:HTTP.FTP.SMTP ...
stm32之Systick（系统时钟）
Systick的两大作用: 1.可以产生精确延时: 2.可以提供给操作系统一个单独的心跳(时钟)节拍: 通常实现Delay(N)函数的方法为: for(i=0;i<x;i++) ; 对于STM3 ...
JS - 删除确认
<a href="javascript:if(confirm('确实要删除吗?'))location='<{:U('Admin/Update/deleteuserinfo', a ...
【Oracle】不安装Oracle客户端直接用PL/SQL连接数据库
1.下载 instantclient_11_2.zip PL/SQL2.解压instantclient_11_2.zip到相应文件夹,比如:E:\oracleclient\instantclient_ ...
IAR之工程配置
参考 : IAR的Workspace顶部下拉菜单中Debug和Release http://blog.csdn.net/yanpingsz/article/details/5588525 ++++++ ...
svn笔记
安装部署 1.yum install subversion 2.创建svn版本库目录 mkdir -p /svn 3.创建版本库 svnadmin create /svn/fengchao/ ...
Spring MVC整体处理流程
一.spring整体结构首先俯视一下spring mvc的整体结构二.处理流程 1.请求处理的第一站就是DispatcherServlet.它是整个spring mvc的控制核心.与大多数的jav ...
ASP.NET MVC 5 学习教程：添加查询
原文 ASP.NET MVC 5 学习教程:添加查询起飞网 ASP.NET MVC 5 学习教程目录: 添加控制器添加视图修改视图和布局页控制器传递数据给视图添加模型创建连接字符串通过控 ...
mybatis-redis项目分析
redis作为现在最优秀的key-value数据库,非常适合提供项目的缓存服务.把redis作为mybatis的查询缓存也是很常见的做法.在网上发现N多人是自己做的Cache,其实在mybatis的g ...

Twenty Newsgroups Classification任务之二seq2sparse

Twenty Newsgroups Classification任务之二seq2sparse的更多相关文章

随机推荐

热门专题