weka特征选择（IG、chi-square)

一、说明

　　IG是information gain 的缩写，中文名称是信息增益，是选择特征的一个很有效的方法（特别是在使用svm分类时）。这里不做详细介绍，有兴趣的可以googling一下。

　　chi-square 是一个常用特征筛选方法，在种子词扩展那篇文章中，有详细说明，这里不再赘述。

二、weka中的使用方法

　　1、特征筛选代码

package com.lvxinjian.alg.models.feature;

import java.nio.charset.Charset;

import java.util.ArrayList;

import weka.attributeSelection.ASEvaluation;

import weka.attributeSelection.AttributeEvaluator;

import weka.attributeSelection.Ranker;

import weka.core.Instances;

import com.iminer.tool.common.util.FileTool;

/**

 * @Description : 使用Weka的特征筛选方法（目前支持IG、Chi-square）

 *

 */

public class FeatureSelectorByWeka {

    /**

     * @function 使用weka内置的算法筛选特征

     * @param eval 特征筛选方法的对象实例

     * @param data arff格式的数据

     * @param maxNumberOfAttribute 支持的最大的特征个数

     * @param outputPath lex输出文件

     * @throws Exception

     */

    public void EvalueAndRank(ASEvaluation eval , Instances data ,int maxNumberOfAttribute , String outputPath) throws Exception

    {

        Ranker rank = new Ranker();

        eval.buildEvaluator(data);

        rank.search(eval, data);

         // 按照特定搜索算法对属性进行筛选 在这里使用的Ranker算法仅仅是属性按照InfoGain/Chi-square的大小进行排序

        int[] attrIndex = rank.search(eval, data);

         // 打印结果信息 在这里我们了属性的排序结果

        ArrayList<String> attributeWords = new ArrayList<String>();

        for (int i = 0; i < attrIndex.length; i++) {

            //如果权重等于0，则跳出循环

            if (((AttributeEvaluator) eval).evaluateAttribute(attrIndex[i]) == 0)

                break;

            if (i >= maxNumberOfAttribute)

                break;

            attributeWords.add(i + "\t"

                    + data.attribute(attrIndex[i]).name() + "\t" + "1");

        }

        FileTool.SaveListToFile(attributeWords, outputPath, false,

                Charset.forName("utf8"));

    }

}

package com.lvxinjian.alg.models.feature;

import java.io.IOException;

import weka.attributeSelection.ASEvaluation;

import weka.attributeSelection.ChiSquaredAttributeEval;

import weka.attributeSelection.InfoGainAttributeEval;

import weka.core.Instances;

import weka.core.converters.ConverterUtils.DataSource;

import com.iminer.alg.models.generatefile.ParameterUtils;

/**

 * @Description : IG、Chi-square特征筛选

 *

 */

public class WekaFeatureSelector extends FeatureSelector{        

    /**

     * 最大的特征个数

     */

    private int maxFeatureNum = 10000;

    /**

     * 特征文件保存路径

     */

    private String outputPath = null;

    /**

     * @Fields rule 对于特征过滤的规则

     */

    private String classname = "CLASS";

    /**

     * 特征筛选方法，默认为IG

     */

    private String selectMethod = "IG";

    private boolean Initialization(String options){

        try {

            String [] paramArrayOfString = options.split(" ");

            //初始化特征最大个数

            String maxFeatureNum = ParameterUtils.getOption("maxFeatureNum", paramArrayOfString);

            if(maxFeatureNum.length() != 0)

                this.maxFeatureNum = Integer.parseInt(maxFeatureNum);

            //初始化类别

            String classname = ParameterUtils.getOption("class", paramArrayOfString);

            if(classname.length() != 0)

                this.classname = classname;

            else{

                System.out.println("use default class name(\"CLASS\")");

            }

            //初始化特征保存路径

            String outputPath = ParameterUtils.getOption("outputPath", paramArrayOfString);

            if(outputPath.length() != 0)

                this.outputPath = outputPath;

            else{

                System.out.println("please initialze output path.");

                return false;

            }

            String selectMethod = ParameterUtils.getOption("selectMethod", paramArrayOfString);

            if(selectMethod.length() != 0)

                this.selectMethod = selectMethod;

            else{

                System.out.println("use default select method(IG)");

            }

        } catch (Exception e) {

            e.printStackTrace();

            return false;

        }

        return true;

    }

    @Override

    public boolean selectFeature(Object obj ,String options) throws IOException {

        try {

            if(!Initialization(options))

                return false;

            Instances data = (Instances)obj;

            data.setClass(data.attribute(this.classname));

            ASEvaluation selector = null;

            if(this.selectMethod.equals("IG"))

                selector = new InfoGainAttributeEval();

            else if(this.selectMethod.equals("CHI"))

                selector = new ChiSquaredAttributeEval();

            FeatureSelectorByWeka attributeSelector = new FeatureSelectorByWeka();

            attributeSelector.EvalueAndRank(selector, data ,this.maxFeatureNum ,this.outputPath);

        } catch (Exception e) {

            // TODO Auto-generated catch block

            e.printStackTrace();

        }

        return true;

    }

    public static void main(String [] args) throws Exception

    {

        String root = "C:\\Users\\Administrator\\Desktop\\12_05\\模型训练\\1219\\";

        WekaFeatureSelector selector = new WekaFeatureSelector();

        Instances data = DataSource.read(root + "train.Bigram.arff");

        String options = "-maxFeatureNum 10000 -outputPath lex.txt";

        selector.selectFeature(data, options);

    }

}

参考：

weka数据挖掘拾遗（二）---- 特征选择（IG、chi-square)

Weka学习四（属性选择）

weka特征选择（IG、chi-square)的更多相关文章

Chi Square Distance
The chi squared distance d(x,y) is, as you already know, a distance between two histograms x=[x_1,.. ...
特征选择之Chi卡方检验
特征选择之Chi卡方检验卡方值越大,说明对原假设的偏离越大,选择的过程也变成了为每个词计算它与类别Ci的卡方值,从大到小排个序(此时开方值越大越相关),取前k个就可以. 针对英文纯文本的实验结果表明 ...
【Machine Learning】wekaの特征选择简介
看过这篇博客的都应该明白,特征选择代码实现应该包括3个部分: 搜索算法: 评估函数: 数据: 因此,代码的一般形式为: AttributeSelection attsel = new Attribut ...
BendFord's law's Chi square test
http://www.siam.org/students/siuro/vol1issue1/S01009.pdf bendford'law e=log10(1+l/n) o=freq of first ...
文本挖掘之特征选择(python 实现)
机器学习算法的空间.时间复杂度依赖于输入数据的规模,维度规约(Dimensionality reduction)则是一种被用于降低输入数据维数的方法.维度规约可以分为两类: 特征选择(feature ...
使用Python的文本挖掘的特征选择/提取
在文本挖掘与文本分类的有关问题中,文本最初始的数据是将文档表示成向量空间模型的一个矩阵,而这个矩阵所拥有的就是不同的词,常采用特征选择方法.原因是文本的特征一般都是单词(term),具有语义信息,使用 ...
scikit-learn：在实际项目中用到过的知识点（总结）
零.全部项目通用的: http://blog.csdn.net/mmc2015/article/details/46851245(数据集格式和预測器) http://blog.csdn.net/mmc ...
NLP-特征选择
文本分类之特征选择 1 研究背景对于高纬度的分类问题,我们在分类之前一般会进行特征降维,特征降维的技术一般会有特征提取和特征选择.而对于文本分类问题,我们一般使用特征选择方法. 特征提取:PCA.线 ...
用R进行市场调查和消费者感知分析
// // 问题到数据理解问题理解客户的问题:谁是客户(某航空公司)?交流,交流,交流! 问题要具体某航空公司: 乘客体验如何?哪方面需要提高? 类别:比较.描述.聚类,判别还是回归需要什么样 ...

随机推荐

PHP自动解压上传的rar文件
PHP自动解压上传的rar文件浏览:383 发布日期:2015/07/20 分类:功能实现关键字: php函数 php扩展大家都知道php有个zip类可直接操作zip压缩文件,可是用户有时候 ...
xenserver+starwind架构布署
主机 CPU 和主板均需支持 INTER-VT/ AMD-VT ,主板默认可能没开进BISO开启下载最新的 xenserver ,授权注册一下轻松得到授权文件 (鄙视一下VMWARE,看人家 ...
Windows Registry
https://msdn.microsoft.com/en-us/library/windows/desktop/ms724871(v=vs.85).aspx https://msdn.microso ...
communication between threads 线程间通信 Programming Concurrent Activities 程序设计中的并发活动 Ada task 任务 Java thread 线程
Computer Science An Overview _J. Glenn Brookshear _11th Edition activation 激活 parallel processing 并行 ...
prototype linkage can reduce object initialization time and memory consumption
//对象是可变的键控集合, //"numbers, strings, booleans (true and false), null, and undefined" 不是对象的解释 ...
Nginx 禁用IP IP段
最近公司网站被竞争对手用爬虫频繁访问,所以我们这边要禁止这些爬虫访问,我们通过nginx 指令就可以实现了方法一:直接在LB机器上封IP 1.在 blocksip.conf 文件中加入要屏蔽的ip或 ...
C# easyui datagrid 复选框填充。
具体效果如下: 首页
JS性能消耗在哪里？
内部原因:构造,递归,循环,拷贝,动态执行,字符串操作等 1.过度的封装(过多的创建“庞大的”对象,但是如果在允许的条件下,面向对象的封装是可以提高维护性,而且符合我们的高内聚低耦合原则): 2. ...
MongoDB聚合查询
1.count:查询记录条数 db.user.count() 它也跟find一样可以有条件的 db.user.count({}) 2.distinct:用来找出给定键的所有不同的值 db.user.d ...
HTML与CSS的关系
1. HTML是网页内容的载体.内容就是网页制作者放在页面上想要让用户浏览的信息,可以包含文字.图片.视频等. 2. CSS样式是表现.就像网页的外衣.比如,标题字体.颜色变化,或为标题加入背景图片. ...

weka特征选择（IG、chi-square)

weka数据挖掘拾遗（二）---- 特征选择（IG、chi-square)

weka特征选择（IG、chi-square)的更多相关文章

随机推荐

热门专题