MapReduce 中的Map后，sort不能对中文的key排序

今天写了一个用mapreduce求平均分的程序，结果是出来了，可是没有按照“学生名字”进行排序，如果是英文名字的话，结果是排好序的。

代码如下：

package com.pro.bq;

import java.io.IOException;

import java.util.StringTokenizer;

import org.apache.hadoop.conf.Configuration;

import org.apache.hadoop.io.IntWritable;

import org.apache.hadoop.io.Text;

import org.apache.hadoop.mapreduce.Job;

import org.apache.hadoop.mapreduce.Mapper;

import org.apache.hadoop.mapreduce.Reducer;

import org.apache.hadoop.mapreduce.lib.input.FileInputFormat;

import org.apache.hadoop.mapreduce.lib.input.TextInputFormat;

import org.apache.hadoop.mapreduce.lib.output.FileOutputFormat;

import org.apache.hadoop.mapreduce.lib.output.TextOutputFormat;

import org.apache.hadoop.util.GenericOptionsParser;

import org.apache.hadoop.fs.Path;

public class AverageScore {

    public static class MapAvg extends Mapper<Object, Text, Text, IntWritable>

    {

        public void map(Object key, Text value,Context context)

                throws IOException, InterruptedException {  
//            String[] lineData=value.toString().split(" ");//split中间如果有很多“ ”的话lineData的长度增加，灵活性差
//            if(lineData.length==2)
//            {        
//                name.set(lineData[0]);
//                score.set(Integer.parseInt(lineData[1]));
//                context.write(name,score);
//            }

            String line=value.toString();

            StringTokenizer tokenizer=new StringTokenizer(line,"\n");

            while(tokenizer.hasMoreElements())

            {

                StringTokenizer token=new StringTokenizer(tokenizer.nextToken());

                Text name=new Text(token.nextToken());

                IntWritable score=new IntWritable(Integer.parseInt(token.nextToken()));

                context.write(name,score);

            }

        }

    }

    public static class ReduceAvg extends Reducer<Text, IntWritable, Text, IntWritable>

    {

        public void reduce(Text key, Iterable<IntWritable> values,Context context)

                throws IOException, InterruptedException {

            // TODO Auto-generated method stub

            int sum=0;

            int cnt=0;

            for(IntWritable val:values)

            {

                sum+=val.get();

                cnt++;

            }

            sum=(Integer)sum/cnt;

            context.write(key, new IntWritable(sum));

        }

    }

    public static void main(String[] args) throws IOException, ClassNotFoundException, InterruptedException {

        Configuration conf=new Configuration();

        String[] hdfsPath=new String[]{"hdfs://localhost:9000/user/haduser/input/averageTest/","hdfs://localhost:9000/user/haduser/output/outAvgScore/"};

        String[] otherArgs=new GenericOptionsParser(conf, hdfsPath).getRemainingArgs();

        if(otherArgs.length!=2)

        {

            System.err.println("<in> <out>!!");

            System.exit(2);

        }

        Job job=new Job();

        job.setJarByClass(AverageScore.class);

        job.setMapperClass(MapAvg.class);

        job.setReducerClass(ReduceAvg.class);

        job.setOutputKeyClass(Text.class);

        job.setOutputValueClass(IntWritable.class);

        FileInputFormat.addInputPath(job, new Path(otherArgs[0]));

        FileOutputFormat.setOutputPath(job,new Path(otherArgs[1]));

        System.exit(job.waitForCompletion(true)?0:1);

    }

}

file1:

zhangsan

lisi

wangwu

zhaoliu 

file2:

张三

李四

王五

赵六    

file3:

zhangsan

lisi

wangwu

zhaoliu 

file4:

李四

张三

王五

赵六

结果如下：

lisi    38

wangwu    49

zhangsan    27

zhaoliu    60

张三    2

李四    1

王五    2

赵六    3

难道不支持中文的排序？？以后学会自己写Partitioner后是不是可以自己写排序的程序？？以后解决...

MapReduce 中的Map后，sort不能对中文的key排序的更多相关文章

MapReduce中的Shuffle和Sort分析
MapReduce 是现今一个非常流行的分布式计算框架,它被设计用于并行计算海量数据.第一个提出该技术框架的是Google 公司,而Google 的灵感则来自于函数式编程语言,如LISP,Scheme ...
Hadoop : MapReduce中的Shuffle和Sort分析
地址 MapReduce 是现今一个非常流行的分布式计算框架,它被设计用于并行计算海量数据.第一个提出该技术框架的是Google 公司,而Google 的灵感则来自于函数式编程语言,如LISP,Sch ...
MapReduce中的map个数
在map阶段读取数据前,FileInputFormat会将输入文件分割成split.split的个数决定了map的个数.影响map个数(split个数)的主要因素有: 1) 文件的大小.当块(dfs. ...
mapreduce中一个map多个输入路径
package duogemap; import java.io.IOException; import java.util.ArrayList; import java.util.List; imp ...
Hadoop框架下MapReduce中的map个数如何控制
控制map个数的核心源码 long minSize = Math.max(getFormatMinSplitSize(), getMinSplitSize(job)); //getFormatMinS ...
list中依据map<String,Object>的某个值排序
private void sort(List<Map<String, Object>> list) { Collections.sort(list, new Comparato ...
MapReduce中combine、partition、shuffle的作用是什么
http://www.aboutyun.com/thread-8927-1-1.html Mapreduce在hadoop中是一个比較难以的概念.以下须要用心看,然后自己就能总结出来了. 概括: co ...
Java Map 键值对排序按key排序和按Value排序
一.理论准备 Map是键值对的集合接口,它的实现类主要包括:HashMap,TreeMap,Hashtable以及LinkedHashMap等. TreeMap:基于红黑树(Red-Black tre ...
mapreduce 中 map数量与文件大小的关系
学习mapreduce过程中, map第一个阶段是从hdfs 中获取文件的并进行切片,我自己在好奇map的启动的数量和文件的大小有什么关系,进过学习得知map的数量和文件切片的数量有关系,那文件的大小 ...

随机推荐

多线程中，static函数与非static函数的区别？
最近在学习多线程,刚入门,好多东西不懂,下面这段代码今天想了半天也没明白,希望看到的兄弟姐妹能解释下. public class NotThreadSafeCounter extends Thread ...
java中封装
.什么是封装? 封装就是将属性私有化,提供公有的方法访问私有属性. 做法就是:修改属性的可见性来限制对属性的访问,并为每个属性创建一对取值(getter)方法和赋值(setter)方法,用于对这些属性 ...
从InputStream到String_写成函数
String result = readFromInputStream(inputStream);//调用处 //将输入流InputStream变为String public String readF ...
动画(Animation) 之 (闪烁、左右摇摆、上下晃动等效果)
左右晃动的效果: (这边显示没那么流畅) 一.续播 (不知道取什么名字好,就是先播放动画A, 接着播放动画B) 有两种方式. 第一种,分别动画两个动画,A和B, 然后先播放动画A,设置A 的 Ani ...
NodeJS - Express 3.0下ejs模板使用 partial展现片段视图
如果你也在看Node.js开发指南,如果你也在一步一步实现 microBlog 项目!也许你会遇到本文提到的问题,如果你用的是Express 3.0 本书实例背景是 Express 2.0 而如今升级 ...
机器学习(Machine Learning)&深度学习(Deep Learning)资料【转】
转自:机器学习(Machine Learning)&深度学习(Deep Learning)资料 <Brief History of Machine Learning> 介绍:这是一 ...
怎么查看其它apk里面的布局代码及资源
今天才看到的好方法, 将你要的apk文件的后缀名改为zip,解压就可以了. --------------------------------- 提示:有时候系统会自动隐藏你的后缀名的,这时候就需要你将 ...
使用Putty连接VirtualBox的Ubuntu
从vbox中安装了ubuntu server,然后用ssh连过去,发现有一个错误:server unexpectedly closed network connection.猛然发现,ssh没有安装. ...
w3c_html_study_note_5.26
xhtml+css 正确的说法 “DIV+CSS”叫法将网页制作者引入两大误区 [误区一]网页中用了Table,页面就不标准,甚至觉着用Table丢人,Table成为了判定页面是否标准的关键点. [误 ...
ubuntu 安装git
问题描述: ubuntu安装git 问题解决: (1)ubuntu下载git 注: 使用命令apt-get install git安装 (2)查看g ...

MapReduce 中的Map后，sort不能对中文的key排序

MapReduce 中的Map后，sort不能对中文的key排序的更多相关文章

随机推荐

热门专题