MapReduce编程实例4

MapReduce编程实例：

排序，比较简单，上代码，代码中有注释，欢迎交流。

总体是利用MapReduce本身对Key进行排序的特性和按key值有序的分配到不同的partition。Mapreduce默认会对每个reduce按text类型key按字母顺序排序，对intwritable类型按大小进行排序。

package com.t.hadoop;
import java.io.IOException;
import org.apache.hadoop.conf.Configuration;
import org.apache.hadoop.fs.Path;
import org.apache.hadoop.io.IntWritable;
import org.apache.hadoop.io.Text;
import org.apache.hadoop.mapreduce.Job;
import org.apache.hadoop.mapreduce.Mapper;
import org.apache.hadoop.mapreduce.Partitioner;
import org.apache.hadoop.mapreduce.Reducer;
import org.apache.hadoop.mapreduce.lib.input.FileInputFormat;
import org.apache.hadoop.mapreduce.lib.output.FileOutputFormat;
import org.apache.hadoop.util.GenericOptionsParser;
/**
* 排序
* 利用MapReduce默认的对Key进行排序
* 继承Partitioner类，重写getPartition使Mapper结果整体有序分到相应的Partition，输入到Reduce分别排序。
* 利用全局变量统计位置
* @author daT dev.tao@gmail.com
*
*/
public class Sort {
public static class SortMapper extends Mapper<Object, Text, IntWritable, IntWritable>{
//直接输出key,value，key为需要排序的值，value任意
@Override
protected void map(Object key, Text value,
Context context)throws IOException, InterruptedException {
System.out.println("Key: "+key+" "+"Value: "+value);
context.write(new IntWritable(Integer.valueOf(value.toString())),new IntWritable(1));
}
}
public static class SortReducer extends Reducer<IntWritable, IntWritable, IntWritable, IntWritable>{
public static IntWritable lineNum = new IntWritable(1);//记录该数据的位置
//查询value的个数，有多少个就输出多少个Key值。
@Override
protected void reduce(IntWritable key, Iterable<IntWritable> value,
Context context) throws IOException, InterruptedException {
System.out.println("lineNum: "+lineNum);
for(IntWritable i:value){
context.write(lineNum, key);
}
lineNum = new IntWritable(lineNum.get()+1);
}
}
public static class SortPartitioner extends Partitioner<IntWritable, IntWritable>{
//根据key对数据进行分派
@Override
public int getPartition(IntWritable key, IntWritable value, int partitionNum) {
System.out.println("partitionNum: "+partitionNum);
int maxnum = 23492;//输入的最大值，自己定义的。mapreduce 自带的有采样算法和partition的实现可以用，此例没有用。
int bound = maxnum/partitionNum;
int keyNum = key.get();
for(int i=0;i<partitionNum;i++){
if(keyNum>bound*i&&keyNum<=bound*(i+1)){
return i;
}
}
return -1;
}
}
public static void main(String[] args) throws IOException, ClassNotFoundException, InterruptedException{
Configuration conf = new Configuration();
String[] otherArgs = new GenericOptionsParser(conf, args).getRemainingArgs();
if(otherArgs.length<2){
System.out.println("input parameters errors");
System.exit(2);
}
Job job= new Job(conf);
job.setJarByClass(Sort.class);
job.setMapperClass(SortMapper.class);
job.setPartitionerClass(SortPartitioner.class);//此例不需要combiner，需要设置Partitioner
job.setReducerClass(SortReducer.class);
job.setOutputKeyClass(IntWritable.class);
job.setOutputValueClass(IntWritable.class);
FileInputFormat.addInputPath(job, new Path(otherArgs[0]));
FileOutputFormat.setOutputPath(job, new Path(otherArgs[1]));
System.exit(job.waitForCompletion(true)?0:1);
}
}

MapReduce编程实例4的更多相关文章

MapReduce编程实例6
前提准备: 1.hadoop安装运行正常.Hadoop安装配置请参考:Ubuntu下 Hadoop 1.2.1 配置安装 2.集成开发环境正常.集成开发环境配置请参考 :Ubuntu 搭建Hadoop ...
MapReduce编程实例5
前提准备: 1.hadoop安装运行正常.Hadoop安装配置请参考:Ubuntu下 Hadoop 1.2.1 配置安装 2.集成开发环境正常.集成开发环境配置请参考 :Ubuntu 搭建Hadoop ...
MapReduce编程实例3
MapReduce编程实例: MapReduce编程实例(一),详细介绍在集成环境中运行第一个MapReduce程序 WordCount及代码分析 MapReduce编程实例(二),计算学生平均成绩 ...
MapReduce编程实例2
MapReduce编程实例: MapReduce编程实例(一),详细介绍在集成环境中运行第一个MapReduce程序 WordCount及代码分析 MapReduce编程实例(二),计算学生平均成绩 ...
三、MapReduce编程实例
前文一.CentOS7 hadoop3.3.1安装(单机分布式.伪分布式.分布式二.JAVA API实现HDFS MapReduce编程实例 @ 目录前文 MapReduce编程实例前言注意 ...
hadoop2.2编程：使用MapReduce编程实例（转）
原文链接:http://www.cnblogs.com/xia520pi/archive/2012/06/04/2534533.html 从网上搜到的一篇hadoop的编程实例,对于初学者真是帮助太大 ...
MapReduce编程实例
MapReduce常见编程实例集锦. WordCount单词统计数据去重倒排索引 1. WordCount单词统计 (1) 输入输出输入数据: file1.csv内容 hellod world ...
hadoop之mapreduce编程实例(系统日志初步清洗过滤处理)
刚刚开始接触hadoop的时候,总觉得必须要先安装hadoop集群才能开始学习MR编程,其实并不用这样,当然如果你有条件有机器那最好是自己安装配置一个hadoop集群,这样你会更容易理解其工作原理.我 ...
Hadoop--mapreduce编程实例1
前提准备: 1.hadoop安装运行正常.Hadoop安装配置请参考:Ubuntu下 Hadoop 1.2.1 配置安装 2.集成开发环境正常.集成开发环境配置请参考 :Ubuntu 搭建Hadoop ...

随机推荐

gmock学习01---Linux配置gmock
本文目的本文主要介绍gmock 1.6.0版本在Linux上如何部署和使用. gmock是做什么的? 使用C++手动编写mock对象将会是一件十分耗时,易于出错,枯燥乏味的事情.gmock提供一整套 ...
C# http Post 方法
摘自: http://geekswithblogs.net/rakker/archive/2006/04/21/76044.aspx Http Post in C# Searched out on t ...
Azkaban配置
1,新建azkaban目录,用于安置azkaban程序 2,azkaban web服务器安装解压 azkaban-web-server-2.5.0.tar.gz tar -zvxf azkaban ...
浏览器开发调试工具的秘密 - Secrets of the Browser Developer Tools
来源:GBin1.com 如果你是一个前端开发人员的话,正确的了解和使用浏览器开发工具是一个必须的技能. Secrets of the Browser Developer Tools是一个帮助大家了解 ...
【转】C语言中不同的结构体类型的指针间的强制转换详解
C语言中不同类型的结构体的指针间可以强制转换,很自由,也很危险.只要理解了其内部机制,你会发现C是非常灵活的. 一. 结构体声明如何内存的分布, 结构体指针声明结构体的首地址, 结构体成员声明该成员在 ...
[Unity3D]Unity3D游戏开发之Lua与游戏的不解之缘终结篇：UniLua热更新全然解读
---------------------------------------------------------------------------------------------------- ...
Unity3D Gamecenter 得分上传失败的处理
Unity3DGamecenter 得分上传失败的处理. 常常会有种情况是,玩家在地铁或者飞机上没法通过wifi或者3G连接到服务器,从而会导致上传分数失败解决方法: 这时候需要把得分存在Playe ...
<译>Zookeeper官方文档
apache原文地址:http://zookeeper.apache.org/doc/trunk/zookeeperOver.html ZooKeeper ZooKeeper: A Distribut ...
创建了几个String对象？
String str = "a"; 1个,在常量池中创建了一个字符串对象. String str = new String("a"); 2个,在常量池中创建了一 ...
CentOS7 升级到7.4
2 升级CentOS7.4 自己电脑上的系统还是CentOS7.2,服务器是CentOS7.3, 打算统统升级到最新版升级前查看 > lsb_release -a LSB Version: : ...

MapReduce编程实例4

MapReduce编程实例4的更多相关文章

随机推荐

热门专题