reduce的数目到底和哪些因素有关

1、我们知道map的数量和文件数、文件大小、块大小、以及split大小有关,而reduce的数量跟哪些因素有关呢?

设置mapred.tasktracker.reduce.tasks.maximum的大小可以决定单个tasktracker一次性启动reduce的数目,但是不能决定总的reduce数目。

conf.setNumReduceTasks(4);JobConf对象的这个方法可以用来设定总的reduce的数目,看下Job Counters的统计:

  1. Job Counters
  2. Data-local map tasks=2
  3. Total time spent by all maps waiting after reserving slots (ms)=0
  4. Total time spent by all reduces waiting after reserving slots (ms)=0
  5. SLOTS_MILLIS_MAPS=10695
  6. SLOTS_MILLIS_REDUCES=29502
  7. Launched map tasks=2
  8. Launched reduce tasks=4

确实启动了4个reduce:看下输出:

  1. diegoball@diegoball:~/IdeaProjects/test/build/classes$ hadoop fs -ls  /user/diegoball/join_ou1123
  2. 11/03/25 15:28:45 INFO security.Groups: Group mapping impl=org.apache.hadoop.security.ShellBasedUnixGroupsMapping; cacheTimeout=300000
  3. 11/03/25 15:28:45 WARN conf.Configuration: mapred.task.id is deprecated. Instead, use mapreduce.task.attempt.id
  4. Found 5 items
  5. -rw-r--r--   1 diegoball supergroup          0 2011-03-25 15:28 /user/diegoball/join_ou1123/_SUCCESS
  6. -rw-r--r--   1 diegoball supergroup        124 2011-03-25 15:27 /user/diegoball/join_ou1123/part-00000
  7. -rw-r--r--   1 diegoball supergroup          0 2011-03-25 15:27 /user/diegoball/join_ou1123/part-00001
  8. -rw-r--r--   1 diegoball supergroup        214 2011-03-25 15:28 /user/diegoball/join_ou1123/part-00002
  9. -rw-r--r--   1 diegoball supergroup          0 2011-03-25 15:28 /user/diegoball/join_ou1123/part-00003

只有2个reduce在干活。为什么呢?

shuffle的过程,需要根据key的值决定将这条<K,V> (map的输出),送到哪一个reduce中去。送到哪一个reduce中去靠调用默认的org.apache.hadoop.mapred.lib.HashPartitioner的getPartition()方法来实现。
HashPartitioner类:

  1. package org.apache.hadoop.mapred.lib;
  2. import org.apache.hadoop.classification.InterfaceAudience;
  3. import org.apache.hadoop.classification.InterfaceStability;
  4. import org.apache.hadoop.mapred.Partitioner;
  5. import org.apache.hadoop.mapred.JobConf;
  6. /** Partition keys by their {@link Object#hashCode()}.
  7. */
  8. @InterfaceAudience.Public
  9. @InterfaceStability.Stable
  10. public class HashPartitioner<K2, V2> implements Partitioner<K2, V2> {
  11. public void configure(JobConf job) {}
  12. /** Use {@link Object#hashCode()} to partition. */
  13. public int getPartition(K2 key, V2 value,
  14. int numReduceTasks) {
  15. return (key.hashCode() & Integer.MAX_VALUE) % numReduceTasks;
  16. }
  17. }

numReduceTasks的值在JobConf中可以设置。默认的是1:显然太小。
   这也是为什么默认的设置中总启动一个reduce的原因。

返回与运算的结果和numReduceTasks求余。

Mapreduce根据这个返回结果决定将这条<K,V>,送到哪一个reduce中去。

key传入的是LongWritable类型,看下这个LongWritable类的hashcode()方法:

  1. public int hashCode() {
  2. return (int)value;
  3. }

简简单单的返回了原值的整型值。

因为getPartition(K2 key, V2 value,int numReduceTask)返回的结果只有2个不同的值,所以最终只有2个reduce在干活。

HashPartitioner是默认的partition类,我们也可以自定义partition类 :

  1. package com.alipay.dw.test;
  2. import org.apache.hadoop.io.IntWritable;
  3. import org.apache.hadoop.mapred.JobConf;
  4. import org.apache.hadoop.mapred.Partitioner;
  5. /**
  6. * Created by IntelliJ IDEA.
  7. * User: diegoball
  8. * Date: 11-3-10
  9. * Time: 下午5:26
  10. * To change this template use File | Settings | File Templates.
  11. */
  12. public class MyPartitioner implements Partitioner<IntWritable, IntWritable> {
  13. public int getPartition(IntWritable key, IntWritable value, int numPartitions) {
  14. /* Pretty ugly hard coded partitioning function. Don't do that in practice, it is just for the sake of understanding. */
  15. int nbOccurences = key.get();
  16. if (nbOccurences > 20051210)
  17. return 0;
  18. else
  19. return 1;
  20. }
  21. public void configure(JobConf arg0) {
  22. }
  23. }

仅仅需要覆盖getPartition()方法就OK。通过:
conf.setPartitionerClass(MyPartitioner.class);
可以设置自定义的partition类。
同样由于之返回2个不同的值0,1,不管conf.setNumReduceTasks(4);设置多少个reduce,也同样只会有2个reduce在干活。

由于每个reduce的输出key都是经过排序的,上述自定义的Partitioner还可以达到排序结果集的目的:

  1. 11/03/25 15:24:49 WARN conf.Configuration: mapred.task.id is deprecated. Instead, use mapreduce.task.attempt.id
  2. Found 5 items
  3. -rw-r--r--   1 diegoball supergroup          0 2011-03-25 15:23 /user/diegoball/opt.del/_SUCCESS
  4. -rw-r--r--   1 diegoball supergroup      24546 2011-03-25 15:23 /user/diegoball/opt.del/part-00000
  5. -rw-r--r--   1 diegoball supergroup      10241 2011-03-25 15:23 /user/diegoball/opt.del/part-00001
  6. -rw-r--r--   1 diegoball supergroup          0 2011-03-25 15:23 /user/diegoball/opt.del/part-00002
  7. -rw-r--r--   1 diegoball supergroup          0 2011-03-25 15:23 /user/diegoball/opt.del/part-00003

part-00000和part-00001是这2个reduce的输出,由于使用了自定义的MyPartitioner,所有key小于20051210的的<K,V>都会放到第一个reduce中处理,key大于20051210就会被放到第二个reduce中处理。
每个reduce的输出key又是经过key排序的,所以最终的结果集降序排列。

但是如果使用上面自定义的partition类,又conf.setNumReduceTasks(1)的话,会怎样? 看下Job Counters:

  1. Job Counters
  2. Data-local map tasks=2
  3. Total time spent by all maps waiting after reserving slots (ms)=0
  4. Total time spent by all reduces waiting after reserving slots (ms)=0
  5. SLOTS_MILLIS_MAPS=16395
  6. SLOTS_MILLIS_REDUCES=3512
  7. Launched map tasks=2
  8. Launched reduce tasks=1

只启动了一个reduce。
  (1)、 当setNumReduceTasks( int a) a=1(即默认值),不管Partitioner返回不同值的个数b为多少,只启动1个reduce,这种情况下自定义的Partitioner类没有起到任何作用。
  (2)、 若a!=1:
   a、当setNumReduceTasks( int a)里 a设置小于Partitioner返回不同值的个数b的话:

  1. public int getPartition(IntWritable key, IntWritable value, int numPartitions) {
  2. /* Pretty ugly hard coded partitioning function. Don't do that in practice, it is just for the sake of understanding. */
  3. int nbOccurences = key.get();
  4. if (nbOccurences < 20051210)
  5. return 0;
  6. if (nbOccurences >= 20051210 && nbOccurences < 20061210)
  7. return 1;
  8. if (nbOccurences >= 20061210 && nbOccurences < 20081210)
  9. return 2;
  10. else
  11. return 3;
  12. }

同时设置setNumReduceTasks( 2)。

于是抛出异常:

  1. 11/03/25 17:03:41 INFO mapreduce.Job: Task Id : attempt_201103241018_0023_m_000000_1, Status : FAILED
  2. ava.io.IOException: Illegal partition for 20110116 (3)
  3. at org.apache.hadoop.mapred.MapTask$MapOutputBuffer.collect(MapTask.java:900)
  4. at org.apache.hadoop.mapred.MapTask$OldOutputCollector.collect(MapTask.java:508)
  5. at com.alipay.dw.test.KpiMapper.map(Unknown Source)
  6. at com.alipay.dw.test.KpiMapper.map(Unknown Source)
  7. at org.apache.hadoop.mapred.MapRunner.run(MapRunner.java:54)
  8. at org.apache.hadoop.mapred.MapTask.runOldMapper(MapTask.java:397)
  9. at org.apache.hadoop.mapred.MapTask.run(MapTask.java:330)
  10. at org.apache.hadoop.mapred.Child$4.run(Child.java:217)
  11. at java.security.AccessController.doPrivileged(Native Method)
  12. at javax.security.auth.Subject.doAs(Subject.java:396)
  13. at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:742)
  14. at org.apache.hadoop.mapred.Child.main(Child.java:211)

某些key没有找到所对应的reduce去处。原因是只启动了a个reduce。
 
   b、当setNumReduceTasks( int a)里 a设置大于Partitioner返回不同值的个数b的话,同样会启动a个reduce,但是只有b个redurce上会得到数据。启动的其他的a-b个reduce浪费了。

c、理想状况是a=b,这样可以合理利用资源,负载更均衡。

reduce的数目到底和哪些因素有关的更多相关文章

  1. 019_Map Task数目的确定和Reduce Task数目的指定

    注意标题:Map Task数目的确定和Reduce Task数目的指定————自然得到结论,前者是后者决定的,后者是人为指定的.查看源码可以很容易看懂 1.MapReduce作业中Map Task数目 ...

  2. reduce个数问题

    reduce的数目到底和哪些因素有关 1.我们知道map的数量和文件数.文件大小.块大小.以及split大小有关,而reduce的数量跟哪些因素有关呢?  设置mapred.tasktracker.r ...

  3. reduce个数究竟和哪些因素有关

    reduce的数目究竟和哪些因素有关 1.我们知道map的数量和文件数.文件大小.块大小.以及split大小有关,而reduce的数量跟哪些因素有关呢?  设置mapred.tasktracker.r ...

  4. MapReduce剖析笔记之五:Map与Reduce任务分配过程

    在上一节分析了TaskTracker和JobTracker之间通过周期的心跳消息获取任务分配结果的过程.中间留了一个问题,就是任务到底是怎么分配的.任务的分配自然是由JobTracker做出来的,具体 ...

  5. Hadoop Map/Reduce教程

    原文地址:http://hadoop.apache.org/docs/r1.0.4/cn/mapred_tutorial.html 目的 先决条件 概述 输入与输出 例子:WordCount v1.0 ...

  6. 一步一步跟我学习hadoop(5)----hadoop Map/Reduce教程(2)

    Map/Reduce用户界面 本节为用户採用框架要面对的各个环节提供了具体的描写叙述,旨在与帮助用户对实现.配置和调优进行具体的设置.然而,开发时候还是要相应着API进行相关操作. 首先我们须要了解M ...

  7. MapReduce流程、如何统计任务数目以及Partitioner

    核心功能描述 应用程序通常会通过提供map和reduce来实现 Mapper和Reducer接口,它们组成作业的核心. Map是一类将输入记录集转换为中间格式记录集的独立任务. 这种转换的中间格式记录 ...

  8. 分布式基础学习(2)分布式计算系统(Map/Reduce)

    二. 分布式计算(Map/Reduce) 分 布式式计算,同样是一个宽泛的概念,在这里,它狭义的指代,按Google Map/Reduce框架所设计的分布式框架.在Hadoop中,分布式文件 系统,很 ...

  9. hadoop学习WordCount+Block+Split+Shuffle+Map+Reduce技术详解

    转自:http://blog.csdn.net/yczws1/article/details/21899007 纯干货:通过WourdCount程序示例:详细讲解MapReduce之Block+Spl ...

随机推荐

  1. 通过UserAgent判断智能手机(设备,Android,IOS)

    转:http://free0007.iteye.com/blog/2017329 /// 根据 Agent 判断是否是智能手机 ///</summary> ///<returns&g ...

  2. opencv矩阵总结

    OpenCV 矩阵操作 CvMat 转自:http://hi.baidu.com/xiaoduo170/blog/item/10fe5e3f0fd252e455e72380.html 每回用矩阵都要查 ...

  3. What is Proguard?

    When packaging an apk, all classes of all libraries used by the program will be included, this makes ...

  4. stl的仿函数adapter

    Stl的一点思考 编程语言是为编译器写一份策略,如果将这份策略模板化那就是泛型编程了 bind1st bind2nd not1 not2 adapter并不改变仿函数接口,只是将参数引入其他的运算流程

  5. CentOS6.4-RMAN定时任务备份 on 11GR2

    1.rman备份脚本位置: /home/oracle ./scripts/ ./bin                    -----存放rman脚本 ./log                   ...

  6. StackTrace,Trim

    一: Environment.StackTrace 可能我们看到最多的就是catch中的e参数,里面会有一个StackTrace,然后不可否认的这玩意太有用了,它会把调用堆栈 中的信息输出出来,有了它 ...

  7. SharePoint入门识记

    SharePoint站点层次结构: 1.Web Application: 一般创建后对应一个IIS Web Site, 默认创建后是打不开的,因为网站没有任何内容. 2.Site Collection ...

  8. MongoDB管理与开发精要 书摘

    摘自:<MongoDB管理与开发精要>         性能优化 创建索引 限定返回结果条数 只查询使用到的字段,而不查询所有字段 采用capped collection 采用Server ...

  9. 谷歌 analytics.js 部分解密版

    源:http://www.google-analytics.com/analytics.js (function(){var aa=encodeURIComponent,f=window,ba=set ...

  10. SmartDo数据挖掘思路

    SmartDo数据挖掘思路 数据挖掘部分: 数据挖掘的主要网址为: https://www.amazon.com/Best-Sellers/zgbs 挖掘部分为网址左边的入口,大约20多个,其中页面分 ...