mapreduce入门之wordcount注释详解

mapreduce版本：0.2.0之前

说明：　　

　　该注释为之前学习时找到的一篇，现在只是在入门以后对该注释做了一些修正以及添加。

　　由于版本问题，该代码并没有在集群环境中运行，只将其做为理解mapreduce的参考吧。

　　切记，该版本是0.2.0之前的版本，请分辨清楚！

正文：

package org.apache.hadoop.examples;

import java.io.IOException;

import java.util.Iterator;

import java.util.StringTokenizer;

import org.apache.hadoop.fs.Path;

import org.apache.hadoop.io.IntWritable;

import org.apache.hadoop.io.LongWritable;

import org.apache.hadoop.io.Text;

import org.apache.hadoop.mapred.FileInputFormat;

import org.apache.hadoop.mapred.FileOutputFormat;

import org.apache.hadoop.mapred.JobClient;

import org.apache.hadoop.mapred.JobConf;

import org.apache.hadoop.mapred.MapReduceBase;

import org.apache.hadoop.mapred.Mapper;

import org.apache.hadoop.mapred.OutputCollector;

import org.apache.hadoop.mapred.Reducer;

import org.apache.hadoop.mapred.Reporter;

import org.apache.hadoop.mapred.TextInputFormat;

import org.apache.hadoop.mapred.TextOutputFormat;

public class WordCount

{

    //Map类继承自MapReduceBase，并且实现了Mapper接口,此接口是一个规范类型.

    //它有4种形式的参数，分别用来指定map的输入key、value值类型,输出key、value值类型

    public static class Map

    extends MapReduceBase

    implements Mapper<LongWritable, Text, Text, IntWritable>

    {

        private final static IntWritable one = new IntWritable(1);

        private Text word = new Text();

        //实现map方法，对输入值进行处理。（此处用来去掉空格）

        public void map(LongWritable key, Text value,

            OutputCollector<Text, IntWritable> output, Reporter reporter)

            throws IOException

            {

                String line = value.toString();

                StringTokenizer tokenizer = new StringTokenizer(line);

                while (tokenizer.hasMoreTokens())

                {

                    word.set(tokenizer.nextToken());

                    output.collect(word, one);

                }

            }

    }

    /*

    //Reduce类也是继承自MapReduceBase的，需要实现Reducer接口。

    //Reduce类以map的输出作为输入，因此Reduce的输入类型是<Text，Intwritable>。

    //而Reduce的输出是单词和它的数目，因此，它的输出类型是<Text,IntWritable>。

    //Reduce类也要实现reduce方法，在此方法中，reduce函数将输入的key值作为输出的key值，然后将获得多个value值加起来，作为输出的值。

    */

    public static class Reduce

        extends MapReduceBase

        implements Reducer<Text, IntWritable, Text, IntWritable>

    {

        public void reduce(Text key, Iterator<IntWritable> values,

        OutputCollector<Text, IntWritable> output, Reporter reporter)

        throws IOException

        {

            int sum = 0;

            while (values.hasNext())

            {

                sum += values.next().get();

            }

            output.collect(key, new IntWritable(sum));

        }

    }

    public static void main(String[] args) throws Exception

    {

        //1.用JobConf类对 MapReduce job进行初始化

        JobConf conf = new JobConf(WordCount.class);

        //    调用setJobName()方法命名这个Job

        conf.setJobName("wordcount");

        //setup2:设置Job输出结果<key,value>的中key和value数据类型,因为结果是<单词,个数>

        //所以key设置为"Text"类型，相当于Java中String类型。

        conf.setOutputKeyClass(Text.class);

        //Value设置为"IntWritable"，相当于Java中的int类型。

        conf.setOutputValueClass(IntWritable.class);

        //setup3:指定job的MapReduce，以及combiner

        //设置Job处理的Map（拆分）

        conf.setMapperClass(Map.class);

        //设置Job处理的Combiner（中间结果合并，这里用Reduce类来进行Map产生的中间结果合并，避免给网络数据传输产生压力。）

            也可以不用设置（已默认）

        conf.setCombinerClass(Reduce.class);

        //设置Job处理的Reduce（合并）

        conf.setReducerClass(Reduce.class);

        //指定输入输出路径，可在项目上右键->Run As->Run Configuration->arguments->program arguments中配置

            即为main(String[] args)中String[] args赋值

        //指定InputPaths

            eg:hdfs://master:9000/input1/

        FileInputFormat.setInputPaths(conf, new Path(args[0]));

        //指定outputPaths

            eg:hdfs://master:9000/input1/

        FileOutputFormat.setOutputPath(conf, new Path(args[1]));

        JobClient.runJob(conf);

    }

}

mapreduce入门之wordcount注释详解的更多相关文章

JScript中的条件注释详解（转载自网络）
JScript中的条件注释详解-转载这篇文章主要介绍了JScript中的条件注释详解,本文讲解了@cc_on.@if.@set.@_win32.@_win16.@_mac等条件注释语句及可用于条件编 ...
大数据Hadoop核心架构HDFS+MapReduce+Hbase+Hive内部机理详解
微信公众号[程序员江湖] 作者黄小斜,斜杠青年,某985硕士,阿里 Java 研发工程师,于 2018 年秋招拿到 BAT 头条.网易.滴滴等 8 个大厂 offer,目前致力于分享这几年的学习经验. ...
Hadoop核心架构HDFS+MapReduce+Hbase+Hive内部机理详解
转自:http://blog.csdn.net/iamdll/article/details/20998035 分类: 分布式 2014-03-11 10:31 156人阅读评论(0) 收藏举报 ...
Spring 入门 web.xml配置详解
Spring 入门 web.xml配置详解 https://www.cnblogs.com/cczz_11/p/4363314.html https://blog.csdn.net/hellolove ...
爬虫入门之urllib库详解(二)
爬虫入门之urllib库详解(二) 1 urllib模块 urllib模块是一个运用于URL的包 urllib.request用于访问和读取URLS urllib.error包括了所有urllib.r ...
《挑战30天C++入门极限》入门教程：实例详解C++友元
入门教程:实例详解C++友元在说明什么是友元之前,我们先说明一下为什么需要友元与友元的缺点: 通常对于普通函数来说,要访问类的保护成员是不可能的,如果想这么做那么必须把类的成员都生命成为pu ...
MapReduce On Yarn的配置详解和日常维护
MapReduce On Yarn的配置详解和日常维护作者:尹正杰版权声明:原创作品,谢绝转载!否则将追究法律责任. 一.MapReduce运维概述 MapReduce on YARN的运维主要是 ...
Hadoop集群WordCount运行详解（转）
原文链接:Hadoop集群(第6期)_WordCount运行详解 1.MapReduce理论简介 1.1 MapReduce编程模型 MapReduce采用"分而治之"的思想,把对 ...
MapReduce 1工作原理图文详解
MapReduce工作原理图文详解一 MapReduce程序执行流程程序执行流程图如下: 流程分析:1.在客户端启动一个作业.2.向JobTracker请求一个Job ID.3.将运行作业所需要的 ...

随机推荐

shell学习记录002-知识点储备
1.echo "4*0.33" |bc #计算机功能的运用 [root@oc3408554812 shell]# ss=22; [root@oc3408554812 shel ...
使用HTTP访问网络------使用HTTPURLConnection
HTTPURLConnection继承了URLConnection,因此也可用于向指定网站发送GET请求.POST请求.它在URLConnection的基础上提供了如下便捷的方法: 1.int ge ...
[Js]布局转换
为什么要布局转换? 要这样的效果,单写css,只要给每个li浮动就行,不需要绝对定位.但是比如做一些效果(如鼠标移入图片变大),就需要改变位置了.直接给每个li在css上定好位置不方便,也不知道有几个 ...
Oracle中any和all的区别用法
对于any,all的用法,书中说的比较绕口,难以理解,如果通过举例就会比较清晰. any的例子: select * from t_hq_ryxx where gongz > any (selec ...
二模（12） day1
第一题: 题目大意: 求由N个1,M个0组成的排列的个数,要求在排列的任意一个前缀中,1的个数不少于0的个数.N,M<=5000. 解题过程: 1.看到N,M的范围就明确肯定不会是dp,因为起码 ...
Redis系列-存储篇set主要操作函数小结
最近,总是以“太忙“为借口,很久没有blog了,凡事贵在恒,希望我能够坚持不懈,毕竟在blog的时候,也能提升自己.废话不说了,直奔主题”set“ redis set 是string类型对象的无序集合 ...
javascript作用域(Scope),简述上下文（context）和作用域的定义
网页制作Webjx文章简介:这篇文章将正面解决这个问题:简述上下文(context)和作用域的定义,分析可以让我们掌控上下文的两种方法,最后深入一种高效的方案,它能有效解决我所碰到的90%的问题. 作 ...
入門必學NO.1 Android 初學特訓班(第四版) 目錄
第 01 章敲開 Android 的開發大門工欲善其事,必先利其器,要學習 Android 應用程式,先取得功能強大的開發工具,就可讓學習事半功倍. 1.1 Android 是啥米?1.2 建構 A ...
OpenCV函数解读之groupRectangles
不管新版本的CascadeClassifier,还是老版本的HAAR检测函数cvHaarDetectObjects,都使用了groupRectangles函数进行窗口的组合,其函数原型有以下几个: C ...
C# CsvFile 类
using System; using System.Collections.Generic; using System.IO; using System.Linq; using System.Tex ...

mapreduce入门之wordcount注释详解

mapreduce入门之wordcount注释详解的更多相关文章

随机推荐

热门专题