怎样通过Java程序提交yarn的mapreduce计算任务

因为项目需求，须要通过Java程序提交Yarn的MapReduce的计算任务。与一般的通过Jar包提交MapReduce任务不同，通过程序提交MapReduce任务须要有点小变动。详见下面代码。

下面为MapReduce主程序，有几点须要提一下：

1、在程序中，我将文件读入格式设定为WholeFileInputFormat，即不正确文件进行切分。

2、为了控制reduce的处理过程。map的输出键的格式为组合键格式。

与常规的<key,value>不同，这里变为了<TextPair,Value>，TextPair的格式为<key1,key2>。

3、为了适应组合键，又一次设定了分组函数。即GroupComparator。分组规则为，仅仅要TextPair中的key1同样（不要求key2同样），则数据被分配到一个reduce容器中。这样，当同样key1的数据进入reduce容器后，key2起到了一个数据标识的作用。

package web.hadoop;

import java.io.IOException;

import org.apache.hadoop.conf.Configuration;

import org.apache.hadoop.fs.Path;

import org.apache.hadoop.io.BytesWritable;

import org.apache.hadoop.io.WritableComparable;

import org.apache.hadoop.io.WritableComparator;

import org.apache.hadoop.mapred.JobClient;

import org.apache.hadoop.mapred.JobConf;

import org.apache.hadoop.mapred.JobStatus;

import org.apache.hadoop.mapreduce.Job;

import org.apache.hadoop.mapreduce.Partitioner;

import org.apache.hadoop.mapreduce.lib.input.FileInputFormat;

import org.apache.hadoop.mapreduce.lib.output.FileOutputFormat;

import org.apache.hadoop.mapreduce.lib.output.NullOutputFormat;

import util.Utils;

public class GEMIMain {

	public GEMIMain(){

		job = null;

	}

	public Job job;

	public static class NamePartitioner extends

			Partitioner<TextPair, BytesWritable> {

		@Override

		public int getPartition(TextPair key, BytesWritable value,

				int numPartitions) {

			return Math.abs(key.getFirst().hashCode() * 127) % numPartitions;

		}

	}

	/**

	 * 分组设置类。仅仅要两个TextPair的第一个key同样。他们就属于同一组。

他们的Value就放到一个Value迭代器中，

	 * 然后进入Reducer的reduce方法中。

	 *

	 * @author hduser

	 *

	 */

	public static class GroupComparator extends WritableComparator {

		public GroupComparator() {

			super(TextPair.class, true);

		}

		@Override

		public int compare(WritableComparable a, WritableComparable b) {

			TextPair t1 = (TextPair) a;

			TextPair t2 = (TextPair) b;

			// 比較同样则返回0，比較不同则返回-1

			return t1.getFirst().compareTo(t2.getFirst()); // 仅仅要是第一个字段同样的就分成为同一组

		}

	}

	public  boolean runJob(String[] args) throws IOException,

			ClassNotFoundException, InterruptedException {

		Configuration conf = new Configuration();

		// 在conf中设置outputath变量，以在reduce函数中能够获取到该參数的值

		conf.set("outputPath", args[args.length - 1].toString());

		//设置HDFS中，每次任务生成产品的质量文件所在目录。args数组的倒数第二个原数为质量文件所在目录

		conf.set("qualityFolder", args[args.length - 2].toString());

		//假设在Server中执行。则须要获取web项目的根路径；假设以java应用方式调试，则读取/opt/hadoop-2.5.0/etc/hadoop/目录下的配置文件

		//MapReduceProgress mprogress = new MapReduceProgress();

		//String rootPath= mprogress.rootPath;

		String rootPath="/opt/hadoop-2.5.0/etc/hadoop/";

		conf.addResource(new Path(rootPath+"yarn-site.xml"));

		conf.addResource(new Path(rootPath+"core-site.xml"));

		conf.addResource(new Path(rootPath+"hdfs-site.xml"));

		conf.addResource(new Path(rootPath+"mapred-site.xml"));

		this.job = new Job(conf);

		job.setJobName("Job name:" + args[0]);

		job.setJarByClass(GEMIMain.class);

		job.setMapperClass(GEMIMapper.class);

		job.setMapOutputKeyClass(TextPair.class);

		job.setMapOutputValueClass(BytesWritable.class);

		// 设置partition

		job.setPartitionerClass(NamePartitioner.class);

		// 在分区之后依照指定的条件分组

		job.setGroupingComparatorClass(GroupComparator.class);

		job.setReducerClass(GEMIReducer.class);

		job.setInputFormatClass(WholeFileInputFormat.class);

		job.setOutputFormatClass(NullOutputFormat.class);

		// job.setOutputKeyClass(NullWritable.class);

		// job.setOutputValueClass(Text.class);

		job.setNumReduceTasks(8);

		// 设置计算输入数据的路径

		for (int i = 1; i < args.length - 2; i++) {

			FileInputFormat.addInputPath(job, new Path(args[i]));

		}

		// args数组的最后一个元素为输出路径

		FileOutputFormat.setOutputPath(job, new Path(args[args.length - 1]));

		boolean flag = job.waitForCompletion(true);

		return flag;

	}

	@SuppressWarnings("static-access")

	public static void main(String[] args) throws ClassNotFoundException,

			IOException, InterruptedException {	

		String[] inputPaths = new String[] { "normalizeJob",

				"hdfs://192.168.168.101:9000/user/hduser/red1/",

				"hdfs://192.168.168.101:9000/user/hduser/nir1/","quality11111",

				"hdfs://192.168.168.101:9000/user/hduser/test" };

		GEMIMain test = new GEMIMain();

		boolean result = test.runJob(inputPaths);

	}

}

下面为TextPair类

public class TextPair implements WritableComparable<TextPair> {

	private Text first;

	private Text second;

	public TextPair() {

		set(new Text(), new Text());

	}

	public TextPair(String first, String second) {

		set(new Text(first), new Text(second));

	}

	public TextPair(Text first, Text second) {

		set(first, second);

	}

	public void set(Text first, Text second) {

		this.first = first;

		this.second = second;

	}

	public Text getFirst() {

		return first;

	}

	public Text getSecond() {

		return second;

	}

	@Override

	public void write(DataOutput out) throws IOException {

		first.write(out);

		second.write(out);

	}

	@Override

	public void readFields(DataInput in) throws IOException {

		first.readFields(in);

		second.readFields(in);

	}

	@Override

	public int hashCode() {

		return first.hashCode() * 163 + second.hashCode();

	}

	@Override

	public boolean equals(Object o) {

		if (o instanceof TextPair) {

			TextPair tp = (TextPair) o;

			return first.equals(tp.first) && second.equals(tp.second);

		}

		return false;

	}

	@Override

	public String toString() {

		return first + "\t" + second;

	}

	@Override

	/**A.compareTo(B)

	 * 假设比較同样，则比較结果为0

	 * 假设A大于B，则比較结果为1

	 * 假设A小于B。则比較结果为-1

	 *

	 */

	public int compareTo(TextPair tp) {

		int cmp = first.compareTo(tp.first);

		if (cmp != 0) {

			return cmp;

		}

		//此时实现的是升序排列

		return second.compareTo(tp.second);

	}

}

下面为WholeFileInputFormat，其控制数据在mapreduce过程中不被切分

package web.hadoop;

import java.io.IOException;  

import org.apache.hadoop.fs.Path;

import org.apache.hadoop.io.BytesWritable;

import org.apache.hadoop.io.Text;

import org.apache.hadoop.mapreduce.InputSplit;

import org.apache.hadoop.mapreduce.JobContext;

import org.apache.hadoop.mapreduce.RecordReader;

import org.apache.hadoop.mapreduce.TaskAttemptContext;

import org.apache.hadoop.mapreduce.lib.input.FileInputFormat;

public class WholeFileInputFormat extends FileInputFormat<Text, BytesWritable> {  

    @Override

    public RecordReader<Text, BytesWritable> createRecordReader(

            InputSplit arg0, TaskAttemptContext arg1) throws IOException,

            InterruptedException {

        // TODO Auto-generated method stub

        return new WholeFileRecordReader();

    }  

    @Override

    protected boolean isSplitable(JobContext context, Path filename) {

        // TODO Auto-generated method stub

        return false;

    }

}

下面为WholeFileRecordReader类

package web.hadoop;

import java.io.IOException;

import org.apache.hadoop.conf.Configuration;

import org.apache.hadoop.fs.FSDataInputStream;

import org.apache.hadoop.fs.FileSystem;

import org.apache.hadoop.fs.Path;

import org.apache.hadoop.io.BytesWritable;

import org.apache.hadoop.io.IOUtils;

import org.apache.hadoop.io.Text;

import org.apache.hadoop.mapreduce.InputSplit;

import org.apache.hadoop.mapreduce.RecordReader;

import org.apache.hadoop.mapreduce.TaskAttemptContext;

import org.apache.hadoop.mapreduce.lib.input.FileSplit;

public class WholeFileRecordReader extends RecordReader<Text, BytesWritable> {

	private FileSplit fileSplit;

	private FSDataInputStream fis;

	private Text key = null;

	private BytesWritable value = null;

	private boolean processed = false;

	@Override

	public void close() throws IOException {

		// TODO Auto-generated method stub

		// fis.close();

	}

	@Override

	public Text getCurrentKey() throws IOException, InterruptedException {

		// TODO Auto-generated method stub

		return this.key;

	}

	@Override

	public BytesWritable getCurrentValue() throws IOException,

			InterruptedException {

		// TODO Auto-generated method stub

		return this.value;

	}

	@Override

	public void initialize(InputSplit inputSplit, TaskAttemptContext tacontext)

			throws IOException, InterruptedException {

		fileSplit = (FileSplit) inputSplit;

		Configuration job = tacontext.getConfiguration();

		Path file = fileSplit.getPath();

		FileSystem fs = file.getFileSystem(job);

		fis = fs.open(file);

	}

	@Override

	public boolean nextKeyValue() {

		if (key == null) {

			key = new Text();

		}

		if (value == null) {

			value = new BytesWritable();

		}

		if (!processed) {

			byte[] content = new byte[(int) fileSplit.getLength()];

			Path file = fileSplit.getPath();

			System.out.println(file.getName());

			key.set(file.getName());

			try {

				IOUtils.readFully(fis, content, 0, content.length);

				// value.set(content, 0, content.length);

				value.set(new BytesWritable(content));

			} catch (IOException e) {

				// TODO Auto-generated catch block

				e.printStackTrace();

			} finally {

				IOUtils.closeStream(fis);

			}

			processed = true;

			return true;

		}

		return false;

	}

	@Override

	public float getProgress() throws IOException, InterruptedException {

		// TODO Auto-generated method stub

		return processed ? fileSplit.getLength() : 0;

	}

}

怎样通过Java程序提交yarn的mapreduce计算任务的更多相关文章

java程序中线程cpu使用率计算
原文地址:https://www.imooc.com/article/27374 最近确实遇到题目上的刚需,也是花了一段时间来思考这个问题. cpu使用率如何计算计算使用率在上学那会就经常算,不过往 ...
本地idea开发mapreduce程序提交到远程hadoop集群执行
https://www.codetd.com/article/664330 https://blog.csdn.net/dream_an/article/details/84342770 通过idea ...
YARN 中的应用程序提交
YARN 中的应用程序提交本节讨论在应用程序提交到 YARN 集群时,ResourceManager.ApplicationMaster.NodeManagers 和容器如何相互交互.下图显示了一个 ...
Java --本地提交MapReduce作业至集群☞实现 Word Count
还是那句话,看别人写的的总是觉得心累,代码一贴,一打包,扔到Hadoop上跑一遍就完事了????写个测试样例程序(MapReduce中的Hello World)还要这么麻烦!!!?,还本地打Jar包, ...
Spark On Yarn：提交Spark应用程序到Yarn
转载自:http://lxw1234.com/archives/2015/07/416.htm 关键字:Spark On Yarn.Spark Yarn Cluster.Spark Yarn Clie ...
PTA中提交Java程序的一些套路
201708新版改版说明 PTA与2017年8月已升级成新版,域名改为https://pintia.cn/,官方建议使用Firefox与Chrome浏览器. 旧版 PTA 用户首次在新版系统登录时,请 ...
将java开发的wordcount程序提交到spark集群上运行
今天来分享下将java开发的wordcount程序提交到spark集群上运行的步骤. 第一个步骤之前,先上传文本文件,spark.txt,然用命令hadoop fs -put spark.txt /s ...
经典MapReduce作业和Yarn上MapReduce作业运行机制
一.经典MapReduce的作业运行机制如下图是经典MapReduce作业的工作原理: 1.1 经典MapReduce作业的实体经典MapReduce作业运行过程包含的实体: 客户端,提交MapR ...
Spark集群模式&Spark程序提交
Spark集群模式&Spark程序提交 1. 集群管理器 Spark当前支持三种集群管理方式 Standalone-Spark自带的一种集群管理方式,易于构建集群. Apache Mesos- ...

随机推荐

Eralng 小知识点
文件属性提取方法:Module:module_info/1 头文件包含头文件 -include(FileName). %% FileName为绝对路径或相对路径引入库中包含文件 -include ...
JavaScript面向对象(01)--函数
在JavaScript中,函数和对象有区别,也有联系, 首先函数是一个对象,但是和对象存在一些区别如下: 1,不论在java还是js中,如果把一个对象赋值给另一个变量,那么,后者会指向前者对象所在的内 ...
air手势代码
//下列2句谁放上面谁生效要么触控生效,要么手势生效 Multitouch.inputMode = MultitouchInputMode.TOUCH_POINT; Multitouch.inputM ...
air for ios
在 Adobe AIR 中为不同屏幕尺寸的多种设备提供支持使用Flash Builder 4.5进行多平台游戏开发手机屏幕触控技术与提升AIR在Android上的触控体验 AIR Native E ...
RHAS Linux下架构Lotus Domino详解（附视频）
此处下载操作视频:RHAS Linux下架构Lotus Domino 6.5视频教程在rhas下架构Lotus Domino 汉化 650) this.width=650;" o ...
kali系统安装图文教程
工具和原料 1.虚拟机:Oracle VM VirtualBox 下载地址:https://www.virtualbox.org/wiki/Downloads 根据你自己的计算机操作系统下载,其中如果 ...
Codeforces 687C. The Values You Can Make (dp)
题目链接:http://codeforces.com/problemset/problem/687/C 题目大概说给n个各有价值的硬币,要从它们中选出若干个组合成面值k,而要求的是各个方案里这些选出的 ...
URAL 2069 Hard Rock (最短路)
题意:给定 n + m 个街道,问你从左上角走到右下角的所有路的权值最小的中的最大的. 析:我们只要考虑几种情况就好了,先走行再走列和先走列再走行差不多.要么是先横着,再竖着,要么是先横再竖再横,要么 ...
汇编语言程序入门实验二：在dos下建立子目录操作
汇编语言程序入门实验二:在dos下建立子目录操作 1,背景在读此文,并读懂前,建议读者先阅读这两篇博客 1,在dos环境下汇编语言程序设计入门(输出hello world)和masm32的下载.安装 ...
UVa 10004：Bicoloring
这道题要我们判断所给图是否可以用两种颜色进行染色,即"二染色“.已知所给图一定是强连通图. 分析之: 若图中无回路,则该图是一棵树,一定可以二染色. 若图中有回路,但回路有偶数个节点,仍然可 ...

怎样通过Java程序提交yarn的mapreduce计算任务

怎样通过Java程序提交yarn的mapreduce计算任务的更多相关文章

随机推荐

热门专题