【原创】MapReduce编程系列之表连接

问题描述

需要连接的表如下：其中左边是child，右边是parent，我们要做的是找出grandchild和grandparent的对应关系，为此需要进行表的连接。

Tom Lucy

Tom Jim

Lucy David

Lucy Lili

Jim Lilei

Jim SuSan

Lily Green

Lily Bians

Green Well

Green MillShell

Havid James

James LiT

Richard Cheng

Cheng LiHua

思路分析

　　诚然，在写MR程序的时候要结合MR数据处理的一些特性。例如如果我们用默认的TextInputFormat来处理传入的文件数据，传入的格式是key为行号，value为这一行的值（如上例中的第一行，key为0，value为[Tom,Lucy]），在shuffle过程中，我们的值如果有相同的key，会merge到一起（这一点很重要！）。我们利用shuffle阶段的特性，merge到一组的数据够成一组关系，然后我们在这组关系中想办法区分晚辈和长辈，最后对merge里的value一一作处理，分离出grandchild和grandparent的关系。

例如，Tom Lucy传入处理后我们将其反转，成为Lucy Tom输出。当然，输出的时候，为了达到join的效果，我们要输出两份，因为join要两个表，一个表为L1：child parent，一个表为L2：child parent，为了达到关联的目的和利用shuffle阶段的特性，我们需要将L1反转，把parent放在前面，这样L1表中的parent和L2表中的child如果字段是相同的那么在shuffle阶段就能merge到一起。还有，为了区分merge到一起后如何区分child和parent，我们把L1表中反转后的child（未来的 grandchild）字段后面加一个1，L2表中parent（未来的grandparent）字段后加2。

 package com.test.join;

 import java.io.IOException;

 import java.util.ArrayList;

 import java.util.Iterator;

 import org.apache.hadoop.conf.Configuration;

 import org.apache.hadoop.fs.Path;

 import org.apache.hadoop.io.Text;

 import org.apache.hadoop.mapreduce.Job;

 import org.apache.hadoop.mapreduce.Mapper;

 import org.apache.hadoop.mapreduce.Reducer;

 import org.apache.hadoop.mapreduce.lib.input.FileInputFormat;

 import org.apache.hadoop.mapreduce.lib.output.FileOutputFormat;

 public class STJoin {

     public static class STJoinMapper extends Mapper<Object, Text, Text, Text>{

         @Override

         protected void map(Object key, Text value, Context context)

                 throws IOException, InterruptedException {

             // TODO Auto-generated method stub

             String[] rela = value.toString().trim().split(" ",2);

             if(rela.length!=2)

                 return;

             String child = rela[0];

             String parent = rela[1];

             context.write(new Text(parent), new Text((child+"1")));

             context.write(new Text(child), new Text((parent+"2")));

         }

     }

     public static class STJoinReducer extends Reducer<Text, Text, Text, Text>{

         @Override

         protected void reduce(Text arg0, Iterable<Text> arg1,Context context)

                 throws IOException, InterruptedException {

             // TODO Auto-generated method stub

             ArrayList<String> grandParent = new ArrayList<>();

             ArrayList<String> grandChild = new ArrayList<>();

             Iterator<Text> iterator = arg1.iterator();

             while(iterator.hasNext()){

                 String text = iterator.next().toString();

                 if(text.endsWith("1"))

                     grandChild.add(text.substring(0, text.length()-1));

                 if(text.endsWith("2"))

                     grandParent.add(text.substring(0, text.length()-1));

             }

             for(String grandparent:grandParent){

                 for(String grandchild:grandChild){

                     context.write(new Text(grandchild), new Text(grandparent));

                 }

             }

         }

     }

     public static void main(String args[]) throws IOException, ClassNotFoundException, InterruptedException {

         Configuration conf = new Configuration();

         Job job = new Job(conf,"STJoin");

         job.setMapperClass(STJoinMapper.class);

         job.setReducerClass(STJoinReducer.class);

         job.setOutputKeyClass(Text.class);

         job.setOutputValueClass(Text.class);

         FileInputFormat.addInputPath(job, new Path("hdfs://localhost:9000/user/hadoop/STJoin/joinFile"));

         FileOutputFormat.setOutputPath(job, new Path("hdfs://localhost:9000/user/hadoop/STJoin/joinResult"));

         System.exit(job.waitForCompletion(true)?0:1);

     }

 }

结果显示

Richard    LiHua

Lily    Well

Lily    MillShell

Havid    LiT

Tom    Lilei

Tom    SuSan

Tom    Lili

Tom    David

以上代码在hadoop1.0.3平台实现

【原创】MapReduce编程系列之表连接的更多相关文章

Hadoop阅读笔记（三）——深入MapReduce排序和单表连接
继上篇了解了使用MapReduce计算平均数以及去重后,我们再来一探MapReduce在排序以及单表关联上的处理方法.在MapReduce系列的第一篇就有说过,MapReduce不仅是一种分布式的计算 ...
【SqlServer系列】表连接
1 概述 1.1 已发布[SqlServer系列]文章 [SqlServer系列]MYSQL安装教程 [SqlServer系列]数据库三大范式 [SqlServer系列]表单查询 1.2 本篇 ...
MapReduce编程系列 — 5：单表关联
1.项目名称: 2.项目数据: chile parentTom LucyTom JackJone LucyJone JackLucy MaryLucy Ben ...
【原创】MapReduce编程系列之二元排序
普通排序实现普通排序的实现利用了按姓名的排序,调用了默认的对key的HashPartition函数来实现数据的分组.partition操作之后写入磁盘时会对数据进行排序操作(对一个分区内的数据作排序 ...
MapReduce编程系列 — 6：多表关联
1.项目名称: 2.程序代码: 版本一(详细版): package com.mtjoin; import java.io.IOException; import java.util.Iterator; ...
MapReduce编程系列 — 4：排序
1.项目名称: 2.程序代码: package com.sort; import java.io.IOException; import org.apache.hadoop.conf.Configur ...
MapReduce编程系列 — 3：数据去重
1.项目名称: 2.程序代码: package com.dedup; import java.io.IOException; import org.apache.hadoop.conf.Configu ...
MapReduce编程系列 — 2：计算平均分
1.项目名称: 2.程序代码: package com.averagescorecount; import java.io.IOException; import java.util.Iterator ...
MapReduce编程系列 — 1：计算单词
1.代码: package com.mrdemo; import java.io.IOException; import java.util.StringTokenizer; import org.a ...

随机推荐

WebService 学习总结
一.概念 Web Web应用程序 Web服务( Web Serivce), SOAP, WSDL, UDDI .Net 框架 ASP.net IIS C#, 代理(委托) 二.实践 1.创建WebSe ...
[转载]MongoDB学习（三）：MongoDB Shell的使用
MongoDB shell MongoDB自带简洁但功能强大的JavaScript shell.JavaScript shell键入一个变量会将变量的值转换为字符串打印到控制台上. 下面介绍基本的操作 ...
在Unity中高效工作（下）
原地址:http://www.unity蛮牛.com/thread-20005-1-1.html Tips for Creating Better Games and Working More Eff ...
CF Rook, Bishop and King
http://codeforces.com/contest/370/problem/A 题意:车是走直线的,可以走任意多个格子,象是走对角线的,也可以走任意多个格子,而国王可以走直线也可以走对角线,但 ...
linux zip 命令详解
功能说明:压缩文件. 语法:zip [-AcdDfFghjJKlLmoqrSTuvVwXyz$][-b <工作目录>][-ll][-n <字尾字符串>][-t <日期时 ...
Java集合类之Hashtable
package com.test; import java.util.*; public class Demo7_3 { public static void main(String[] args) ...
USB Type-C 连接器规范推出之后,市场很多低质量线材容易损坏设备
USB Type-C 连接器规范推出之后,已有不少行动装置产品使用,其中最知名的产品为 Apple MacBook,机身仅提供一组 Type-C 端口,同时兼具充电与数据传输之用.市面上第三方厂商也开 ...
C语言考试解答十题
学院比较奇葩,大一下期让学的VB,这学期就要学C++了,然后在开学的前三个周没有课,就由老师讲三个周的C语言,每天9:30~11:30听课,除去放假和双休日,实际听课时间一共是12天*2小时,下午是1 ...
Document字段发生变化后，报的错
2016-10-11 15:27:47,828 [ERROR] [main] SpringApplication:838 - Application startup failedorg.springf ...
常用linux命令合集（持续更新中）
我的博客:www.while0.com 开发调试 readelf-a 查看elf文件中的内容 hexdump -C 用16进制查看文件 objdump -d 反汇编目标文件 nm 查看目标文件或者可执 ...

【原创】MapReduce编程系列之表连接

【原创】MapReduce编程系列之表连接的更多相关文章

随机推荐

热门专题