【Hadoop学习之十】MapReduce案例分析二-好友推荐

环境
　　虚拟机：VMware 10
　　Linux版本：CentOS-6.5-x86_64
　　客户端：Xshell4
　　FTP：Xftp4
　　jdk8
　　hadoop-3.1.1

最应该推荐的好友TopN，如何排名？

tom hello hadoop cat

world hadoop hello hive

cat tom hive

mr hive hello

hive cat hadoop world hello mr

hadoop tom hive world

hello tom world hive mr

package test.mr.fof;

import org.apache.hadoop.conf.Configuration;

import org.apache.hadoop.fs.Path;

import org.apache.hadoop.io.IntWritable;

import org.apache.hadoop.io.Text;

import org.apache.hadoop.mapreduce.Job;

import org.apache.hadoop.mapreduce.lib.input.FileInputFormat;

import org.apache.hadoop.mapreduce.lib.output.FileOutputFormat;

import org.apache.hadoop.util.GenericOptionsParser;

public class MyFOF {

    /**

     * 最应该推荐的好友TopN，如何排名？

     * @param args

     * @throws Exception

     */

    public static void main(String[] args) throws Exception {

        Configuration conf = new Configuration(true);

        String[] otherArgs = new GenericOptionsParser(conf, args).getRemainingArgs();

        conf.set("sleep", otherArgs[2]);

        Job job = Job.getInstance(conf,"FOF");

        job.setJarByClass(MyFOF.class);

        //Map

        job.setMapperClass(FMapper.class);

        job.setMapOutputKeyClass(Text.class);

        job.setMapOutputValueClass(IntWritable.class);

        //Reduce

        job.setReducerClass(FReducer.class);

        //HDFS 输入路径

        Path input = new Path(otherArgs[0]);

        FileInputFormat.addInputPath(job, input );

        //HDFS 输出路径

        Path output = new Path(otherArgs[1]);

        if(output.getFileSystem(conf).exists(output)){

            output.getFileSystem(conf).delete(output,true);

        }

        FileOutputFormat.setOutputPath(job, output );

        System.exit(job.waitForCompletion(true) ? 0 :1);

    }

//    tom hello hadoop cat

//    world hadoop hello hive

//    cat tom hive

//    mr hive hello

//    hive cat hadoop world hello mr

//    hadoop tom hive world

//    hello tom world hive mr

}

package test.mr.fof;

import java.io.IOException;

import org.apache.hadoop.io.IntWritable;

import org.apache.hadoop.io.LongWritable;

import org.apache.hadoop.io.Text;

import org.apache.hadoop.mapreduce.Mapper;

import org.apache.hadoop.util.StringUtils;

public class FMapper extends Mapper<LongWritable, Text, Text, IntWritable>{

    Text mkey= new Text();

    IntWritable mval = new IntWritable();

    @Override

    protected void map(LongWritable key, Text value,Context context)

            throws IOException, InterruptedException {

        //value: 0-直接关系  1-间接关系

        //tom       hello hadoop cat   :   hello:hello  1

        //hello     tom world hive mr      hello:hello  0

        String[] strs = StringUtils.split(value.toString(), ' ');

        String user=strs[0];

        String user01=null;

        for(int i=1;i<strs.length;i++){

            //与好友清单中好友属于直接关系

            mkey.set(fof(strs[0],strs[i]));

            mval.set(0);

            context.write(mkey, mval);  

            for (int j = i+1; j < strs.length; j++) {

                Thread.sleep(context.getConfiguration().getInt("sleep", 0));

                //好友列表内 成员之间是间接关系

                mkey.set(fof(strs[i],strs[j]));

                mval.set(1);

                context.write(mkey, mval);

            }

        }

    }

    public static String fof(String str1  , String str2){

        if(str1.compareTo(str2) > 0){

            //hello,hadoop

            return str2+":"+str1;

        }

        //hadoop,hello

        return str1+":"+str2;

    }

}

package test.mr.fof;

import java.io.IOException;

import org.apache.hadoop.io.IntWritable;

import org.apache.hadoop.io.Text;

import org.apache.hadoop.mapreduce.Reducer;

public class FReducer  extends  Reducer<Text, IntWritable, Text, Text> {

    Text rval = new Text();

    @Override

    protected void reduce(Text key, Iterable<IntWritable> vals, Context context)

            throws IOException, InterruptedException

    {

        //是简单的好友列表的差集吗？

        //最应该推荐的好友TopN，如何排名？

        //hadoop:hello  1

        //hadoop:hello  0

        //hadoop:hello  1

        //hadoop:hello  1

        int sum=0;

        int flg=0;

        for (IntWritable v : vals)

        {

            //0为直接关系

            if(v.get()==0){

                //hadoop:hello  0

                flg=1;

            }

            sum += v.get();

        }

        //只有间接关系才会被输出

        if(flg==0){

            rval.set(sum+"");

            context.write(key, rval);

        }

    }

}

【Hadoop学习之十】MapReduce案例分析二-好友推荐的更多相关文章

MapReduce深度分析(二)
MapReduce深度分析(二) 五.JobTracker分析 JobTracker是hadoop的重要的后台守护进程之一,主要的功能是管理任务调度.管理TaskTracker.监控作业执行.运行作业 ...
Hadoop学习笔记—20.网站日志分析项目案例（二）数据清洗
网站日志分析项目案例(一)项目介绍:http://www.cnblogs.com/edisonchou/p/4449082.html 网站日志分析项目案例(二)数据清洗:当前页面网站日志分析项目案例 ...
【Hadoop学习之十三】MapReduce案例分析五-ItemCF
环境虚拟机:VMware 10 Linux版本:CentOS-6.5-x86_64 客户端:Xshell4 FTP:Xftp4 jdk8 hadoop-3.1.1 推荐系统——协同过滤(Collab ...
Hadoop学习笔记—20.网站日志分析项目案例（一）项目介绍
网站日志分析项目案例(一)项目介绍:当前页面网站日志分析项目案例(二)数据清洗:http://www.cnblogs.com/edisonchou/p/4458219.html 网站日志分析项目案例 ...
Hadoop学习笔记—20.网站日志分析项目案例（三）统计分析
网站日志分析项目案例(一)项目介绍:http://www.cnblogs.com/edisonchou/p/4449082.html 网站日志分析项目案例(二)数据清洗:http://www.cnbl ...
[b0012] Hadoop 版hello word mapreduce wordcount 运行(二)
目的: 学习Hadoop mapreduce 开发环境eclipse windows下的搭建环境: Winows 7 64 eclipse 直接连接hadoop运行的环境已经搭建好,结果输出到ecl ...
【第二课】kaggle案例分析二
Evernote Export 推荐系统比赛(常见比赛) 推荐系统分类最能变现的机器学习应用基于应用领域分类:电子商务推荐,社交好友推荐,搜索引擎推荐,信息内容推荐等 **基于设计思想:**基于协 ...
【Hadoop学习之十二】MapReduce案例分析四-TF-IDF
环境虚拟机:VMware 10 Linux版本:CentOS-6.5-x86_64 客户端:Xshell4 FTP:Xftp4 jdk8 hadoop-3.1.1 概念TF-IDF(term fre ...
【Hadoop学习之九】MapReduce案例分析一-天气
环境虚拟机:VMware 10 Linux版本:CentOS-6.5-x86_64 客户端:Xshell4 FTP:Xftp4 jdk8 hadoop-3.1.1 找出每个月气温最高的2天 1949 ...

随机推荐

关于javascript中defineProperty的学习
语法 Object.defineProperty(obj, prop, descriptor) 参数 obj 要在其上定义属性的对象. prop 要定义或修改的属性的名称. descriptor 将被 ...
简单的document操作
1.新增商品:新建文档,建立索引PUT /index/type/id{ "json数据"}例如:PUT /ecommerce/product/1{ "name" ...
Innodb semi-consistent 简介
A type of read operation used for UPDATE statements, that is a combination of read committed and con ...
ios禁止页面下拉
document.querySelector('body').addEventListener('touchmove', function(e) { e.preventDefault(); } ...
linux 查看系统负载：uptime
uptime命令用于查看系统负载,跟 w 命令的输出内容一致 [root@mysql ~]# uptime :: up days, :, user, load average: 1.12, 0.97, ...
what's the 跨期套利
出自 MBA智库百科(https://wiki.mbalib.com/) 跨期套利的定义跨期套利是套利交易中最普遍的一种,是股指期货的跨期套利(Calendar Spread Arbitrage)即 ...
dos2unix命令
dos2unix命令用来将DOS格式的文本文件转换成UNIX格式的(DOS/MAC to UNIX text file format converter).DOS下的文本文件是以\r\n作为断行标志的 ...
TlistView基本使用
//增加 procedure TForm1.Button1Click(Sender: TObject); var lsItem: TListItem; begin lsItem := ListView ...
JavaScript基础学习2
/* 1.把函数作为参数.匿名函数作为参数传递到函数 */ function dogEat(food) { console.log("dog eat " + food); } fu ...
linux中cmake语法的学习
在linux 下进行开发很多人选择编写makefile 文件进行项目环境搭建,而makefile 文件依赖关系复杂,工作量很大,搞的人头很大.常常,写代码,效率才是王道.这里还有自动化的项目构建工具C ...

【Hadoop学习之十】MapReduce案例分析二-好友推荐

【Hadoop学习之十】MapReduce案例分析二-好友推荐的更多相关文章

随机推荐

热门专题