Apache mahout 源码阅读笔记-DataModel之UserBaseRecommender

先来看一下使用流程：

1）拿到DataModel

2）定义相似度计算模型 PearsonCorrelationSimilarity

3）定义用户邻域计算模型 NearestNUserNeighborhood

4）定义推荐模型 GenericUserBasedRecommender

5)进行推荐

  @Test

  public void testHowMany() throws Exception {

    DataModel dataModel = getDataModel(

            new long[] {1, 2, 3, 4, 5},

            new Double[][] {

                    {0.1, 0.2},

                    {0.2, 0.3, 0.3, 0.6},

                    {0.4, 0.4, 0.5, 0.9},

                    {0.1, 0.4, 0.5, 0.8, 0.9, 1.0},

                    {0.2, 0.3, 0.6, 0.7, 0.1, 0.2},

            });

    //用于计算最相似的用户,领域用户

    UserSimilarity similarity = new PearsonCorrelationSimilarity(dataModel);

    UserNeighborhood neighborhood = new NearestNUserNeighborhood(2, similarity, dataModel);

    Recommender recommender = new GenericUserBasedRecommender(dataModel, neighborhood, similarity);

    List<RecommendedItem> fewRecommended = recommender.recommend(1, 2);

    List<RecommendedItem> moreRecommended = recommender.recommend(1, 4);

    for (int i = 0; i < fewRecommended.size(); i++) {

      assertEquals(fewRecommended.get(i).getItemID(), moreRecommended.get(i).getItemID());

    }

    recommender.refresh(null);

    for (int i = 0; i < fewRecommended.size(); i++) {

      assertEquals(fewRecommended.get(i).getItemID(), moreRecommended.get(i).getItemID());

    }

  }

相似度计算，参考上篇的PearsonCorrelationSimilarity。

NearestNUserNeighborhood ，获取最近的N个用户，怎么实现的呢？
~/mahout-core/src/main/java/org/apache/mahout/cf/taste/impl/recommender/GenericUserBasedRecommender.java

  @Override

  public List<RecommendedItem> recommend(long userID, int howMany, IDRescorer rescorer) throws TasteException {

    Preconditions.checkArgument(howMany >= 1, "howMany must be at least 1");

    log.debug("Recommending items for user ID '{}'", userID);

    //根据similarity模型进行计算，计算最相似的N个用户

    long[] theNeighborhood = neighborhood.getUserNeighborhood(userID);

    if (theNeighborhood.length == 0) {

      return Collections.emptyList();

    }

    //获取其他领域用户进行评分而且当前用户所没有进行评分的Item列表，作为推荐的基本池子

    FastIDSet allItemIDs = getAllOtherItems(theNeighborhood, userID);

    //获取池子里面,当前用户偏好最高的TopN进行推荐

    TopItems.Estimator<Long> estimator = new Estimator(userID, theNeighborhood);

    List<RecommendedItem> topItems = TopItems

        .getTopItems(howMany, allItemIDs.iterator(), rescorer, estimator);

    log.debug("Recommendations are: {}", topItems);

    return topItems;

  }

Estimator的实现，是这样的：

  private final class Estimator implements TopItems.Estimator<Long> {

    private final long theUserID;

    private final long[] theNeighborhood;

    Estimator(long theUserID, long[] theNeighborhood) {

      this.theUserID = theUserID;

      this.theNeighborhood = theNeighborhood;

    }

    @Override

    public double estimate(Long itemID) throws TasteException {

      return doEstimatePreference(theUserID, theNeighborhood, itemID);

    }

  }

}

  protected float doEstimatePreference(long theUserID, long[] theNeighborhood, long itemID) throws TasteException {

    //把相似用户对该Item的偏好累加起来,再做平均值,当做当前用户对改Item的偏好

    if (theNeighborhood.length == 0) {

      return Float.NaN;

    }

    DataModel dataModel = getDataModel();

    double preference = 0.0;

    double totalSimilarity = 0.0;

    int count = 0;

    for (long userID : theNeighborhood) {

      if (userID != theUserID) {

        // See GenericItemBasedRecommender.doEstimatePreference() too

        Float pref = dataModel.getPreferenceValue(userID, itemID);

        if (pref != null) {

          double theSimilarity = similarity.userSimilarity(theUserID, userID);

          if (!Double.isNaN(theSimilarity)) {

            preference += theSimilarity * pref;

            totalSimilarity += theSimilarity;

            count++;

          }

        }

      }

    }

    // Throw out the estimate if it was based on no data points, of course, but also if based on

    // just one. This is a bit of a band-aid on the 'stock' item-based algorithm for the moment.

    // The reason is that in this case the estimate is, simply, the user's rating for one item

    // that happened to have a defined similarity. The similarity score doesn't matter, and that

    // seems like a bad situation.

    if (count <= 1) {

      return Float.NaN;

    }

    float estimate = (float) (preference / totalSimilarity);

    if (capper != null) {

      estimate = capper.capEstimate(estimate);

    }

    return estimate;

  }

总结：
1）计算最相似的N个用户
2）从最相似的N个用户中，获取自己没有评分过的Item
3）预计自己对每个Item的偏好
4）取偏好最高的N个Item进行推荐

Apache mahout 源码阅读笔记-DataModel之UserBaseRecommender的更多相关文章

Apache mahout 源码阅读笔记--DataModel之FileDataModel
要做推荐,用户行为数据是基础. 用户行为数据有哪些字段呢? mahout的DataModel支持,用户ID,ItemID是必须的,偏好值(用户对当前Item的评分),时间戳这四个字段 {@code ...
Apache mahout 源码阅读笔记--协同过滤, PearsonCorrelationSimilarity
协同过滤源码路径: ~/project/javaproject/mahout-0.9/core/src $tree main/java/org/apache/mahout/cf/taste/ -L 2 ...
Apache Storm源码阅读笔记
欢迎转载,转载请注明出处. 楔子自从建了Spark交流的QQ群之后,热情加入的同学不少,大家不仅对Spark很热衷对于Storm也是充满好奇.大家都提到一个问题就是有关storm内部实现机理的资料比 ...
Mina源码阅读笔记（四）—Mina的连接IoConnector2
接着Mina源码阅读笔记(四)-Mina的连接IoConnector1,,我们继续: AbstractIoAcceptor: 001 package org.apache.mina.core.rewr ...
CI框架源码阅读笔记5 基准测试 BenchMark.php
上一篇博客(CI框架源码阅读笔记4 引导文件CodeIgniter.php)中,我们已经看到:CI中核心流程的核心功能都是由不同的组件来完成的.这些组件类似于一个一个单独的模块,不同的模块完成不同的功 ...
CI框架源码阅读笔记4 引导文件CodeIgniter.php
到了这里,终于进入CI框架的核心了.既然是“引导”文件,那么就是对用户的请求.参数等做相应的导向,让用户请求和数据流按照正确的线路各就各位.例如,用户的请求url: http://you.host.c ...
CI框架源码阅读笔记3 全局函数Common.php
从本篇开始,将深入CI框架的内部,一步步去探索这个框架的实现.结构和设计. Common.php文件定义了一系列的全局函数(一般来说,全局函数具有最高的加载优先权,因此大多数的框架中BootStrap ...
CI框架源码阅读笔记2 一切的入口 index.php
上一节(CI框架源码阅读笔记1 - 环境准备.基本术语和框架流程)中,我们提到了CI框架的基本流程,这里再次贴出流程图,以备参考: 作为CI框架的入口文件,源码阅读,自然由此开始.在源码阅读的过程中, ...
源码阅读笔记 - 1 MSVC2015中的std::sort
大约寒假开始的时候我就已经把std::sort的源码阅读完毕并理解其中的做法了,到了寒假结尾,姑且把它写出来这是我的第一篇源码阅读笔记,以后会发更多的,包括算法和库实现,源码会按照我自己的代码风格格 ...

随机推荐

结构体内重载小于号< 及构造函数
struct Node { int d, e; bool operator < (const Node x) const { return x.d < d; } Node(int d, i ...
poj3261（后缀数组）
题意:给出一串长度为n的字符,再给出一个k值,要你求重复次数大于等于k次的最长子串长度........ 思路:其实也非常简单,直接求出height值,然后将它分组,二分答案......结果就出来了.. ...
Bootstrap学习笔记（9）--模态框（登录/注册弹框）
说明: 1. 上来一个ul先把登录和注册两个链接扔进去,ul的类nav,navbar-nav是导航条,navbar-right让他固定在右侧.每个li的里面,data-toggle="mod ...
安装expect命令两种方式
yum安装 yum -y install expect 手动安装 expect以及tcl版本 #!/bin/bash oldpath=`pwd` tar -zxf tcl8.4.20-src.tar. ...
Java反射机制在代理模式中的使用
代理模式的核心思路就是一个接口有两个子类,一个子类完成核心的业务操作,另一个子类完成与核心业务有关的辅助性操作. 代理模式分为静态代理模式和动态代理模式. 静态代理模式: //接口类 interfa ...
jquery 排除重复
应用场景——双盒选择器两个select可能会出现重复的情况排除重复代码如下: /** * 删除$fromGroup中与$toGroup重复的option * @param $fromGroup = ...
Spider Studio 新版本 (20140225) - 设置菜单调整 / 提供JQueryContext布局相关的方法
这是年后的第一个新版本, 包含如下: 1. 先前去掉的浏览器设置功能又回来了! 说来惭愧, 去掉了这两个功能之后发现浏览经常会被JS错误打断, 很不方便, 于是乎又把它们给找回来了. :) 2. 为J ...
PHP之PHP文件引用详解
HP的文件引用涉及到四个函数: 文件引用 1.include()2.include_once()3.require()4.require_once() 这四个函数常常会给PHP初学者造成困扰,总的来说 ...
数据库 Proc编程二
#define _CRT_SECURE_NO_WARNINGS #include <stdio.h> #include <stdlib.h> #include <stri ...
Hadoop源码分析之客户端向HDFS写数据
转自:http://www.tuicool.com/articles/neUrmu 在上一篇博文中分析了客户端从HDFS读取数据的过程,下面来看看客户端是怎么样向HDFS写数据的,下面的代码将本地文件 ...

Apache mahout 源码阅读笔记-DataModel之UserBaseRecommender

Apache mahout 源码阅读笔记-DataModel之UserBaseRecommender的更多相关文章

随机推荐

热门专题