This tutorial introduces the concept of pairwise preference used in most ranking problems. I'll use scikit-learn and for learning and matplotlib for visualization.

In the ranking setting, training data consists of lists of items with some order specified between items in each list. This order is typically induced by giving a numerical or ordinal score or a binary judgment (e.g. "relevant" or "not relevant") for each item, so that for any two samples a and b, either a < b, b > a or band a are not comparable.

For example, in the case of a search engine, our dataset consists of results that belong to different queries and we would like to only compare the relevance for results coming from the same query.

This order relation is usually domain-specific. For instance, in information retrieval the set of comparable samples is referred to as a "query id". The goal behind this is to compare only documents that belong to the same query (Joachims 2002). In medical imaging on the other hand, the order of the labels usually depend on the subject so the comparable samples is given by the different subjects in the study (Pedregosa et al 2012).

import itertools

import numpy as np

from scipy import stats

import pylab as pl

from sklearn import svm, linear_model, cross_validation

To start with, we'll create a dataset in which the target values consists of three graded measurements Y = {0, 1, 2} and the input data is a collection of 30 samples, each one with two features.

The set of comparable elements (queries in information retrieval) will consist of two equally sized blocks, X=X1∪X2, where each block is generated using a normal distribution with different mean and covariance. In the pictures, we represent X1 with round markers and X2 with triangular markers.

np.random.seed(0)

theta = np.deg2rad(60)

w = np.array([np.sin(theta), np.cos(theta)])

K = 20

X = np.random.randn(K, 2)

y = [0] * K

for i in range(1, 3):

    X = np.concatenate((X, np.random.randn(K, 2) + i * 4 * w))

    y = np.concatenate((y, [i] * K))

# slightly displace data corresponding to our second partition

X[::2] -= np.array([3, 7])

blocks = np.array([0, 1] * (X.shape[0] / 2))

# split into train and test set

cv = cross_validation.StratifiedShuffleSplit(y, test_size=.5)

train, test = iter(cv).next()

X_train, y_train, b_train = X[train], y[train], blocks[train]

X_test, y_test, b_test = X[test], y[test], blocks[test]

# plot the result

idx = (b_train == 0)

pl.scatter(X_train[idx, 0], X_train[idx, 1], c=y_train[idx],

    marker='^', cmap=pl.cm.Blues, s=100)

pl.scatter(X_train[~idx, 0], X_train[~idx, 1], c=y_train[~idx],

    marker='o', cmap=pl.cm.Blues, s=100)

pl.arrow(0, 0, 8 * w[0], 8 * w[1], fc='gray', ec='gray',

    head_width=0.5, head_length=0.5)

pl.text(0, 1, '$w$', fontsize=20)

pl.arrow(-3, -8, 8 * w[0], 8 * w[1], fc='gray', ec='gray',

    head_width=0.5, head_length=0.5)

pl.text(-2.6, -7, '$w$', fontsize=20)

pl.axis('equal')

pl.show()

In the plot we clearly see that for both blocks there's a common vector w such that the projection onto w gives a list with the correct ordering.

However, because linear considers that output labels live in a metric space it will consider that all pairs are comparable. Thus if we fit this model to the problem above it will fit both blocks at the same time, yielding a result that is clearly not optimal. In the following plot we estimate w^ using an l2-regularized linear model.

ridge = linear_model.Ridge(1.)

ridge.fit(X_train, y_train)

coef = ridge.coef_ / linalg.norm(ridge.coef_)

pl.scatter(X_train[idx, 0], X_train[idx, 1], c=y_train[idx],

    marker='^', cmap=pl.cm.Blues, s=100)

pl.scatter(X_train[~idx, 0], X_train[~idx, 1], c=y_train[~idx],

    marker='o', cmap=pl.cm.Blues, s=100)

pl.arrow(0, 0, 7 * coef[0], 7 * coef[1], fc='gray', ec='gray',

    head_width=0.5, head_length=0.5)

pl.text(2, 0, '$\hat{w}$', fontsize=20)

pl.axis('equal')

pl.title('Estimation by Ridge regression')

pl.show()

To assess the quality of our model we need to define a ranking score. Since we are interesting in a model that ordersthe data, it is natural to look for a metric that compares the ordering of our model to the given ordering. For this, we use Kendall's tau correlation coefficient, which is defined as (P - Q)/(P + Q), being P the number of concordant pairs and Q is the number of discordant pairs. This measure is used extensively in the ranking literature (e.g Optimizing Search Engines using Clickthrough Data).

We thus evaluate this metric on the test set for each block separately.

for i in range(2):

    tau, _ = stats.kendalltau(

        ridge.predict(X_test[b_test == i]), y_test[b_test == i])

    print('Kendall correlation coefficient for block %s: %.5f' % (i, tau))

Kendall correlation coefficient for block 0: 0.71122

Kendall correlation coefficient for block 1: 0.84387

The pairwise transform

As proved in (Herbrich 1999), if we consider linear ranking functions, the ranking problem can be transformed into a two-class classification problem. For this, we form the difference of all comparable elements such that our data is transformed into (x′k,y′k)=(xi−xj,sign(yi−yj)) for all comparable pairs.

This way we transformed our ranking problem into a two-class classification problem. The following plot shows this transformed dataset, and color reflects the difference in labels, and our task is to separate positive samples from negative ones. The hyperplane {x^T w = 0} separates these two classes.

# form all pairwise combinations

comb = itertools.combinations(range(X_train.shape[0]), 2)

k = 0

Xp, yp, diff = [], [], []

for (i, j) in comb:

    if y_train[i] == y_train[j] \

        or blocks[train][i] != blocks[train][j]:

        # skip if same target or different group

        continue

    Xp.append(X_train[i] - X_train[j])

    diff.append(y_train[i] - y_train[j])

    yp.append(np.sign(diff[-1]))

    # output balanced classes

    if yp[-1] != (-1) ** k:

        yp[-1] *= -1

        Xp[-1] *= -1

        diff[-1] *= -1

    k += 1

Xp, yp, diff = map(np.asanyarray, (Xp, yp, diff))

pl.scatter(Xp[:, 0], Xp[:, 1], c=diff, s=60, marker='o', cmap=pl.cm.Blues)

x_space = np.linspace(-10, 10)

pl.plot(x_space * w[1], - x_space * w[0], color='gray')

pl.text(3, -4, '$\{x^T w = 0\}$', fontsize=17)

pl.axis('equal')

pl.show()

As we see in the previous plot, this classification is separable. This will not always be the case, however, in our training set there are no order inversions, thus the respective classification problem is separable.

We will now finally train an Support Vector Machine model on the transformed data. This model is known as RankSVM. We will then plot the training data together with the estimated coefficient w^ by RankSVM.

clf = svm.SVC(kernel='linear', C=.1)

clf.fit(Xp, yp)

coef = clf.coef_.ravel() / linalg.norm(clf.coef_)

pl.scatter(X_train[idx, 0], X_train[idx, 1], c=y_train[idx],

    marker='^', cmap=pl.cm.Blues, s=100)

pl.scatter(X_train[~idx, 0], X_train[~idx, 1], c=y_train[~idx],

    marker='o', cmap=pl.cm.Blues, s=100)

pl.arrow(0, 0, 7 * coef[0], 7 * coef[1], fc='gray', ec='gray',

    head_width=0.5, head_length=0.5)

pl.arrow(-3, -8, 7 * coef[0], 7 * coef[1], fc='gray', ec='gray',

    head_width=0.5, head_length=0.5)

pl.text(1, .7, '$\hat{w}$', fontsize=20)

pl.text(-2.6, -7, '$\hat{w}$', fontsize=20)

pl.axis('equal')

pl.show()

Finally we will check that as expected, the ranking score (Kendall tau) increases with the RankSVM model respect to linear regression.

for i in range(2):

    tau, _ = stats.kendalltau(

        np.dot(X_test[b_test == i], coef), y_test[b_test == i])

    print('Kendall correlation coefficient for block %s: %.5f' % (i, tau))

Kendall correlation coefficient for block 0: 0.83627

Kendall correlation coefficient for block 1: 0.84387

This is indeed higher than the values (0.71122, 0.84387) obtained in the case of linear regression.

Original ipython notebook for this blog post can be found here

"Large Margin Rank Boundaries for Ordinal Regression", R. Herbrich, T. Graepel, and K. Obermayer. Advances in Large Margin Classifiers, 115-132, Liu Press, 2000 ↩
"Optimizing Search Engines Using Clickthrough Data", T. Joachims. Proceedings of the ACM Conference on Knowledge Discovery and Data Mining (KDD), ACM, 2002. ↩
"Learning to rank from medical imaging data", Pedregosa et al. [arXiv] ↩
"Efficient algorithms for ranking with SVMs", O. Chapelle and S. S. Keerthi, Information Retrieval Journal, Special Issue on Learning to Rank, 2009 ↩

Comments !

转：pairwise 代码参考的更多相关文章

Session id实现通过Cookie来传输方法及代码参考
1. Web中的Session指的就是用户在浏览某个网站时,从进入网站到浏览器关闭所经过的这段时间,也就是用户浏览这个网站所花费的时间.因此从上述的定义中我们可以看到,Session实际上是一个特定的 ...
Jquery 代码参考
jquery 代码参考 jQuery(document).ready(function($){}); jQuery(window).on('load', function(){}); $('.vide ...
php 修改后端代码参考
后端代码参考:
【原创】C#模拟Post请求，正文为json数据的代码参考
由于之前一直在做键值对post数据的提交,没遇到过json正文的提交,遇到的问题截图: 对于此种情况的post,我用谷歌插件 PostMan 模拟试了下成功了,截图如下: Postman插件在你选择 ...
公共代码参考（Volley）
Volley 是google提供的一个网络库,相对于自己写httpclient确实方便很多,本文参考部分网上例子整理如下,以作备忘: 定义一个缓存类: public class BitmapCache ...
固定表头/锁定前几列的代码参考[JS篇]
引语:做有难度的事情,才是成长最快的时候.前段时间,接了一个公司的稍微大点的项目,急着赶进度,本人又没有独立带过队,因此,把自己给搞懵逼了.总是没有多余的时间来做自己想做的事,而且,经常把工作带入生活 ...
C语言实现冒泡排序法和选择排序法代码参考
为了易用,我编写排序函数,这和直接在主调函数中用是差不多的. 我认为选择排序法更好理解!请注意 i 和 j ,在写代码时别弄错了,不然很难找到错误! 冒泡排序法 void sort(int * ar, ...
.OpenWrt驱动程序Makefile的分析概述、驱动程序代码参考、以及测试程序代码参考
# # # include $(TOPDIR)/rules.mk //一般在 Makefile 的开头 include $(INCLUDE_DIR)/kernel.mk // 文件对于软件包为内核时 ...
Flex组件参考代码参考汇总
1:tourdeflex快速熟悉各种组件用法的参考http://www.adobe.com/devnet/flex/tourdeflex.html在线:http://www.adobe.com/dev ...

随机推荐

【BZOJ4803】逆欧拉函数
[BZOJ4803]逆欧拉函数题面 bzoj 题解题目是给定你$\varphi(n)$要求前$k$小的$n$. 设$n=\prod_{i=1}^k{p_i}^{c_i}$ 则\(\ ...
Arduino 101/Genuino101使用-第2篇
1. Arduino 101编程只是在ARC的核心上进行,其具体架构为ARCv2EM.. 2. 而Quark核心,从目前可知的信息来看,其应该运行着名为Zephyr的RTOS 3.101并没有EEPR ...
【JUC源码解析】FutureTask
简介 FutureTask, 一个支持取消行为的异步任务执行器. 概述 FutureTask实现了Future,提供了start, cancel, query等功能,并且实现了Runnable接口,可 ...
Python闭包相关问题
闭包的概念一直是似懂非懂,看过了原理,却不知道怎么实际应用. 刚好看到Python的late binding问题,记录如下,以备后续增补. >>> def create_multip ...
Linux 安装FastDFS<单机版>(使用Mac远程访问)
阅读本文需要先阅读安装FastDFS<准备> 一编译环境 yum install gcc-c++ yum -y install libevent yum install -y pcre ...
这才是球王应有的技艺，他就是C罗
四年一度的世界杯在本周四拉开了帷幕,俄罗斯以5:0碾压沙特阿拉伯,让我们惊呼战斗名族的强大,其后的摩洛哥VS伊朗,摩洛哥前锋布哈杜兹将足球顶入自家球门,这......咳,咳,本来是为了解围,没想到成就 ...
Selenium WebDriver 下 plugin container for firefox has stopped working
用selenium 的webdriver 和 firefox 浏览器做自动化测试,经常会出现 plugin container for firefox has stopped working 如下图所 ...
阿里与ShopRunner达成协议联手在国内推出服务
阿里巴巴集团与美国在线零售商 ShopRunner 达成协议,将帮助后者在中国大陆销售商品和履行订单交付产品. ShopRunner 首席战略官菲奥娜·迪亚斯(Fiona Dias)周三接受媒体采访时 ...
pyextend库-accepts函数参数检查
pyextend - python extend lib accepts(exception=TypeError, **types) 参数: exception: 检查失败时的抛出异常类型 **typ ...
第六次作业psp
psp 进度条代码累积折线图博文累积折线图 psp饼状图

转：pairwise 代码参考

Learning to rank with scikit-learn: the pairwise transform

The pairwise transform

Comments !

转：pairwise 代码参考的更多相关文章

随机推荐

热门专题