Chi Square Distance
The chi squared distance d(x,y) is, as you already know, a distance between two histograms x=[x_1,..,x_n] and y=[y_1,...,y_n] having n bins both. Moreover, both histograms are normalized, i.e. their entries sum up to one.
The distance measure d is usually defined (although alternative definitions exist) as d(x,y) = sum( (xi-yi)^2 / (xi+yi) ) / 2 . It is often used in computer vision to compute distances between some bag-of-visual-word representations of images.
The name of the distance is derived from Pearson's chi squared test statistic X²(x,y) = sum( (xi-yi)^2 / xi) for comparing discrete probability distributions (i.e histograms). However, unlike the test statistic, d(x,y) is symmetric wrt. x and y, which is often useful in practice, e.g., when you want to construct a kernel out of the histogram distances.
Chi-Square Distance
Consider a frequency table with n rows and p columns, it is possible to calculate row profiles and column profiles. Let us then plot the n or p points from each profile. We can define the distances between these points. The Euclidean distance between the components of the profiles, on which a weighting is defined (each term has a weight that is the inverse of its frequency), is called the chi-square distance. The name of the distance is derived from the fact that the mathematical expression defining the distance is identical to that encountered in the elaboration of the chi square goodness of fit test.
MATHEMATICAL ASPECTS
![]() |
where
| f i. | is the sum of the components of the ith row; |
| f .j | is the sum of the components of the jth column; |
![]() |
is the ith row profile for j = 1,2,...,p. |
![]() |
where
is the jth column profile for j = 1,...,n.
DOMAINS AND LIMITATIONS
The chi-square distance incorporates a weight that is inversely proportional to the total of each row (or column), which increases the importance of small deviations in the rows (or columns) which have a small sum with respect to those with more important sum package.
The chi-square distance has the property of distributional equivalence, meaning that it ensures that the distances between rows and columns are invariant when two columns (or two rows) with identical profiles are aggregated.
EXAMPLES
Consider a contingency table charting how satisfied employees working for three different businesses are. Let us establish a distance table using the chi-square distance.
Values for the studied variable X can fall into one of three categories:
- X 1: high satisfaction;
- X 2: medium satisfaction;
- X 3: low satisfaction.
The observations collected from samples of individuals from the three businesses are given below:
|
Business 1 |
Business 2 |
Business 3 |
Total |
|
|---|---|---|---|---|
|
X 1 |
20 |
55 |
30 |
105 |
|
X 2 |
18 |
40 |
15 |
73 |
|
X 3 |
12 |
5 |
5 |
22 |
|
Total |
50 |
100 |
50 |
200 |
The relative frequency table is obtained by dividing all of the elements of the table by 200, the total number of observations:
|
Business 1 |
Business 2 |
Business 3 |
Total |
|
|---|---|---|---|---|
|
X 1 |
0.1 |
0.275 |
0.15 |
0.525 |
|
X 2 |
0.09 |
0.2 |
0.075 |
0.365 |
|
X 3 |
0.06 |
0.025 |
0.025 |
0.11 |
|
Total |
0.25 |
0.5 |
0.25 |
1 |
We can calculate the difference in employee satisfaction between the the 3 enterprises. The column profile matrix is given below:
|
Business 1 |
Business 2 |
Business 3 |
Total |
|
|---|---|---|---|---|
|
X 1 |
0.4 |
0.55 |
0.6 |
1.55 |
|
X 2 |
0.36 |
0.4 |
0.3 |
1.06 |
|
X 3 |
0.24 |
0.05 |
0.1 |
0.39 |
|
Total |
1 |
1 |
1 |
3 |
![]() |
We can calculate d(1,3) and d(2,3) in a similar way. The distances obtained are summarized in the following distance table:
|
Business 1 |
Business 2 |
Business 3 |
|
|---|---|---|---|
|
Business 1 |
0 |
0.613 |
0.514 |
|
Business 2 |
0.613 |
0 |
0.234 |
|
Business 3 |
0.514 |
0.234 |
0 |
We can also calculate the distances between the rows, in other words the difference in employee satisfaction; to do this we need the line profile table:
|
Business 1 |
Business 2 |
Business 3 |
Total |
|
|---|---|---|---|---|
|
X 1 |
0.19 |
0.524 |
0.286 |
1 |
|
X 2 |
0.246 |
0.548 |
0.206 |
1 |
|
X 3 |
0.546 |
0.227 |
0.227 |
1 |
|
Total |
0.982 |
1.299 |
0.719 |
3 |
![]() |
We can calculate d(1,3) and d(2,3) in a similar way. The differences between the degrees of employee satisfaction are finally summarized in the following distance table:
|
X 1 |
X 2 |
X 3 |
|
|---|---|---|---|
|
X 1 |
0 |
0.198 |
0.835 |
|
X 2 |
0.198 |
0 |
0.754 |
|
X 3 |
0.835 |
0.754 |
0 |
http://www.researchgate.net/post/What_is_chi-squared_distance_I_need_help_with_the_source_code
http://www.springerreference.com/docs/html/chapterdbid/60817.html
Chi Square Distance的更多相关文章
- BestCoder Round #87 1002 Square Distance[DP 打印方案]
Square Distance Accepts: 73 Submissions: 598 Time Limit: 4000/2000 MS (Java/Others) Memory Limit ...
- HDU 5903 Square Distance (贪心+DP)
题意:一个字符串被称为square当且仅当它可以由两个相同的串连接而成. 例如, "abab", "aa"是square, 而"aaa", ...
- hdu 5903 Square Distance(dp)
Problem Description A string is called a square string if it can be obtained by concatenating two co ...
- [HDU5903]Square Distance(DP)
题意:给一个字符串t ,求与这个序列刚好有m个位置字符不同的由两个相同的串拼接起来的字符串 s,要求字典序最小的答案. 分析:按照贪心的想法,肯定在前面让字母尽量小,尽可能的填a,但问题是不知道前面填 ...
- BendFord's law's Chi square test
http://www.siam.org/students/siuro/vol1issue1/S01009.pdf bendford'law e=log10(1+l/n) o=freq of first ...
- HDU 5903 - Square Distance [ DP ] ( BestCoder Round #87 1002 )
题意: 给一个字符串t ,求与这个序列刚好有m个位置字符不同的由两个相同的串拼接起来的字符串 s, 要求字典序最小的答案 分析: 把字符串折半,分成0 - n/2-1 和 n/2 - n-1 d ...
- HDU 5903 Square Distance
$dp$预处理,贪心. 因为$t$串前半部分和后半部分是一样的,所以只要构造前一半就可以了. 因为要求字典序最小,所以肯定是从第一位开始贪心选择,$a,b,c,d,...z$,一个一个尝试过去,如果发 ...
- 生成式模型之 GAN
生成对抗网络(Generative Adversarial Networks,GANs),由2014年还在蒙特利尔读博士的Ian Goodfellow引入深度学习领域.2016年,GANs热潮席卷AI ...
- Scoring and Modeling—— Underwriting and Loan Approval Process
https://www.fdic.gov/regulations/examinations/credit_card/ch8.html Types of Scoring FICO Scores V ...
随机推荐
- POJ2479,2593: 两段maximum-subarray问题
虽然是两个水题,但是一次AC的感觉真心不错 这个问题算是maximum-subarray问题的升级版,不过主要算法思想不变: 1. maximum-subarray问题 maximum-subarra ...
- CF29D - Ant on the Tree(DFS)
题目大意 给定一棵树,要求你按给定的叶子节点顺序对整棵树进行遍历,并且恰好经过2*n-1个点,输出任意一条符合要求的路径 题解 每次从叶子节点开始遍历到上一个叶子节点就OK了, 这个就是符合要求的路径 ...
- RIA算法解决最小覆盖圆问题
一.概念引入 最小包围圆问题:对于给定的平面上甩个点所组成的一个集合P,求出P的最小包围圆,即包含P中所有点.半径最小的那个圆.也就是求出这个最小 包围圆的圆心位置和半径. ...
- 利用系统镜像文件安装.Net框架的方式
最近重装系统之后,在安装部分程序时需要.NET3.5框架,在线安装时间较长,网上搜到了一个很好的解决办法.利用windows系统镜像.首先将镜像加载到驱动中比如L,然后在cmd中输入 dism.exe ...
- maven分模块间依赖注意事项
1.被依赖模块应该先通过 maven -install 命令将该模块打包为jar发布到本地仓库 2.引用的模块通过在pom.xml文件中添加dependence引用 maven -package 将项 ...
- [Git]git常用命令总结
git clone url 将远程库复制到本地 git status 查看本地库的状态 git add filename.filetype 将库中被修改的文件标记为添加状态 git diff 查看库中 ...
- iOS开发:创建真机调试证书
关于苹果iOS开发,笔者也是从小白过来的,经历过各种困难和坑,其中就有关于开发证书,生产证书,in_house证书,add_Hoc证书申请过程中的问题,以及上架发布问题.今天就着重说一下关于针对于苹果 ...
- jQuery获取鼠标移动方向2
(function($) { $.fn.extend({ show: function(div) { var w = this.width(), h = this.height(), xpos = w ...
- jenkens构建脚本
Build Root POM Goals and options Command # consts SERVER="192.168.60.209" DEPLOY=" ...
- psd via fft and pwelch
%fft and pwelch方法求取功率谱load x.mat Fs = 1; t = (0:1/Fs:1-1/Fs).'; Nx = length(x); % Window data w = ha ...




