今天遇到一个性能问题,再调优过程中发现耗时最久的计划是exist 部分涉及的三个表。

然后计划用left join 来替换exist,然后查询了很多资料,大部分都说exist和left join 性能差不多。 为了验证这一结论进行了如下实验

步骤如下

1、创建测试表

drop table app_family;

CREATE TABLE app_family (

"family_id" character varying(32 char) NOT NULL,

"application_id" character varying(32 char) NULL,

"family_number" character varying(50 char) ,

"household_register_number" character varying(50 char),

"poverty_reason" character varying(32 char),

CONSTRAINT "pk_app_family_idpk" PRIMARY KEY (family_id));

insert into app_family select generate_series(1,1000000),generate_series(1,1000000),'aaaa','aaa','bbb' from dual ;

create table app_family2 as select * from app_family;

create table app_memeber as select * from app_family;

2、验证两张表join和exist 性能对比

语句1、两张表exist

explain analyze select a1.application_id,a1.family_id from app_family a1 where

a1.family_id >1000 and

EXISTS(

SELECT

1

FROM

app_family2 a2

WHERE

a2.application_id=a1.application_id

and a2.family_id > 500000

)

总计用时646.203 ms

 ----------------------------------------------------------------------------------------------------------------------------------------------------
Gather (cost=16927.11..44466.84 rows=111111 width=12) (actual time=354.314..621.714 rows=500000 loops=1)
Workers Planned: 2
Workers Launched: 2
-> Parallel Hash Semi Join (cost=15927.11..32355.74 rows=46296 width=12) (actual time=355.657..512.049 rows=166667 loops=3)
Hash Cond: ((a1.application_id)::text = (a2.application_id)::text)
-> Parallel Seq Scan on app_family a1 (cost=0.00..13648.00 rows=138889 width=12) (actual time=0.222..111.618 rows=333000 loops=3)
Filter: ((family_id)::integer > 1000)
Rows Removed by Filter: 333
-> Parallel Hash (cost=13648.00..13648.00 rows=138889 width=6) (actual time=149.203..149.204 rows=166667 loops=3)
Buckets: 131072 Batches: 8 Memory Usage: 3520kB
-> Parallel Seq Scan on app_family2 a2 (cost=0.00..13648.00 rows=138889 width=6) (actual time=48.576..109.251 rows=166667 loops=3)
Filter: ((family_id)::integer > 500000)
Rows Removed by Filter: 166667
Planning Time: 0.145 ms
Execution Time: 645.095 ms
(15 rows) Time: 646.203 ms
kingbase=#

语句2 两张表join

explain analyze select a1.application_id,a1.family_id from app_family a1 LEFT JOIN app_family2 a2 ON a2.application_id=a1.application_id

WHERE a1.family_id >1000 AND a2.family_id > 500000

总计执行时间624.211 ms

---------------------------------------------------------------------------------------------------------------------------------------------------
Gather (cost=16927.11..44300.95 rows=111111 width=12) (actual time=349.752..601.304 rows=500000 loops=1)
Workers Planned: 2
Workers Launched: 2
-> Parallel Hash Join (cost=15927.11..32189.85 rows=46296 width=12) (actual time=337.548..508.139 rows=166667 loops=3)
Hash Cond: ((a1.application_id)::text = (a2.application_id)::text)
-> Parallel Seq Scan on app_family a1 (cost=0.00..13648.00 rows=138889 width=12) (actual time=0.087..111.949 rows=333000 loops=3)
Filter: ((family_id)::integer > 1000)
Rows Removed by Filter: 333
-> Parallel Hash (cost=13648.00..13648.00 rows=138889 width=6) (actual time=131.718..131.719 rows=166667 loops=3)
Buckets: 131072 Batches: 8 Memory Usage: 3488kB
-> Parallel Seq Scan on app_family2 a2 (cost=0.00..13648.00 rows=138889 width=6) (actual time=31.730..90.917 rows=166667 loops=3)
Filter: ((family_id)::integer > 500000)
Rows Removed by Filter: 166667
Planning Time: 0.093 ms
Execution Time: 623.465 ms
(15 rows) Time: 624.211 ms

两张表场景总结

针对两张表的对比可以发现join还相对满了10几ms但是总的来说两边 差异不大。所以再两张表的关联情况下 join和exist 性能相近。

3、验证3张表join和exist 性能对比

语句1 三张表exist

本场景最开始执行时 exit 用户6 s多,原因时用到了内存排序,后来调整了work_mem 排除了内存排序的影响,最终执行时间

2911.146 ms

explain analyze select a1.application_id,a1.family_id from app_family a1 ,app_family2 a2 where

a1.family_id >1000 and a2.family_id < 900000 and

EXISTS(

SELECT

1

FROM

app_memeber m

WHERE

m.application_id=a1.application_id

and m.family_id=a2.family_id

)

------------------------------------------------------------------------------------------------------------------------------------------------------
--
Gather (cost=61282.11..88664.67 rows=111111 width=12) (actual time=2112.079..2847.233 rows=898999 loops=1)
Workers Planned: 2
Workers Launched: 2
-> Parallel Hash Join (cost=60282.11..76553.57 rows=46296 width=12) (actual time=2119.345..2705.935 rows=299666 loops=3)
Hash Cond: ((m.family_id)::text = (a2.family_id)::text)
-> Hash Join (cost=44898.00..60455.72 rows=138889 width=18) (actual time=1885.923..2264.850 rows=333000 loops=3)
Hash Cond: ((a1.application_id)::text = (m.application_id)::text)
-> Parallel Seq Scan on app_family a1 (cost=0.00..13648.00 rows=138889 width=12) (actual time=0.091..109.196 rows=333000 loops=3)
Filter: ((family_id)::integer > 1000)
Rows Removed by Filter: 333
-> Hash (cost=32398.00..32398.00 rows=1000000 width=12) (actual time=1880.027..1880.028 rows=1000000 loops=3)
Buckets: 1048576 Batches: 1 Memory Usage: 52897kB
-> HashAggregate (cost=22398.00..32398.00 rows=1000000 width=12) (actual time=957.973..1382.683 rows=1000000 loops=3)
Group Key: (m.application_id)::text, (m.family_id)::text
-> Seq Scan on app_memeber m (cost=0.00..17398.00 rows=1000000 width=12) (actual time=0.047..247.902 rows=1000000 loops=3
)
-> Parallel Hash (cost=13648.00..13648.00 rows=138889 width=6) (actual time=231.705..231.706 rows=300000 loops=3)
Buckets: 1048576 (originally 524288) Batches: 1 (originally 1) Memory Usage: 47552kB
-> Parallel Seq Scan on app_family2 a2 (cost=0.00..13648.00 rows=138889 width=6) (actual time=0.039..100.756 rows=300000 loops=3)
Filter: ((family_id)::integer < 900000)
Rows Removed by Filter: 33334
Planning Time: 0.359 ms
Execution Time: 2911.146 ms
(22 rows)

语句2 三张表join

为了保证语句的一致性,三张表的join顺序保持和语句1的执行计划中的顺序一致,join总计用时1476.651 ms

explain analyze select a1.application_id,a1.family_id from app_family a1

left join app_memeber m on a1.application_id = m.application_id LEFT JOIN app_family2 a2 ON m.family_id = a2.family_id

WHERE a1.family_id >1000 AND a2.family_id < 900000

 Gather  (cost=32990.22..64898.93 rows=111111 width=12) (actual time=993.681..1436.895 rows=898999 loops=1)
Workers Planned: 2
Workers Launched: 2
-> Parallel Hash Join (cost=31990.22..52787.83 rows=46296 width=12) (actual time=982.512..1241.385 rows=299666 loops=3)
Hash Cond: ((m.application_id)::text = (a1.application_id)::text)
-> Parallel Hash Join (cost=15927.11..34245.98 rows=138889 width=6) (actual time=377.411..635.945 rows=300000 loops=3)
Hash Cond: ((m.family_id)::text = (a2.family_id)::text)
-> Parallel Seq Scan on app_memeber m (cost=0.00..11564.67 rows=416667 width=12) (actual time=0.034..59.470 rows=333333 loops=3)
-> Parallel Hash (cost=13648.00..13648.00 rows=138889 width=6) (actual time=232.286..232.287 rows=300000 loops=3)
Buckets: 131072 (originally 131072) Batches: 16 (originally 8) Memory Usage: 3296kB
-> Parallel Seq Scan on app_family2 a2 (cost=0.00..13648.00 rows=138889 width=6) (actual time=0.030..104.370 rows=300000 loops=
3)
Filter: ((family_id)::integer < 900000)
Rows Removed by Filter: 33334
-> Parallel Hash (cost=13648.00..13648.00 rows=138889 width=12) (actual time=271.185..271.185 rows=333000 loops=3)
Buckets: 131072 (originally 131072) Batches: 16 (originally 8) Memory Usage: 4032kB
-> Parallel Seq Scan on app_family a1 (cost=0.00..13648.00 rows=138889 width=12) (actual time=0.091..129.188 rows=333000 loops=3)
Filter: ((family_id)::integer > 1000)
Rows Removed by Filter: 333
Planning Time: 0.140 ms
Execution Time: 1475.305 ms
(20 rows) Time: 1476.651 ms (00:01.477)

总结三张表场景

在三张表的场景下exist用时2911.146 ms ,join用时1476.651 ms 可见 join的顺序明显优于exist。

在三张表的场景下可以看到,针对中间表appmember扫描时, exist语句用到HashAggregate 并做了 Group Key,所以导致exist 执行时间增加。如果work_mem 配置不合适时间会更长。

exist和left join 性能对比的更多相关文章

  1. Go 字符串连接+=与strings.Join性能对比

    Go字符串连接 对于字符串的连接大致有两种方式: 1.通过+号连接 func StrPlus1(a []string) string { var s, sep string for i := 0; i ...

  2. 自己写的轻量级PHP框架trig与laravel5.1,yii2性能对比

    看了下当前最热门的php开发框架,想对比一下自己写的框架与这些框架的性能对比.先看下当前流行框架的投票情况. 看结果对比,每个测试脚本做了一个数据库的联表查询并进行print_r输出,查询的sql语句 ...

  3. SQL点滴10—使用with语句来写一个稍微复杂sql语句,附加和子查询的性能对比

    原文:SQL点滴10-使用with语句来写一个稍微复杂sql语句,附加和子查询的性能对比 今天偶尔看到sql中也有with关键字,好歹也写了几年的sql语句,居然第一次接触,无知啊.看了一位博主的文章 ...

  4. python3下multiprocessing、threading和gevent性能对比----暨进程池、线程池和协程池性能对比

    python3下multiprocessing.threading和gevent性能对比----暨进程池.线程池和协程池性能对比   标签: python3 / 线程池 / multiprocessi ...

  5. SQL Server-聚焦IN VS EXISTS VS JOIN性能分析(十九)

    前言 本节我们开始讲讲这一系列性能比较的终极篇IN VS EXISTS VS JOIN的性能分析,前面系列有人一直在说场景不够,这里我们结合查询索引列.非索引列.查询小表.查询大表来综合分析,简短的内 ...

  6. [原] KVM 环境下MySQL性能对比

    KVM 环境下MySQL性能对比 标签(空格分隔): Cloud2.0 [TOC] 测试目的 对比MySQL在物理机和KVM环境下性能情况 压测标准 压测遵循单一变量原则,所有的对比都是只改变一个变量 ...

  7. 浅谈C++之冒泡排序、希尔排序、快速排序、插入排序、堆排序、基数排序性能对比分析之后续补充说明(有图有真相)

    如果你觉得我的有些话有点唐突,你不理解可以想看看前一篇<C++之冒泡排序.希尔排序.快速排序.插入排序.堆排序.基数排序性能对比分析>. 这几天闲着没事就写了一篇<C++之冒泡排序. ...

  8. Java--Stream,NIO ByteBuffer,NIO MappedByteBuffer性能对比

    目前Java中最IO有多种文件读取的方法,本文章对比Stream,NIO ByteBuffer,NIO MappedByteBuffer的性能,让我们知道到底怎么能写出性能高的文件读取代码. pack ...

  9. C正则库做DNS域名验证时的性能对比

    C正则库做DNS域名验证时的性能对比   本文对C的正则库regex和pcre在做域名验证的场景下做评测. 验证DNS域名的正则表达式为: "^[0-9a-zA-Z_-]+(\\.[0-9a ...

  10. 开发语言性能对比,C++、Java、Python、LUA、TCC

    一直想做开发语言性能对比,刚好有时间都做了给大家参考一下, 编译类:C++和Java表现还不错 脚本类:TCC脚本动态运行C语言,性能比其他脚本快好多... 想玩TCC的同学下载测试包,TCC目录下修 ...

随机推荐

  1. XML和JSON的比较

    XML和JSON的比较 XML与JSON都可以用来描述或者存储数据,两者都有各自的优点,使用场景取决于需求. 描述 XML 可扩展标记语言Extensible Markup Language,是一种用 ...

  2. 微信小程序获取本日、本周、本月、本年时间段

    原文链接 https://cslaoxu.vip/110.html 说明 最近需要用到统计不同时间段内的记录数,所以找了一下现成的工具类.下面就演示一下如何引用到实际项目中. 详细用法请参考:http ...

  3. nginx新增conf文件

    说明 最近租了一台美国vps,通过nginx反向代理设置搞谷歌镜像.因为BxxDx搜索太垃圾.中间涉及到添加反向代理配置. 操作步骤 1.在conf.d文件下新增配置 cd /etc/nginx/co ...

  4. html+css:小米顶部菜单+二级菜单

    1.源码 <!DOCTYPE html> <html lang="en"> <head> <meta charset="UTF- ...

  5. Git 分支管理参考模型

    一个值得参考的Git分支管理模型如下: master 生产主分支,发布到生产环境使用这个分支,由hotfix或者release分支合并过来,不直接提交代码. release 预发布分支, 基于feat ...

  6. django学习第八天--多表操作删除和修改,子查询连表查询,双下划线跨表查询,聚合查询,分组查询,F查询,Q查询

    orm多条操作 删除和修改 修改 在一对一和一对多关系时,和单表操作是一样的 一对一 一个作者对应一个信息 ad_obj = models.AuthorDetail.objects.get(id=1) ...

  7. 【Java复健指南07】OOP中级02-重写与多态思想

    前情提要:https://www.cnblogs.com/DAYceng/category/2227185.html 重写 注意事项和使用细节 方法重写也叫方法覆法,需要满足下面的条件 1.子类的方法 ...

  8. 【Azure Key Vault】使用Azure CLI获取Key Vault 机密遇见问题后使用curl命令来获取机密内容

    问题描述 在使用Azure Key Vault的过程中,遇见无法获取机密信息,在不方便直接写代码的情况下,快速使用Azure CLI指令来验证当前使用的认证是否可以获取到正确的机密值. 使用CLI的指 ...

  9. 【Azure 应用服务】用App Service部署运行 Vue.js 编写的项目,应该怎么部署运行呢?

    问题描述 用App Service部署运行 Vue.js 编写的项目,应该怎么部署运行呢? 问题解答 VUE通常是运行在客户端侧的JS框架. App Service 在这种场景中是以静态文件的形式提供 ...

  10. C++ //count_if //按条件统计元素个数 //自定义和 内置

    1 //按条件统计元素个数 2 //count_if 3 4 #include <iostream> 5 #include<string> 6 #include<vect ...