WITH RECURSIVE and MySQL

If you have been using certain DBMSs, or reading recent versions of the SQL standard, you are probably aware of the so-called “WITH clause” of SQL. Some call it Subquery Factoring. Others call it Common Table Expression. A form of the WITH CLAUSE, “WITH RECURSIVE”, allows to design a recursive query: a query which repeats itself again and again, each time using the results of the previous iteration. This can be quite useful to produce reports based on hierarchical data. And thus is an alternative to Oracle’s CONNECT BY. MySQL does not natively support WITH RECURSIVE, but it is easy to emulate it with a generic, reusable stored procedure. Read the full article here…

http://guilhembichot.blogspot.co.uk/2013/11/with-recursive-and-mysql.html

If you have been using certain DBMSs, or reading recent versions of the SQL standard, you are probably aware of the so-called "WITH clause" of SQL. Some call it Subquery Factoring. Others call it Common Table Expression. In its simplest form, this feature is a kind of "boosted derived table".
Assume that a table T1 has three columns:

CREATE TABLE T1(

YEAR INT, # 2000, 2001, 2002 ...

MONTH INT, # January, February, ...

SALES INT # how much we sold on that month of that year

);

Now I want to know the sales trend (increase/decrease), year after year:

SELECT D1.YEAR, (CASE WHEN D1.S>D2.S THEN 'INCREASE' ELSE 'DECREASE' END) AS TREND

FROM

  (SELECT YEAR, SUM(SALES) AS S FROM T1 GROUP BY YEAR) AS D1,

  (SELECT YEAR, SUM(SALES) AS S FROM T1 GROUP BY YEAR) AS D2

WHERE D1.YEAR = D2.YEAR-1;

Both derived tables are based on the same subquery text, but usually a DBMS is not smart enough to recognize it. Thus, it will evaluate "SELECT YEAR, SUM(SALES)... GROUP BY YEAR" twice! A first time to fill D1, a second time to fill D2. This limitation is sometimes stated as "it's not possible to refer to a derived table twice in the same query". Such double evaluation can lead to a serious performance problem. Using WITH, this limitation does not exist, and the following statement evaluates the subquery only once:

WITH D AS (SELECT YEAR, SUM(SALES) AS S FROM T1 GROUP BY YEAR)

SELECT D1.YEAR, (CASE WHEN D1.S>D2.S THEN 'INCREASE' ELSE 'DECREASE' END) AS TREND

FROM

 D AS D1,

 D AS D2

WHERE D1.YEAR = D2.YEAR-1;

This already demonstrates one benefit of WITH. In MySQL, WITH is not yet supported. But it can be emulated with a view:

CREATE VIEW D AS (SELECT YEAR, SUM(SALES) AS S FROM T1 GROUP BY YEAR);

SELECT D1.YEAR, (CASE WHEN D1.S>D2.S THEN 'INCREASE' ELSE 'DECREASE' END) AS TREND

FROM

 D AS D1,

 D AS D2

WHERE D1.YEAR = D2.YEAR-1;

DROP VIEW D;

Instead of a view, I could as well create D as a normal table. But not as a temporary table, because in MySQL a temporary table cannot be referred twice in the same query, as mentioned in the manual.
After this short introduction, showing the simplest form of WITH, I would like to turn to the more complex form of WITH: the RECURSIVE form. According to the SQL standard, to use the recursive form, you should write WITH RECURSIVE. However, looking at some other DBMSs, they seem to not require the RECURSIVE word. WITH RECURSIVE is a powerful construct. For example, it can do the same job as Oracle's CONNECT BY clause (you can check out some example conversions between both constructs). Let's walk through an example, to understand what WITH RECURSIVE does.
Assume you have a table of employees (this is a very classical example of WITH RECURSIVE):

CREATE TABLE EMPLOYEES (

ID INT PRIMARY KEY,

NAME VARCHAR(100),

MANAGER_ID INT,

INDEX (MANAGER_ID),

FOREIGN KEY (MANAGER_ID) REFERENCES EMPLOYEES(ID)

);

INSERT INTO EMPLOYEES VALUES

(333, "Yasmina", NULL),

(198, "John", 333),

(29, "Pedro", 198),

(4610, "Sarah", 29),

(72, "Pierre", 29),

(692, "Tarek", 333);

In other words, Yasmina is CEO, John and Tarek report to her. Pedro reports to John, Sarah and Pierre report to Pedro. In a big company, they would be thousands of rows in this table.
Now, let's say that you would like to know, for each employee: "how many people are, directly and indirectly, reporting to him/her"? Here is how I would do it. First, I would make a list of people who are not managers: with a subquery I get the list of all managers, and using NOT IN (subquery) I get the list of all non-managers:

SELECT ID, NAME, MANAGER_ID, 0 AS REPORTS

FROM EMPLOYEES

WHERE ID NOT IN (SELECT MANAGER_ID FROM EMPLOYEES WHERE MANAGER_ID IS NOT NULL);

Then I would insert the results into a new table named EMPLOYEES_EXTENDED; EXTENDED stands for "extended with more information", the new information being the fourth column named REPORTS: it is a count of people who are reporting directly or indirectly to the employee. Because we have listed people who are not managers, they have a value of 0 in the REPORTS column. Then, we can produce the rows for "first level" managers (the direct managers of non-managers):

SELECT M.ID, M.NAME, M.MANAGER_ID, SUM(1+E.REPORTS) AS REPORTS

FROM EMPLOYEES M JOIN EMPLOYEES_EXTENDED E ON M.ID=E.MANAGER_ID

GROUP BY M.ID, M.NAME, M.MANAGER_ID;

Explanation: for a row of M (that is, for an employee), the JOIN will produce zero or more rows, one per non-manager directly reporting to the employee. Each such non-manager contributes to the value of REPORTS for his manager, through two numbers: 1 (the non-manager himself), and the number of direct/indirect reports of the non-manager (i.e. the value of REPORTS for the non-manager). Then I would empty EMPLOYEES_EXTENDED, and fill it with the rows produced just above, which describe the first level managers. Then the same query should be run again, and it would produce information about the "second level" managers. And so on. Finally, at one point Yasmina will be the only row of EMPLOYEES_EXTENDED, and when we run the above SELECT again, the JOIN will produce no rows, because E.MANAGER_ID will be NULL (she's the CEO). We are done.
It's time for a recap: EMPLOYEES_EXTENDED has been a kind of "temporary buffer", which has successively held non-managers, first level managers, second level managers, etc. We have used recursion. The answer to the original problem is: the union of all the successive content of EMPLOYEES_EXTENDED. Non-managers have been the start of the recursion, which is usually called "the anchor member" or "the seed". The SELECT query which moves from one step of recursion to the next one, is the "recursive member". The complete statement looks like this:

WITH RECURSIVE

# The temporary buffer, also used as UNION result:

EMPLOYEES_EXTENDED

AS

(

  # The seed:

  SELECT ID, NAME, MANAGER_ID, 0 AS REPORTS

  FROM EMPLOYEES

  WHERE ID NOT IN (SELECT MANAGER_ID FROM EMPLOYEES WHERE MANAGER_ID IS NOT NULL)

UNION ALL

  # The recursive member:

  SELECT M.ID, M.NAME, M.MANAGER_ID, SUM(1+E.REPORTS) AS REPORTS

  FROM EMPLOYEES M JOIN EMPLOYEES_EXTENDED E ON M.ID=E.MANAGER_ID

  GROUP BY M.ID, M.NAME, M.MANAGER_ID

)

# what we want to do with the complete result (the UNION):

SELECT * FROM EMPLOYEES_EXTENDED;

MySQL does not yet support WITH RECURSIVE, but it is possible to code a generic stored procedure which can easily emulate it. Here is how you would call it:

CALL WITH_EMULATOR(

"EMPLOYEES_EXTENDED",

"

  SELECT ID, NAME, MANAGER_ID, 0 AS REPORTS

  FROM EMPLOYEES

  WHERE ID NOT IN (SELECT MANAGER_ID FROM EMPLOYEES WHERE MANAGER_ID IS NOT NULL)

",

"

  SELECT M.ID, M.NAME, M.MANAGER_ID, SUM(1+E.REPORTS) AS REPORTS

  FROM EMPLOYEES M JOIN EMPLOYEES_EXTENDED E ON M.ID=E.MANAGER_ID

  GROUP BY M.ID, M.NAME, M.MANAGER_ID

",

"SELECT * FROM EMPLOYEES_EXTENDED",

0,

""

);

You can recognize, as arguments of the stored procedure, every member of the WITH standard syntax: name of the temporary buffer, query for the seed, query for the recursive member, and what to do with the complete result. The last two arguments - 0 and the empty string - are details which you can ignore for now.
Here is the result returned by this stored procedure:

+------+---------+------------+---------+

| ID   | NAME    | MANAGER_ID | REPORTS |

+------+---------+------------+---------+

|   72 | Pierre  |         29 |       0 |

|  692 | Tarek   |        333 |       0 |

| 4610 | Sarah   |         29 |       0 |

|   29 | Pedro   |        198 |       2 |

|  333 | Yasmina |       NULL |       1 |

|  198 | John    |        333 |       3 |

|  333 | Yasmina |       NULL |       4 |

+------+---------+------------+---------+

7 rows in set

Notice how Pierre, Tarek and Sarah have zero reports, Pedro has two, which looks correct... However, Yasmina appears in two rows! Odd? Yes and no. Our algorithm starts from non-managers, the "leaves" of the tree (Yasmina being the root of the tree). Then our algorithm looks at first level managers, the direct parents of leaves. Then at second level managers. But Yasmina is both a first level manager (of the nonmanager Tarek) and a third level manager (of the nonmanagers Pierre, Tarek and Sarah). That's why she appears twice in the final result: once for the "tree branch" which ends at leaf Tarek, once for the tree branch which ends at leaves Pierre, Tarek and Sarah. The first tree branch contributes 1 direct/indirect report. The second tree branch contributes 4. The right number, which we want, is the sum of the two: 5. Thus we just need to change the final query, in the CALL:

CALL WITH_EMULATOR(

"EMPLOYEES_EXTENDED",

"

  SELECT ID, NAME, MANAGER_ID, 0 AS REPORTS

  FROM EMPLOYEES

  WHERE ID NOT IN (SELECT MANAGER_ID FROM EMPLOYEES WHERE MANAGER_ID IS NOT NULL)

",

"

  SELECT M.ID, M.NAME, M.MANAGER_ID, SUM(1+E.REPORTS) AS REPORTS

  FROM EMPLOYEES M JOIN EMPLOYEES_EXTENDED E ON M.ID=E.MANAGER_ID

  GROUP BY M.ID, M.NAME, M.MANAGER_ID

",

"

  SELECT ID, NAME, MANAGER_ID, SUM(REPORTS)

  FROM EMPLOYEES_EXTENDED

  GROUP BY ID, NAME, MANAGER_ID

",

0,

""

);

And here is finally the proper result:

+------+---------+------------+--------------+

| ID   | NAME    | MANAGER_ID | SUM(REPORTS) |

+------+---------+------------+--------------+

|   29 | Pedro   |        198 |            2 |

|   72 | Pierre  |         29 |            0 |

|  198 | John    |        333 |            3 |

|  333 | Yasmina |       NULL |            5 |

|  692 | Tarek   |        333 |            0 |

| 4610 | Sarah   |         29 |            0 |

+------+---------+------------+--------------+

6 rows in set

Let's finish by showing the body of the stored procedure. You will notice that it does heavy use of dynamic SQL, thanks to prepared statements. Its body does not depend on the particular problem to solve, it's reusable as-is for other WITH RECURSIVE use cases. I have added comments inside the body, so it should be self-explanatory. If it's not, feel free to drop a comment on this post, and I will explain further. Note that it uses temporary tables internally, and the first thing it does is dropping any temporary tables with the same names.

# Usage: the standard syntax:

#   WITH RECURSIVE recursive_table AS

#    (initial_SELECT

#     UNION ALL

#     recursive_SELECT)

#   final_SELECT;

# should be translated by you to

# CALL WITH_EMULATOR(recursive_table, initial_SELECT, recursive_SELECT,

#                    final_SELECT, 0, "").

# ALGORITHM:

# 1) we have an initial table T0 (actual name is an argument

# "recursive_table"), we fill it with result of initial_SELECT.

# 2) We have a union table U, initially empty.

# 3) Loop:

#   add rows of T0 to U,

#   run recursive_SELECT based on T0 and put result into table T1,

#   if T1 is empty

#      then leave loop,

#      else swap T0 and T1 (renaming) and empty T1

# 4) Drop T0, T1

# 5) Rename U to T0

# 6) run final select, send relult to client

# This is for *one* recursive table.

# It would be possible to write a SP creating multiple recursive tables.

delimiter |

CREATE PROCEDURE WITH_EMULATOR(

recursive_table varchar(100), # name of recursive table

initial_SELECT varchar(65530), # seed a.k.a. anchor

recursive_SELECT varchar(65530), # recursive member

final_SELECT varchar(65530), # final SELECT on UNION result

max_recursion int unsigned, # safety against infinite loop, use 0 for default

create_table_options varchar(65530) # you can add CREATE-TABLE-time options

# to your recursive_table, to speed up initial/recursive/final SELECTs; example:

# "(KEY(some_column)) ENGINE=MEMORY"

)

BEGIN

  declare new_rows int unsigned;

  declare show_progress int default 0; # set to 1 to trace/debug execution

  declare recursive_table_next varchar(120);

  declare recursive_table_union varchar(120);

  declare recursive_table_tmp varchar(120);

  set recursive_table_next  = concat(recursive_table, "_next");

  set recursive_table_union = concat(recursive_table, "_union");

  set recursive_table_tmp   = concat(recursive_table, "_tmp");

  # Cleanup any previous failed runs

  SET @str =

    CONCAT("DROP TEMPORARY TABLE IF EXISTS ", recursive_table, ",",

    recursive_table_next, ",", recursive_table_union,

    ",", recursive_table_tmp);

  PREPARE stmt FROM @str;

  EXECUTE stmt;

 # If you need to reference recursive_table more than

  # once in recursive_SELECT, remove the TEMPORARY word.

  SET @str = # create and fill T0

    CONCAT("CREATE TEMPORARY TABLE ", recursive_table, " ",

    create_table_options, " AS ", initial_SELECT);

  PREPARE stmt FROM @str;

  EXECUTE stmt;

  SET @str = # create U

    CONCAT("CREATE TEMPORARY TABLE ", recursive_table_union, " LIKE ", recursive_table);

  PREPARE stmt FROM @str;

  EXECUTE stmt;

  SET @str = # create T1

    CONCAT("CREATE TEMPORARY TABLE ", recursive_table_next, " LIKE ", recursive_table);

  PREPARE stmt FROM @str;

  EXECUTE stmt;

  if max_recursion = 0 then

    set max_recursion = 100; # a default to protect the innocent

  end if;

  recursion: repeat

    # add T0 to U (this is always UNION ALL)

    SET @str =

      CONCAT("INSERT INTO ", recursive_table_union, " SELECT * FROM ", recursive_table);

    PREPARE stmt FROM @str;

    EXECUTE stmt;

    # we are done if max depth reached

    set max_recursion = max_recursion - 1;

    if not max_recursion then

      if show_progress then

        select concat("max recursion exceeded");

      end if;

      leave recursion;

    end if;

    # fill T1 by applying the recursive SELECT on T0

    SET @str =

      CONCAT("INSERT INTO ", recursive_table_next, " ", recursive_SELECT);

    PREPARE stmt FROM @str;

    EXECUTE stmt;

    # we are done if no rows in T1

    select row_count() into new_rows;

    if show_progress then

      select concat(new_rows, " new rows found");

    end if;

    if not new_rows then

      leave recursion;

    end if;

    # Prepare next iteration:

    # T1 becomes T0, to be the source of next run of recursive_SELECT,

    # T0 is recycled to be T1.

    SET @str =

      CONCAT("ALTER TABLE ", recursive_table, " RENAME ", recursive_table_tmp);

    PREPARE stmt FROM @str;

    EXECUTE stmt;

    # we use ALTER TABLE RENAME because RENAME TABLE does not support temp tables

    SET @str =

      CONCAT("ALTER TABLE ", recursive_table_next, " RENAME ", recursive_table);

    PREPARE stmt FROM @str;

    EXECUTE stmt;

    SET @str =

      CONCAT("ALTER TABLE ", recursive_table_tmp, " RENAME ", recursive_table_next);

    PREPARE stmt FROM @str;

    EXECUTE stmt;

    # empty T1

    SET @str =

      CONCAT("TRUNCATE TABLE ", recursive_table_next);

    PREPARE stmt FROM @str;

    EXECUTE stmt;

  until 0 end repeat;

  # eliminate T0 and T1

  SET @str =

    CONCAT("DROP TEMPORARY TABLE ", recursive_table_next, ", ", recursive_table);

  PREPARE stmt FROM @str;

  EXECUTE stmt;

  # Final (output) SELECT uses recursive_table name

  SET @str =

    CONCAT("ALTER TABLE ", recursive_table_union, " RENAME ", recursive_table);

  PREPARE stmt FROM @str;

  EXECUTE stmt;

  # Run final SELECT on UNION

  SET @str = final_SELECT;

  PREPARE stmt FROM @str;

  EXECUTE stmt;

  # No temporary tables may survive:

  SET @str =

    CONCAT("DROP TEMPORARY TABLE ", recursive_table);

  PREPARE stmt FROM @str;

  EXECUTE stmt;

  # We are done :-)

END|

delimiter ;

In the SQL Standard, WITH RECURSIVE allows some nice additional tweaks (depth-first or breadth-first ordering, cycle detection). In future posts I will show how to emulate them too.

WITH RECURSIVE and MySQL的更多相关文章

MySQL: Tree-Hierarchical query
http://dba.stackexchange.com/questions/30021/mysql-tree-hierarchical-query No problem. ...
如何将MySQL help contents的内容有层次的输出
经常会遇到这种情况,在一个不能上网的环境通过MySQL客户端登录数据库,想执行一个操作,却忘了操作的具体语法,各种不方便. 其实,MySQL数据库内置了帮助文档,通过help contents即可查看 ...
PHP定时备份MySQL，mysqldump语法大全
几个常用操作: 1.备份 # 只导出表结构 d:/PHP/xampp/mysql/bin/mysqldump -h127.0.0.1 -P3306 -uroot -p123456 snsgou_sns ...
Using Recursive Common table expressions to represent Tree structures
http://www.postgresonline.com/journal/archives/131-Using-Recursive-Common-table-expressions-to-repre ...
MySql Error: Can't update table in stored function/trigger
MySql Error: Can't update table in stored function/trigger because it is already used by statement w ...
MySQL入门笔记
MySQL入门笔记版本选择: 5.x.20 以上版本比较稳定一．MySQL的三种安装方式: 安装MySQL的方式常见的有三种: · rpm包形式 · 通用二进制 ...
QSqlDatabase的进一步封装（多线程支持+更加简单的操作）——同时支持MySQL, SQL Server和Sqlite
开发背景: 1.直接用QSqlDatabase我觉得太麻烦了: 2.对于某些数据库,多个线程同时使用一个QSqlDatabase的时候会崩溃: 3.这段时间没什么干货放出来觉得浑身不舒服,就想写一个. ...
MySQL · 引擎特性 · InnoDB 同步机制
前言现代操作系统以及硬件基本都支持并发程序,而在并发程序设计中,各个进程或者线程需要对公共变量的访问加以制约,此外,不同的进程或者线程需要协同工作以完成特征的任务,这就需要一套完善的同步机制,在Li ...
Virtual Box虚拟机Ubuntu18.X系统安装及Mysql基本开发配置
Linux简介什么是 Linux? Linux:世界上不仅只有一个 Windows 操作系统,还有 Linux.mac.Unix 等操作系统.桌面操作系统下 Windows 是霸主,而 Linux ...

随机推荐

php版的redis操作库predis操作大全
转载于:http://www.itxuexiwang.com/a/shujukujishu/redis/2016/0216/146.html predis是php连接redis的操作库,由于它完全使用 ...
爱上MVC系列~过滤器实现对响应流的处理
回到目录 MVC的过滤器相信大家都用过,一般用来作权限控制,因为它可以监视你的Action从进入到最后View的渲染,整个过程ActionFilter这个过滤器都参与了,而这给我们的开发带来了更多的好 ...
Servlet字符编码过滤器，实现图书信息的添加功能，避免产生文字乱码现象的产生
同样的代码,网上可以找到和我一模一样的代码和配置,比我的更加详细,但是我重新写一个博客的原因自是把错误的原因写出来,因为这就是个坑,我弄了一天,希望对你们有所帮助.只为初学者发现错误不知道怎么解决有所 ...
Liferay7 BPM门户开发之44: 集成Activiti展示流程列表
处理依赖关系集成Activiti之前,必须搞清楚其中的依赖关系,才能在Gradle里进行配置. 依赖关系: 例如,其中activiti-engine依赖于activiti-bpmn-converte ...
lufylegend游戏引擎
lufylegend游戏引擎介绍:click 这个链接我觉得已经很详细的介绍了这个引擎. 所以以下我只说说一些简单的游戏代码过程. 首先从canvas做游戏叙述起: 这是一个让人很熟悉的简单小游戏,网 ...
asp.net的简易的参数化查询
protected void btnInsert_Click(object sender, EventArgs e) { string sql = "insert into contactg ...
SQL Server的日期和时间类型
Sql Server使用 Date 表示日期,time表示时间,使用datetime和datetime2表示日期和时间. 1,秒的精度是指使用多少位小数表示秒 DateTime数据类型秒的精度是3,D ...
LLBL Gen Pro 4.2 Lite 免费的对象关系映射开发框架与工具
LLBL Gen Pro是一款优秀的对象关系映射开发框架,自2003年发布以来,一直有广泛的客户群.LLBL Gen Pro有几个标志性的版本,2.5/2.6是一个很稳定的版本,公司的一些旧的项目仍然 ...
JS中实现数组和对象的深拷贝和浅拷贝
数组的拷贝 > 数组的深拷贝,两层 var arr = [[1,2,3],[4,5,6],[7,8,9]]; var arr2 = []; 循环第一层数组 for(var i=0,len=arr ...
JSP网站开发基础总结《十一》
继上一篇关于过滤器连总结后,本篇为大家详细介绍一下过滤器中过滤规则的dispatcher属性的使用,在servlet2.5中dispatcher的属性有四种,其中上一篇已经为大家介绍了error属性的 ...

WITH RECURSIVE and MySQL

WITH RECURSIVE and MySQL

WITH RECURSIVE and MySQL的更多相关文章

随机推荐

热门专题