Spark修炼之道（高级篇）——Spark源代码阅读：第十二节 Spark SQL 处理流程分析

作者：周志湖

以下的代码演示了通过Case Class进行表Schema定义的样例：

// sc is an existing SparkContext.

val sqlContext = new org.apache.spark.sql.SQLContext(sc)

// this is used to implicitly convert an RDD to a DataFrame.

import sqlContext.implicits._

// Define the schema using a case class.

// Note: Case classes in Scala 2.10 can support only up to 22 fields. To work around this limit,

// you can use custom classes that implement the Product interface.

case class Person(name: String, age: Int)

// Create an RDD of Person objects and register it as a table.

val people = sc.textFile("examples/src/main/resources/people.txt").map(_.split(",")).map(p => Person(p(0), p(1).trim.toInt)).toDF()

people.registerTempTable("people")

// SQL statements can be run by using the sql methods provided by sqlContext.

val teenagers = sqlContext.sql("SELECT name, age FROM people WHERE age >= 13 AND age <= 19")

// The results of SQL queries are DataFrames and support all the normal RDD operations.

// The columns of a row in the result can be accessed by field index:

teenagers.map(t => "Name: " + t(0)).collect().foreach(println)

// or by field name:

teenagers.map(t => "Name: " + t.getAs[String]("name")).collect().foreach(println)

// row.getValuesMap[T] retrieves multiple columns at once into a Map[String, T]

teenagers.map(_.getValuesMap[Any](List("name", "age"))).collect().foreach(println)

// Map("name" -> "Justin", "age" -> 19)

（1）sql方法返回DataFrame

  def sql(sqlText: String): DataFrame = {

    DataFrame(this, parseSql(sqlText))

  }

当中parseSql(sqlText)方法生成对应的LogicalPlan得到，该方法源代码例如以下：

//依据传入的sql语句，生成LogicalPlan

protected[sql] def parseSql(sql: String): LogicalPlan = ddlParser.parse(sql, false)

ddlParser对象定义例如以下：

protected[sql] val sqlParser = new SparkSQLParser(getSQLDialect().parse(_))

protected[sql] val ddlParser = new DDLParser(sqlParser.parse(_))

（2）然后调用DataFrame的apply方法

private[sql] object DataFrame {

  def apply(sqlContext: SQLContext, logicalPlan: LogicalPlan): DataFrame = {

    new DataFrame(sqlContext, logicalPlan)

  }

}

能够看到，apply方法參数有两个，各自是SQLContext和LogicalPlan，调用的是DataFrame的构造方法，详细源代码例如以下：

//DataFrame构造方法。该构造方法会自己主动对LogicalPlan进行分析，然后返回QueryExecution对象

def this(sqlContext: SQLContext, logicalPlan: LogicalPlan) = {

    this(sqlContext, {

      val qe = sqlContext.executePlan(logicalPlan)

      //推断是否已经创建。假设是则抛异常

      if (sqlContext.conf.dataFrameEagerAnalysis) {

        qe.assertAnalyzed()  // This should force analysis and throw errors if there are any

      }

      qe

    })

  }

（3）val qe = sqlContext.executePlan(logicalPlan) 返回QueryExecution， sqlContext.executePlan方法源代码例如以下：

protected[sql] def executePlan(plan: LogicalPlan) =

    new sparkexecution.QueryExecution(this, plan)

QueryExecution类中表达了Spark运行SQL的主要工作流程，详细例如以下

class QueryExecution(val sqlContext: SQLContext, val logical: LogicalPlan) {

  @VisibleForTesting

  def assertAnalyzed(): Unit = sqlContext.analyzer.checkAnalysis(analyzed)

  lazy val analyzed: LogicalPlan = sqlContext.analyzer.execute(logical)

  lazy val withCachedData: LogicalPlan = {

    assertAnalyzed()

    sqlContext.cacheManager.useCachedData(analyzed)

  }

  lazy val optimizedPlan: LogicalPlan = sqlContext.optimizer.execute(withCachedData)

  // TODO: Don't just pick the first one...

  lazy val sparkPlan: SparkPlan = {

    SparkPlan.currentContext.set(sqlContext)

    sqlContext.planner.plan(optimizedPlan).next()

  }

  // executedPlan should not be used to initialize any SparkPlan. It should be

  // only used for execution.

  lazy val executedPlan: SparkPlan = sqlContext.prepareForExecution.execute(sparkPlan)

  /** Internal version of the RDD. Avoids copies and has no schema */

  //调用toRDD方法运行任务将结果转换为RDD

  lazy val toRdd: RDD[InternalRow] = executedPlan.execute()

  protected def stringOrError[A](f: => A): String =

    try f.toString catch { case e: Throwable => e.toString }

  def simpleString: String = {

    s"""== Physical Plan ==

       |${stringOrError(executedPlan)}

      """.stripMargin.trim

  }

  override def toString: String = {

    def output =

      analyzed.output.map(o => s"${o.name}: ${o.dataType.simpleString}").mkString(", ")

    s"""== Parsed Logical Plan ==

       |${stringOrError(logical)}

       |== Analyzed Logical Plan ==

       |${stringOrError(output)}

       |${stringOrError(analyzed)}

       |== Optimized Logical Plan ==

       |${stringOrError(optimizedPlan)}

       |== Physical Plan ==

       |${stringOrError(executedPlan)}

       |Code Generation: ${stringOrError(executedPlan.codegenEnabled)}

    """.stripMargin.trim

  }

}

能够看到，SQL的运行流程为

1.Parsed Logical Plan：LogicalPlan

2.Analyzed Logical Plan：

lazy val analyzed: LogicalPlan = sqlContext.analyzer.execute(logical)

3.Optimized Logical Plan：lazy val optimizedPlan: LogicalPlan = sqlContext.optimizer.execute(withCachedData)

4. Physical Plan：lazy val executedPlan: SparkPlan = sqlContext.prepareForExecution.execute(sparkPlan)

能够调用results.queryExecution方法查看，代码例如以下：

scala> results.queryExecution

res1: org.apache.spark.sql.SQLContext#QueryExecution =

== Parsed Logical Plan ==

'Project [unresolvedalias('name)]

 'UnresolvedRelation [people], None

== Analyzed Logical Plan ==

name: string

Project [name#0]

 Subquery people

  LogicalRDD [name#0,age#1], MapPartitionsRDD[4] at createDataFrame at <console>:47

== Optimized Logical Plan ==

Project [name#0]

 LogicalRDD [name#0,age#1], MapPartitionsRDD[4] at createDataFrame at <console>:47

== Physical Plan ==

TungstenProject [name#0]

 Scan PhysicalRDD[name#0,age#1]

Code Generation: true

（4）然后调用DataFrame的主构造器完毕DataFrame的构造

class DataFrame private[sql](

    @transient val sqlContext: SQLContext,

    @DeveloperApi @transient val queryExecution: QueryExecution) extends Serializable

（5）

当调用DataFrame的collect等方法时，便会触发运行executedPlan

  def collect(): Array[Row] = withNewExecutionId {

    queryExecution.executedPlan.executeCollect()

  }

比如：

scala> results.collect

res6: Array[org.apache.spark.sql.Row] = Array([Michael], [Andy], [Justin])

总体流程图例如以下：

Spark修炼之道（高级篇）——Spark源代码阅读：第十二节 Spark SQL 处理流程分析的更多相关文章

Spark修炼之道——Spark学习路线、课程大纲
课程内容 Spark修炼之道(基础篇)--Linux基础(15讲).Akka分布式编程(8讲) Spark修炼之道(进阶篇)--Spark入门到精通(30讲) Spark修炼之道(实战篇)--Spar ...
【转】【技术博客】Spark性能优化指南——高级篇
http://mp.weixin.qq.com/s?__biz=MjM5NjQ5MTI5OA==&mid=2651745207&idx=1&sn=3d70d59cede236e ...
Spark性能优化指南——高级篇
本文转载自:https://tech.meituan.com/spark-tuning-pro.html 美团技术点评团队) Spark性能优化指南——高级篇李雪蕤 ·2016-05-12 14:4 ...
Spark性能优化指南-高级篇(spark shuffle)
Spark性能优化指南-高级篇(spark shuffle) 非常好的讲解
Spark修炼之道（进阶篇）——Spark入门到精通：第九节 Spark SQL执行流程解析
1.总体执行流程使用下列代码对SparkSQL流程进行分析.让大家明确LogicalPlan的几种状态,理解SparkSQL总体执行流程 // sc is an existing SparkCont ...
【转载】Spark性能优化指南——高级篇
前言数据倾斜调优调优概述数据倾斜发生时的现象数据倾斜发生的原理如何定位导致数据倾斜的代码查看导致数据倾斜的key的数据分布情况数据倾斜的解决方案解决方案一:使用Hive ETL预处理数 ...
Spark性能优化指南——高级篇（转载）
前言继基础篇讲解了每个Spark开发人员都必须熟知的开发调优与资源调优之后,本文作为<Spark性能优化指南>的高级篇,将深入分析数据倾斜调优与shuffle调优,以解决更加棘手的性能问 ...
Spark性能优化指南-高级篇
转自https://tech.meituan.com/spark-tuning-pro.html,感谢原作者的贡献前言继基础篇讲解了每个Spark开发人员都必须熟知的开发调优与资源调优之后,本文作 ...
Spark性能调优-高级篇
前言继基础篇讲解了每个Spark开发人员都必须熟知的开发调优与资源调优之后,本文作为<Spark性能优化指南>的高级篇,将深入分析数据倾斜调优与shuffle调优,以解决更加棘手的性能问 ...

随机推荐

c/s结构的自动化——pyautogui
环境:Python 3.5.3 pip install pyautogui -i http://pypi.douban.com/simple --trusted-host pypi.douban.co ...
BZOJ 3529 [Sdoi2014]数表 (莫比乌斯反演+树状数组+离线)
题目大意:有一张$n*m$的数表,第$i$行第$j$列的数是同时能整除$i,j$的所有数之和,求数表内所有不大于A的数之和先是看错题了...接着看对题了发现不会做了...刚了大半个下午无果看了Po ...
WPF 封装 dotnet remoting 调用其他进程
原文:WPF 封装 dotnet remoting 调用其他进程本文告诉大家一个封装好的库,使用这个库可以快速搭建多进程相互使用. 目录创建端口调用软件运行的类运行C++程序通道使用在 ...
win系统安装node出现这个2503和2502解决办法
一: 今天在公司的新电脑要安装appium,所以要搭建appium的环境,所以在安装到node的时候,出现了内部错误2503和2502,安装中断. 这种错误可能是权限不足导致,一般“.exe”程序可以 ...
Linux学习总结（10）——Linux查看CPU和内存使用情况
在系统维护的过程中,随时可能有需要查看 CPU 使用率,并根据相应信息分析系统状况的需要.在 CentOS 中,可以通过 top 命令来查看 CPU 使用状况.运行 top 命令后,CPU 使用状态会 ...
spring boot启动STS 运行报错 java.lang.NoClassDefFoundError: ch/qos/logback/classic/LoggerContext
spring boot启动STS 运行报错 java.lang.NoClassDefFoundError: ch/qos/logback/classic/LoggerContext 学习了: http ...
Qt 3D教程（二）初步显示3D的内容
Qt3D教程(二)初步显示3D的内容前一篇很easy,全然就没有牵涉到3D的内容,它仅仅是我们搭建3D应用的基本框架而已,而这一篇.我们将要利用它来初步地显示3D的内容了! 本次目的是将程序中间的内 ...
冒泡，简单选择，直接插入排序（Java版）
冒泡.简单选择,直接插入这三种排序都是简单排序. 工具类 package Utils; import java.util.Arrays; public class SortUtils { public ...
Servlet监听器及在线用户
Servlet中的监听器分为三种类型Ⅰ 监听ServletContext.Request.Session作用域的创建和销毁 (1)ServletContextListener (2)HttpSessi ...
springMVC接受前台传值
今天,用ajax向springMVC的控制器传参数,是一个json对象.({"test":"test","test1":"test ...

Spark修炼之道（高级篇）——Spark源代码阅读：第十二节 Spark SQL 处理流程分析

Spark修炼之道（高级篇）——Spark源代码阅读：第十二节 Spark SQL 处理流程分析的更多相关文章

随机推荐

热门专题