Kafka 0.8: 多日志文件夹机制

kafka 0.7.2 中对log.dir的定义如下：

log.dir none Specifies the root directory in which all log data is kept.

在kafka 0.8 中将log.dir 修改为 log.dirs，官方文档说明如下：

log.dirs

/tmp/kafka-logs

A comma-separated list of one or more directories in which Kafka data is stored. Each new partition that is created will be placed in the directory which currently has the fewest partitions.

从0.8开始，支持配置多个日志文件夹，文件夹之间使用逗号隔开即可，这样做在实际项目中有非常大的好处，那就是支持多硬盘。

下面从源码着手来浅析一下多日志文件夹是怎么工作的

1. 首先broker启动时会加载指定的配置文件，并把property对象传入KafkaConfig对象中

object Kafka extends Logging {

    try {

      val props = Utils.loadProps(args(0))

      val serverConfig = new KafkaConfig(props)

2. 在kafkaConfig 中会解析log.dirs字符串，将其通过逗号隔开，形成Set，调用split方法时传入"\\s*,\\s*"，表示逗号前后的空格都会被忽略

 /* the directories in which the log data is kept */

  val logDirs = Utils.parseCsvList(props.getString("log.dirs", props.getString("log.dir", "/tmp/kafka-logs")))

  require(logDirs.size > 0)

  /**

   * Parse a comma separated string into a sequence of strings.

   * Whitespace surrounding the comma will be removed.

   */

  def parseCsvList(csvList: String): Seq[String] = {

    if(csvList == null || csvList.isEmpty)

      Seq.empty[String]

    else {

      csvList.split("\\s*,\\s*").filter(v => !v.equals(""))

    }

  }

3. 在KafkaServer中生成LogManager对象时传入 [(dir_path_1,File(dir_path_1)), (dir_path_2,File(dir_path_2)) ]

    new LogManager(logDirs = config.logDirs.map(new File(_)).toArray,

                   topicConfigs = configs,

                   defaultConfig = defaultLogConfig,

                   cleanerConfig = cleanerConfig,

                   flushCheckMs = config.logFlushSchedulerIntervalMs,

                   flushCheckpointMs = config.logFlushOffsetCheckpointIntervalMs,

                   retentionCheckMs = config.logCleanupIntervalMs,

                   scheduler = kafkaScheduler,

                   time = time)

4.LogManager首先对传入的dir进行下列验证：是否存在相同的文件夹、文件夹是否存在（不存在则创建）、是否为可读的文件夹

  /**

   * Create and check validity of the given directories, specifically:

   * <ol>

   * <li> Ensure that there are no duplicates in the directory list

   * <li> Create each directory if it doesn't exist

   * <li> Check that each path is a readable directory

   * </ol>

   */

  private def createAndValidateLogDirs(dirs: Seq[File]) {

    if(dirs.map(_.getCanonicalPath).toSet.size < dirs.size)

      throw new KafkaException("Duplicate log directory found: " + logDirs.mkString(", "))

    for(dir <- dirs) {

      if(!dir.exists) {

        info("Log directory '" + dir.getAbsolutePath + "' not found, creating it.")

        val created = dir.mkdirs()

        if(!created)

          throw new KafkaException("Failed to create data directory " + dir.getAbsolutePath)

      }

      if(!dir.isDirectory || !dir.canRead)

        throw new KafkaException(dir.getAbsolutePath + " is not a readable log directory.")

    }

  }

5. LogManager 对所有的文件夹获取文件锁，防止其他进行对该文件夹进行操作

  /**

   * Lock all the given directories

   */

  private def lockLogDirs(dirs: Seq[File]): Seq[FileLock] = {

    dirs.map { dir =>

      val lock = new FileLock(new File(dir, LockFile))

      if(!lock.tryLock())

        throw new KafkaException("Failed to acquire lock on file .lock in " + lock.file.getParentFile.getAbsolutePath +

                               ". A Kafka instance in another process or thread is using this directory.")

      lock

    }

  }

6. 通过文件夹下面的recovery-point-offset-checkpoint 恢复加载每个目录下面的partition文件

  /**

   * Recover and load all logs in the given data directories

   */

  private def loadLogs(dirs: Seq[File]) {

    for(dir <- dirs) {

      val recoveryPoints = this.recoveryPointCheckpoints(dir).read

      /* load the logs */

      val subDirs = dir.listFiles()

      if(subDirs != null) {

        //当kafka退出时，正常关闭的日志文件都会在该日志文件下生成.kafka_cleanshutdown为后缀的文件，该文件的作用是，在下次启动时，此日志文件可以不进行恢复流程

        val cleanShutDownFile = new File(dir, Log.CleanShutdownFile)

        if(cleanShutDownFile.exists())

          info("Found clean shutdown file. Skipping recovery for all logs in data directory '%s'".format(dir.getAbsolutePath))

        for(dir <- subDirs) {

          if(dir.isDirectory) {

            info("Loading log '" + dir.getName + "'")

            val topicPartition = Log.parseTopicPartitionName(dir.getName)

            val config = topicConfigs.getOrElse(topicPartition.topic, defaultConfig)

            val log = new Log(dir,

                              config,

                              recoveryPoints.getOrElse(topicPartition, 0L),

                              scheduler,

                              time)

            val previous = this.logs.put(topicPartition, log)

            if(previous != null)

              throw new IllegalArgumentException("Duplicate log directories found: %s, %s!".format(log.dir.getAbsolutePath, previous.dir.getAbsolutePath))

          }

        }

        cleanShutDownFile.delete()

      }

    }

  }

7. 当需要创建新的日志文件时，会在日志文件比较少的文件夹下去创建，源码中的注释很详细

  /**

   * Choose the next directory in which to create a log. Currently this is done

   * by calculating the number of partitions in each directory and then choosing the

   * data directory with the fewest partitions.

   */

  private def nextLogDir(): File = {

    if(logDirs.size == 1) {

      logDirs(0)

    } else {

      // count the number of logs in each parent directory (including 0 for empty directories

      val logCounts = allLogs.groupBy(_.dir.getParent).mapValues(_.size)

      val zeros = logDirs.map(dir => (dir.getPath, 0)).toMap

      //下面代码的主要作用是，对没有日志文件的文件夹设置size为0

      var dirCounts = (zeros ++ logCounts).toBuffer

      // choose the directory with the least logs in it

      val leastLoaded = dirCounts.sortBy(_._2).head

      new File(leastLoaded._1)

    }

  }

Kafka 0.8: 多日志文件夹机制的更多相关文章

hololens DEP2220: 无法删除目标计算机“127.0.0.1”上的文件夹
Hololens开发调试的过程中,可能会出现 “DEP2220: 无法删除目标计算机“127.0.0.1”上的文件夹“ 的错误导致无法部署,解决办法是进入项目属性页——调试——启动选项,勾选“卸载并重 ...
eas之日志文件夹
F:\ThisIs_MyWork\kingdee\eas\server\profiles\server1\logs 服务端的日志文件夹 F:\ThisIs_MyWork\kingdeecusto ...
oracle 10g/11g 命令对照，日志文件夹对照
oracle 10g/11g 命令对照,日志文件夹对照 oracle 11g 中不再建议使用的命令 Deprecated Command Replacement Commands crs_st ...
CI3.0控制器下面建文件夹访问一直404 的解决方法
在单入口文件(框架目录下面的index.php)最下面的require_once BASEPATH.'core/CodeIgniter.php';这行上面设置一个路径,是相对于conrollers文件 ...
Kafka 入门（二）--数据日志、副本机制和消费策略
一.Kafka 数据日志 1.主题 Topic Topic 是逻辑概念. 主题类似于分类,也可以理解为一个消息的集合.每一条发送到 Kafka 的消息都会带上一个主题信息,表明属于哪个主题. Kafk ...
IIS下众多网站，如何快速定位某站点日志在哪个文件夹？
windows2008,iis 多站点, 日志.应用程序池都是默认设置, 没有分开………… Logs目录里面有W3SVC43,W3SVC44,W3SVC45,W3SVC46.....等等日志文件夹. ...
asp 中创建日志打印文件夹
string FilePath = HttpRuntime.BinDirectory.ToString(); string FileName = FilePath + "日志" + ...
iis7下查看站点日志对应文件夹
原文:iis7下查看站点日志对应文件夹 IIS7下面默认日志文件的存放路径:%SystemDrive%\inetpub\logs\LogFiles 查看方法:点击对应网站 -> 右侧功能视图 - ...
kafka 0.10.2 cetos6.5 集群部署
安装 zookeeper http://www.cnblogs.com/xiaojf/p/6572351.html安装 scala http://www.cnblogs.com/xiaojf/p/65 ...

随机推荐

linux下跨服务器文件文件夹的复制
文件的复制:scp –P (端口号) ./authorized_keys berchina@hadoop002:/home/berchina 文件夹的复制:scp -r -P (端口号) /home/ ...
Android开发之ADT中无Annotation Processin的解决办法
使用ButterKnife的时候,进入ADT中设置的时候发现在Java Compiler展开后无Annotation Processin 解决办法: 安装插件:Juno - http://downlo ...
制作LiveCD
1) 需要的工具Redhat9.0.VMware虚拟机,选择用grub作loader 2) 制作ramdisk A) cd /usr/local && mk ...
Windows系统中IIS 6.0+Tomcat服务器环境的整合配置过程
IIS6.0+Tomcat整合 1.首先准备工作 Windows IIS 6.0 apache-tomcat-7.0.26.exe tomcat-connectors-1.2.33-windows-i ...
C#图片处理之: 另存为压缩质量可自己控制的JPEG
处理图片时常用的过程是:读入图片文件并转化为Bitmap -> 处理此Bitmap的每个点以得到需要的效果 -> 保存新的Bitmap到文件使用C#很方便的就可以把多种格式的图片文件读到B ...
linux cross toolsChain 交叉编译 ARM（转）
转载请注明出处:http://blog.csdn.net/mybelief321/article/details/9076583 安装环境 Linux版本:Ubuntu 12.04 内核版本:L ...
openlayer调用geoserver发布的地图实现地图的基本功能
转自:http://starting.iteye.com/blog/1039809 主要实现的功能有放大,缩小,获取地图大小,平移,线路测量,面积测量,拉宽功能,显示标注,移除标注,画多边形获取经纬度 ...
SQL SERVER 2008查询其他数据库
1.访问本地的其他数据库 --启用Ad Hoc Distributed Queries-- reconfigure reconfigure -- 使用完成后,关闭Ad Hoc Distributed ...
jps 显示process information unavailable解决方法
jps 显示process information unavailable解决办法jps时出现如下信息: 4791 -- process information unavailable 解决办法: 进 ...
PLS-00306:错误解决思路 - OracleHelper 执行Oracle函数的坑
如果你是像我一样初次使用Net+Oracle的结合,我想你会跟我一样,有很大的概率碰到这个问题 ==================================================== ...

Kafka 0.8: 多日志文件夹机制

Kafka 0.8: 多日志文件夹机制的更多相关文章

随机推荐

热门专题