FileInputFormat 的实现之TextInputFormat

说明

TextInputFormat默认是按行切分记录record，本篇在于理解，对于同一条记录record，如果被切分在不同的split时是怎么处理的。首先getSplits是在逻辑上划分，并没有物理切分，也就是只是记录每个split从文件的个位置读到哪个位置，文件还是一个整体。所以在LineRecordReader中，它的处理方式是每个split多读一行，也就是读到下一个split的第一行。然后除了每个文件的第一个split，其他split都跳过第一行，进而避免重复读取，这种方式去处理。

FileInputFomat 之 getSplits

TextInputFormat 继承TextInputFormat，并没有重写getSplits，而是沿用父类的getSplits方法，下面看下该方法的源码

public List<InputSplit> getSplits(JobContext job) throws IOException {

    StopWatch sw = new StopWatch().start();

    //getFormatMinSplitSize() == 1，getMinSplitSize(job)为用户设置的切片最小值，默认1。 job.getConfiguration().getLong("mapreduce.input.fileinputformat.split.minsize", 1L);

    long minSize = Math.max(getFormatMinSplitSize(), getMinSplitSize(job));

    // getMaxSplitSize(job)为用户设置的切片最大值，context.getConfiguration().getLong("mapreduce.input.fileinputformat.split.maxsize", Long.MAX_VALUE);

    long maxSize = getMaxSplitSize(job);

    // generate splits

    List<InputSplit> splits = new ArrayList<InputSplit>();

    List<FileStatus> files = listStatus(job);

    for (FileStatus file: files) {

      Path path = file.getPath();

      long length = file.getLen();

      if (length != 0) {

        BlockLocation[] blkLocations;

        //LocatedFileStatus带有blockLocation信息

        if (file instanceof LocatedFileStatus) {

          blkLocations = ((LocatedFileStatus) file).getBlockLocations();

        } else {

          FileSystem fs = path.getFileSystem(job.getConfiguration());

          blkLocations = fs.getFileBlockLocations(file, 0, length);

        }

        //判断文件是否可切分

        if (isSplitable(job, path)) {

          long blockSize = file.getBlockSize();

          //真正的切片设置大小判断，computeSplitSize方法中的实现，返回值 Math.max(minSize, Math.min(maxSize, blockSize));

          long splitSize = computeSplitSize(blockSize, minSize, maxSize);

          long bytesRemaining = length;

          while (((double) bytesRemaining)/splitSize > SPLIT_SLOP) {

            int blkIndex = getBlockIndex(blkLocations, length-bytesRemaining);

            splits.add(makeSplit(path, length-bytesRemaining, splitSize,

                        blkLocations[blkIndex].getHosts(),

                        blkLocations[blkIndex].getCachedHosts()));

            bytesRemaining -= splitSize;

          }

          if (bytesRemaining != 0) {

            int blkIndex = getBlockIndex(blkLocations, length-bytesRemaining);

            splits.add(makeSplit(path, length-bytesRemaining, bytesRemaining,

                       blkLocations[blkIndex].getHosts(),

                       blkLocations[blkIndex].getCachedHosts()));

          }

        } else { // not splitable

          if (LOG.isDebugEnabled()) {

            // Log only if the file is big enough to be splitted

            if (length > Math.min(file.getBlockSize(), minSize)) {

              LOG.debug("File is not splittable so no parallelization "

                  + "is possible: " + file.getPath());

            }

          }

          splits.add(makeSplit(path, 0, length, blkLocations[0].getHosts(),

                      blkLocations[0].getCachedHosts()));

        }

      } else {

        //Create empty hosts array for zero length files

        splits.add(makeSplit(path, 0, length, new String[0]));

      }

    }

    // Save the number of input files for metrics/loadgen

    job.getConfiguration().setLong(NUM_INPUT_FILES, files.size());

    sw.stop();

    if (LOG.isDebugEnabled()) {

      LOG.debug("Total # of splits generated by getSplits: " + splits.size()

          + ", TimeTaken: " + sw.now(TimeUnit.MILLISECONDS));

    }

    return splits;

  }

FileInputFomat 之 createRecordReader，主要是看LineRecordReader

public RecordReader<LongWritable, Text>

    createRecordReader(InputSplit split,

                       TaskAttemptContext context) {

    //设置record的分隔符

    String delimiter = context.getConfiguration().get(

        "textinputformat.record.delimiter");

    byte[] recordDelimiterBytes = null;

    if (null != delimiter)

      recordDelimiterBytes = delimiter.getBytes(Charsets.UTF_8);

    return new LineRecordReader(recordDelimiterBytes);

  }

LineRecordReader的方法initialize和nextKeyValue方法

public void initialize(InputSplit genericSplit,

                         TaskAttemptContext context) throws IOException {

    FileSplit split = (FileSplit) genericSplit;

    Configuration job = context.getConfiguration();

    this.maxLineLength = job.getInt(MAX_LINE_LENGTH, Integer.MAX_VALUE);

    start = split.getStart();

    end = start + split.getLength();

    final Path file = split.getPath();

    // open the file and seek to the start of the split

    final FileSystem fs = file.getFileSystem(job);

    fileIn = fs.open(file);

    //判断是否压缩，赋值对应的SplitLineReader

    CompressionCodec codec = new CompressionCodecFactory(job).getCodec(file);

    if (null!=codec) {

      isCompressedInput = true;

      decompressor = CodecPool.getDecompressor(codec);

      if (codec instanceof SplittableCompressionCodec) {

        final SplitCompressionInputStream cIn =

          ((SplittableCompressionCodec)codec).createInputStream(

            fileIn, decompressor, start, end,

            SplittableCompressionCodec.READ_MODE.BYBLOCK);

        in = new CompressedSplitLineReader(cIn, job,

            this.recordDelimiterBytes);

        start = cIn.getAdjustedStart();

        end = cIn.getAdjustedEnd();

        filePosition = cIn;

      } else {

        in = new SplitLineReader(codec.createInputStream(fileIn,

            decompressor), job, this.recordDelimiterBytes);

        filePosition = fileIn;

      }

    } else {

      fileIn.seek(start);

      in = new UncompressedSplitLineReader(

          fileIn, job, this.recordDelimiterBytes, split.getLength());

      filePosition = fileIn;

    }

    //这句是关键，由于getSplits的时候，并不能保证一条record记录，不被切分到不同的split。所以处理方式是，除了每个文件的第一个split，其他每个split多读一行

    //所以避免重复读，不是开始的split都跳过第一行。

    // If this is not the first split, we always throw away first record

    // because we always (except the last split) read one extra line in

    // next() method.

    if (start != 0) {

      start += in.readLine(new Text(), 0, maxBytesToConsume(start));

    }

    this.pos = start;

  }

接下来是nextKeyValue

public boolean nextKeyValue() throws IOException {

    if (key == null) {

      key = new LongWritable();

    }

    key.set(pos);

    if (value == null) {

      value = new Text();

    }

    int newSize = 0;

    // We always read one extra line, which lies outside the upper

    // split limit i.e. (end - 1)

    //这个in具体看是CompressedSplitLineReader还是UncompressedSplitLineReader，重写了其中的readerLine方法

    while (getFilePosition() <= end || in.needAdditionalRecordAfterSplit()) {

      if (pos == 0) {

        //跳过utf的开头

        newSize = skipUtfByteOrderMark();

      } else {

        //readerLine有两种实现方法，一种readCustomLine这种是自己定义了record的分隔符，还有一种是readDefaultLine，这种是没有自定义分隔符，默认的读取数据的方式，用\r,\n或者\r\n分割

        newSize = in.readLine(value, maxLineLength, maxBytesToConsume(pos));

        pos += newSize;

      }

      if ((newSize == 0) || (newSize < maxLineLength)) {

        break;

      }

      // line too long. try again

      LOG.info("Skipped line of size " + newSize + " at pos " +

               (pos - newSize));

    }

    if (newSize == 0) {

      key = null;

      value = null;

      return false;

    } else {

      return true;

    }

  }

FileInputFormat 的实现之TextInputFormat的更多相关文章

MapReduce ：基于 FileInputFormat 的 mapper 数量控制
本篇分两部分,第一部分分析使用 java 提交 mapreduce 任务时对 mapper 数量的控制,第二部分分析使用 streaming 形式提交 mapreduce 任务时对 mapper 数量 ...
MR 的 mapper 数量问题
看到群里面一篇文章涨了贱识 http://www.cnblogs.com/xuxm2007/archive/2011/09/01/2162011.html 之前关注过 reduceer 的数量问题,还 ...
（转）通过input分片的大小来设置map的个数
摘要通过input分片的大小来设置map的个数 map inputsplit hadoop 前言:在具体执行Hadoop程序的时候,我们要根据不同的情况来设置Map的个数.除了设置固定的每个节点上可 ...
Spark大数据针对性问题。
1.海量日志数据,提取出某日访问百度次数最多的那个IP. 解决方案:首先是将这一天,并且是访问百度的日志中的IP取出来,逐个写入到一个大文件中.注意到IP是32位的,最多有个2^32个IP.同样可以采 ...
通过inputSplit分片size控制map数目
前言:在具体执行Hadoop程序的时候,我们要根据不同的情况来设置Map的个数.除了设置固定的每个节点上可运行的最大map个数外,我们还需要控制真正执行Map操作的任务个数. 1.如何控制实际运行的m ...
MapReduce框架原理-InputFormat数据输入
InputFormat简介 InputFormat:管控MR程序文件输入到Mapper阶段,主要做两项操作:怎么去切片?怎么将切片数据转换成键值对数据. InputFormat是一个抽象类,没有实现怎 ...
旧版API的TextInputFormat源码分析
TextInputFormat类 package org.apache.hadoop.mapred; import java.io.*; import org.apache.hadoop.fs.*; ...
FileInputFormat
MapReduce框架要处理数据的文件类型 FileInputFormat这个类决定. TextInputFormat是框架默认的文件类型,可以处理Text文件类型,如果你要处理的文件类型不是Text ...
MapReduce中TextInputFormat分片和读取分片数据源码级分析
InputFormat主要用于描述输入数据的格式(我们只分析新API,即org.apache.hadoop.mapreduce.lib.input.InputFormat),提供以下两个功能: (1) ...

随机推荐

【转载】C#通过IndexOf方法获取某一列在DataTable中的索引位置
在C#中的Datatable数据变量的操作过程中,有时候需要知道某一个列名在DataTable中的索引位置信息,此时可以通过DataTable变量的Columns属性来获取到所有的列信息,然后通过Co ...
hibernate Criteria中多个or和and的用法 and ( or or)
/s筛选去除无效数据 /* detachedCriteria.add( Restrictions.or( Restrictions.like("chanpin", &qu ...
Java集合学习(8)：LinkedList
一.概述 LinkedList和ArrayList一样,都实现了List接口,但其内部的数据结构有本质的不同.LinkedList是基于链表实现的(通过名字也能区分开来),所以它的插入和删除操作比Ar ...
Unity基础功能：粒子特效（Shuriken）
版权申明: 本文原创首发于以下网站: 博客园『优梦创客』的空间:https://www.cnblogs.com/raymondking123 优梦创客的官方博客:https://91make.top ...
Swagger 学习资料
Swagger 学习资料网址 Spring Boot中使用Swagger2构建强大的RESTful API文档 http://blog.didispace.com/springbootswagger ...
动态创建form 完成form 提交
document.body.appendChild(jForm) won't work because jForm is not a dom element, it is a jQuery objec ...
19、Python标准库: 日期和时间
一.time时间模块 import time 1 .时间戳时间戳(timestamp):时间戳表示的是从1970年1月1日00:00:00开始按秒计算的偏移量. time_stamp = tim ...
python tkinter txt窗口，开线程两个卡死
一个线程可以,两个卡死 #!/usr/bin/python # coding: utf-8 from Tkinter import * import threading class MyThread2 ...
jQuery扩展$.fn、$.extend jQery命名方法扩展练习总结
<script> $.fn.hello = function(){ //扩展jQuery实例的自定义方法,基于$.fn的jq方法扩展 this.click(function(){ ...
c语言的#和##的用法
#include <stdio.h> #define ADD(A,B) printf(#A " + " #B " = %d\n",((A)+(B)) ...

FileInputFormat 的实现之TextInputFormat

说明

FileInputFomat 之 getSplits

FileInputFomat 之 createRecordReader，主要是看LineRecordReader

LineRecordReader的方法initialize和nextKeyValue方法

接下来是nextKeyValue

FileInputFormat 的实现之TextInputFormat的更多相关文章

随机推荐

热门专题