【Nutch2.2.1源代码分析之5】索引的基本流程

一、各个主要类之间的关系

SolrIndexerJob extends IndexerJob

1、IndexerJob：主要完成

2、SolrIndexerJob：主要完成

3、IndexUtil：主要只有一个方法public NutchDocument index(String key, WebPage page)，用于根据网页信息，返回一个solr的Document对象。

二、程序调用流程

查看Nutch中的执行脚本--nutch，得到以下信息：

elif [ "$COMMAND" = "solrindex" ] ; then

CLASS=org.apache.nutch.indexer.solr.SolrIndexerJob

因此程序入口位于SolrIndexerJob类中。

（一）org.apache.nutch.indexer.SolrIndexerJob

1、程序入口

  public static void main(String[] args) throws Exception {

    final int res = ToolRunner.run(NutchConfiguration.create(),

        new SolrIndexerJob(), args);

    System.exit(res);

  }

使用了ToolRunner.run()来执行程序，可参考：使用ToolRunner运行Hadoop程序基本原理分析。

其中第一个参数主要是加载了nutch相关的参数，主要包括hadoop的core-default.xml、core-site.xml以及nutch的nutch-default.xml、nutch-site.xml。

第二个参数指明了运行SolrIndexerJob的run(String[])方法.

2、执行SolrIndexerJob类的run(String[])方法

  public int run(String[] args) throws Exception {

    if (args.length < 2) {

      System.err.println("Usage: SolrIndexerJob <solr url> (<batchId> | -all | -reindex) [-crawlId <id>]");

      return -1;

    }

    if (args.length == 4 && "-crawlId".equals(args[2])) {

      getConf().set(Nutch.CRAWL_ID_KEY, args[3]);

    }

    try {

      indexSolr(args[0], args[1]);

      return 0;

    } catch (final Exception e) {

      LOG.error("SolrIndexerJob: " + StringUtils.stringifyException(e));

      return -1;

    }

  }

先判断参数的合理性，然后执行执行indexSolr(String,String)方法。

3、执行indexSolr(String,String)方法

public void indexSolr(String solrUrl, String batchId) throws Exception {

    LOG.info("SolrIndexerJob: starting");

    run(ToolUtil.toArgMap(

        Nutch.ARG_SOLR, solrUrl,

        Nutch.ARG_BATCH, batchId));

    // do the commits once and for all the reducers in one go

    getConf().set(SolrConstants.SERVER_URL,solrUrl);

    SolrServer solr = SolrUtils.getCommonsHttpSolrServer(getConf());

    if (getConf().getBoolean(SolrConstants.COMMIT_INDEX, true)) {

      solr.commit();

    }

    LOG.info("SolrIndexerJob: done.");

  }

4、执行run(Map<...>）方法

@Override

  public Map<String,Object> run(Map<String,Object> args) throws Exception {

    String solrUrl = (String)args.get(Nutch.ARG_SOLR);

    String batchId = (String)args.get(Nutch.ARG_BATCH);

    NutchIndexWriterFactory.addClassToConf(getConf(), SolrWriter.class);

    getConf().set(SolrConstants.SERVER_URL, solrUrl);

    currentJob = createIndexJob(getConf(), "solr-index", batchId);

    currentJob.waitForCompletion(true);

    ToolUtil.recordJobStatus(null, currentJob, results);

    return results;

  }

（二）org.apache.nutch.indexer.IndexerJob

1、执行createIndexJob()方法。

  protected Job createIndexJob(Configuration conf, String jobName, String batchId)

  throws IOException, ClassNotFoundException {

    conf.set(GeneratorJob.BATCH_ID, batchId);

    Job job = new NutchJob(conf, jobName);

    // TODO: Figure out why this needs to be here

    job.getConfiguration().setClass("mapred.output.key.comparator.class",

        StringComparator.class, RawComparator.class);

    Collection<WebPage.Field> fields = getFields(job);

    StorageUtils.initMapperJob(job, fields, String.class, NutchDocument.class,

        IndexerMapper.class);

    job.setNumReduceTasks(0);

    job.setOutputFormatClass(IndexerOutputFormat.class);

    return job;

  }

}

2、执行map相关的方法，包括setup()，map()，cleanup()

  public static class IndexerMapper

      extends GoraMapper<String, WebPage, String, NutchDocument> {

    public IndexUtil indexUtil;

    public DataStore<String, WebPage> store;

    protected Utf8 batchId;

    @Override

    public void setup(Context context) throws IOException {

      Configuration conf = context.getConfiguration();

      batchId = new Utf8(conf.get(GeneratorJob.BATCH_ID, Nutch.ALL_BATCH_ID_STR));

      indexUtil = new IndexUtil(conf);

      try {

        store = StorageUtils.createWebStore(conf, String.class, WebPage.class);

      } catch (ClassNotFoundException e) {

        throw new IOException(e);

      }

    }

    protected void cleanup(Context context) throws IOException ,InterruptedException {

      store.close();

    };

    @Override

    public void map(String key, WebPage page, Context context)

    throws IOException, InterruptedException {

      ParseStatus pstatus = page.getParseStatus();

      if (pstatus == null || !ParseStatusUtils.isSuccess(pstatus)

          || pstatus.getMinorCode() == ParseStatusCodes.SUCCESS_REDIRECT) {

        return; // filter urls not parsed

      }

      Utf8 mark = Mark.UPDATEDB_MARK.checkMark(page);

      if (!batchId.equals(REINDEX)) {

        if (!NutchJob.shouldProcess(mark, batchId)) {

          if (LOG.isDebugEnabled()) {

            LOG.debug("Skipping " + TableUtil.unreverseUrl(key) + "; different batch id (" + mark + ")");

          }

          return;

        }

      }

      NutchDocument doc = indexUtil.index(key, page);

      if (doc == null) {

        return;

      }

      if (mark != null) {

        Mark.INDEX_MARK.putMark(page, Mark.UPDATEDB_MARK.checkMark(page));

        store.put(key, page);

      }

      context.write(key, doc);

    }

  }

3、调用context.write()

由于 job.setOutputFormatClass(IndexerOutputFormat.class);

所以写入index？？

（三）public class IndexUtil

1、调用index()方法

  public NutchDocument index(String key, WebPage page) {

    NutchDocument doc = new NutchDocument();

    doc.add("id", key);

    doc.add("digest", StringUtil.toHexString(page.getSignature()));

    if (page.getBatchId() != null) {

      doc.add("batchId", page.getBatchId().toString());

    }

    String url = TableUtil.unreverseUrl(key);

    if (LOG.isDebugEnabled()) {

      LOG.debug("Indexing URL: " + url);

    }

    try {

      doc = filters.filter(doc, url, page);

    } catch (IndexingException e) {

      LOG.warn("Error indexing "+key+": "+e);

      return null;

    }

    // skip documents discarded by indexing filters

    if (doc == null) return null;

    float boost = 1.0f;

    // run scoring filters

    try {

      boost = scoringFilters.indexerScore(url, doc, page, boost);

    } catch (final ScoringFilterException e) {

      LOG.warn("Error calculating score " + key + ": " + e);

      return null;

    }

    doc.setScore(boost);

    // store boost for use by explain and dedup

    doc.add("boost", Float.toString(boost));

    return doc;

  }

三、plugin中的字段索引

1、关于basic字段的索引在public class BasicIndexingFilter implements IndexingFilter 中

【Nutch2.2.1源代码分析之5】索引的基本流程的更多相关文章

【Nutch2.2.1源代码分析之4】Nutch加载配置文件的方法
小结: (1)在nutch中,一般通过ToolRunner来运行hadoop job,此方法可以方便的通过ToolRunner.run(Configuration conf,Tool tool,Str ...
【原创】k8s源代码分析-----kubelet（1）主要流程
本人空间链接http://user.qzone.qq.com/29185807/blog/1460015727 源代码为k8s v1.1.1稳定版本号 kubelet代码比較复杂.主要是由于其担负的任 ...
LIRe 源代码分析 4：建立索引（DocumentBuilder）[以颜色布局为例]
===================================================== LIRe源代码分析系列文章列表: LIRe 源代码分析 1:整体结构 LIRe 源代码分析 ...
转：SDL2源代码分析
1:初始化(SDL_Init()) SDL简介有关SDL的简介在<最简单的视音频播放示例7:SDL2播放RGB/YUV>以及<最简单的视音频播放示例9:SDL2播放PCM>中 ...
Hadoop源代码分析
http://wenku.baidu.com/link?url=R-QoZXhc918qoO0BX6eXI9_uPU75whF62vFFUBIR-7c5XAYUVxDRX5Rs6QZR9hrBnUdM ...
Spark SQL 源代码分析之 In-Memory Columnar Storage 之 in-memory query
/** Spark SQL源代码分析系列文章*/ 前面讲到了Spark SQL In-Memory Columnar Storage的存储结构是基于列存储的. 那么基于以上存储结构,我们查询cache ...
新秀nginx源代码分析数据结构篇（四）红黑树ngx_rbtree_t
新秀nginx源代码分析数据结构篇(四)红黑树ngx_rbtree_t Author:Echo Chen(陈斌) Email:chenb19870707@gmail.com Blog:Blog.csd ...
HBase源代码分析之HRegion上MemStore的flsuh流程（一）
了解HBase架构的用户应该知道,HBase是一种基于LSM模型的分布式数据库.LSM的全称是Log-Structured Merge-Trees.即日志-结构化合并-树. 相比于Oracle普通索引 ...
SDL2源代码分析3：渲染器（SDL_Renderer）
===================================================== SDL源代码分析系列文章列表: SDL2源代码分析1:初始化(SDL_Init()) SDL ...

随机推荐

Ajax实现三级联动(0520)
查询数据库中的chinastates表,通过父级代号查询相应省市区. 实现界面: 在js页面实现三级联动在JQuery中调用Ajax方法(引用JQuery文件一定放在最上面) 用插件的形式,创建三个 ...
MYSQL常见出错mysql_errno()代码解析
如题,今天遇到怎么一个问题, 在理论上代码是不会有问题的,但是还是报了如上的错误,把sql打印出來放到DB中却可以正常执行.真是郁闷,在百度里面渡了很久没有相关的解释,到时找到几个没有人回复的 & ...
在PHP中开启CURL扩展，使其支持curl()函数
在用PHP开发CMS的时候,要用到PHP的curl函数,默认状态下,这个函数需要开启CURL扩展,有主机使用权的,可通过PHP.ini文件开启本扩展,方法如下: 1.打开php.ini,定位到;ext ...
php 中_set()_get()实例解析
<?php class Person { // 下面是人的成员属性, 都是封装的私有成员 private $name; // 人的名子 private $sex; // 人的性别 private ...
破解官方recovery的签名验证
步骤简述1.解包recovery.img,2.反编译/sbin/recovery,用ida64plus3.在反编译出来的文本中查找:signature 4.简单的看一下指令流程,CBZ下面是faile ...
cf C. Insertion Sort
http://codeforces.com/contest/362/problem/C #include <cstdio> #include <cstring> #includ ...
Powershell 设置数值格式 1
设置数值格式 1 6 6月, 2013 在 Powershell tagged 字符串 / 数字 / 文本 / 日期 / 格式化 by Mooser Lee 格式化操作符 -f 可以将数值插入到字符 ...
list append 总是复制前面的参数，而不复制最后一个参数
append 总是复制前面的参数,而不复制最后一个参数 (define a1 '(1 2 3)) (define a2 '(a b c)) (define x (append a1 a2)) x ; ...
Linux tr 命令使用
man tr: TR(1) User Commands TR(1) NAME tr - translate or delete characters SYNOPSIS tr [OPTION]... S ...
HDU4893--Wow! Such Sequence! （线段树延迟标记）
Wow! Such Sequence! Time Limit: 10000/5000 MS (Java/Others) Memory Limit: 65536/65536 K (Java/Oth ...

【Nutch2.2.1源代码分析之5】索引的基本流程

【Nutch2.2.1源代码分析之5】索引的基本流程的更多相关文章

随机推荐

热门专题