【Nutch2.2.1基础教程之3】Nutch2.2.1配置文件

nutch-site.xml

在nutch2.2.1中，有两份配置文件：nutch-default.xml与nutch-site.xml。

其中前者是nutch自带的默认属性，一般情况下不要修改。

如果需要修改默认属性，可以在nutch-site.xml中增加一个同名的属性，并修改其值。nutch-site.xml中的属性值会覆盖nutch-default.xml中的值。

1、db.ignore.external.links

若为true，则只抓取本域名内的网页，忽略外部链接。

可以在 regex-urlfilter.txt中增加过滤器达到同样效果，但如果过滤器过多，如几千个，则会大大影响nutch的性能。

<property>

  <name>db.ignore.external.links</name>

  <value>true</value>

  <description>If true, outlinks leading from a page to external hosts

  will be ignored. This is an effective way to limit the crawl to include

  only initially injected hosts, without creating complex URLFilters.

  </description>

</property>

2、fetcher.parse

能否在抓取的同时进行解释：可以，但不建议这样做。

<property>

  <name>fetcher.parse</name>

  <value>false</value>

  <description>If true, fetcher will parse content. NOTE: previous releases would

  default to true. Since 2.0 this is set to false as a safer default.</description>

</property>

官方解释

N.B. In a parsing fetcher, outlinks are processed in the reduce phase (at least when outlinks are followed). If a fetcher's reducer stalls you may run out of memory or disk space,
usually after a very long reduce job. Behaviour typical to this is usually
observed in this situation.

In summary, if it is possible, users are advised not to use a parsing fetcher as it is heavy on IO and often leads to the above outcome.

3、db.max.outlinks.per.page

默认情况下，Nutch只抓取某个网页的100个外部链接，导致部分链接无法抓取。若要改变此情况，可以修改此配置项。

<property>

  <name>db.max.outlinks.per.page</name>

  <value>100</value>

  <description>The maximum number of outlinks that we'll process for a page.  If this value is nonnegative (>=0), at most db.max.outlinks.per.page outlinks  will be processed for a page; otherwise, all outlinks will be processed.

  </description>

</property>

官方说明如下：http://wiki.apache.org/nutch/FAQ/

Nutch doesn't crawl relative URLs? Some pages are not indexed but my regex file and everything else is okay - what is going on?

The crawl tool has a default limitation of 100 outlinks of one page that are being fetched. To overcome this limitation change thedb.max.outlinks.per.page property to a higher
value or simply -1 (unlimited).

file: conf/nutch-default.xml

 <property>

   <name>db.max.outlinks.per.page</name>

   <value>-1</value>

   <description>The maximum number of outlinks that we'll process for a page.

   If this value is nonnegative (>=0), at most db.max.outlinks.per.page outlinks

   will be processed for a page; otherwise, all outlinks will be processed.

   </description>

 </property>

4、file.content.limit http.content.limit ftp.content.limit

默认情况下，nutch只抓取网页的前65536个字节，之后的内容将被丢弃。

但对于某些大型网站，首页的内容远远不止65536个字节，甚至前面65536个字节里面均是一些布局信息，并没有任何的超链接。

因此修改默认值如下：

<property>

  <name>file.content.limit</name>

  <value>-1</value>

  <description>The length limit for downloaded content using the file

   protocol, in bytes. If this value is nonnegative (>=0), content longer

   than it will be truncated; otherwise, no truncation at all. Do not

   confuse this setting with the http.content.limit setting.

  </description>

</property>

<property>

  <name>http.content.limit</name>

  <value>-1</value>

  <description>The length limit for downloaded content using the http

  protocol, in bytes. If this value is nonnegative (>=0), content longer

  than it will be truncated; otherwise, no truncation at all. Do not

  confuse this setting with the file.content.limit setting.

  </description>

</property>

<property>

  <name>ftp.content.limit</name>

  <value>-1</value>

  <description>The length limit for downloaded content, in bytes.

  If this value is nonnegative (>=0), content longer than it will be truncated;

  otherwise, no truncation at all.

  Caution: classical ftp RFCs never defines partial transfer and, in fact,

  some ftp servers out there do not handle client side forced close-down very

  well. Our implementation tries its best to handle such situations smoothly.

  </description>

</property>

【Nutch2.2.1基础教程之3】Nutch2.2.1配置文件的更多相关文章

【Nutch2.2.1基础教程之2.2】集成Nutch/Hbase/Solr构建搜索引擎之二：内容分析
请先参见"集成Nutch/Hbase/Solr构建搜索引擎之一:安装及运行",搭建测试环境 http://blog.csdn.net/jediael_lu/article/deta ...
【Nutch2.2.1基础教程之3】Nutch2.2.1配置文件分类： H3_NUTCH 2014-08-18 16:33 1376人阅读评论(0) 收藏
nutch-site.xml 在nutch2.2.1中,有两份配置文件:nutch-default.xml与nutch-site.xml. 其中前者是nutch自带的默认属性,一般情况下不要修改. 如 ...
【Nutch2.2.1基础教程之6】Nutch2.2.1抓取流程
一.抓取流程概述 1.nutch抓取流程当使用crawl命令进行抓取任务时,其基本流程步骤如下: (1)InjectorJob 开始第一个迭代 (2)GeneratorJob (3)FetcherJ ...
【Nutch2.2.1基础教程之1】nutch相关异常
1.在任务一开始运行,注入Url时即出现以下错误. InjectorJob: Injecting urlDir: urls InjectorJob: Using class org.apache.go ...
【Nutch2.2.1基础教程之2.1】集成Nutch/Hbase/Solr构建搜索引擎之一：安装及运行【单机环境】
1.下载相关软件,并解压版本号如下: (1)apache-nutch-2.2.1 (2) hbase-0.90.4 (3)solr-4.9.0 并解压至/usr/search 2.Nutch的配置 ...
【Nutch2.2.1基础教程之6】Nutch2.2.1抓取流程分类： H3_NUTCH 2014-08-15 21:39 2530人阅读评论(1) 收藏
一.抓取流程概述 1.nutch抓取流程当使用crawl命令进行抓取任务时,其基本流程步骤如下: (1)InjectorJob 开始第一个迭代 (2)GeneratorJob (3)FetcherJ ...
【Nutch2.2.1基础教程之1】nutch相关异常分类： H3_NUTCH 2014-08-08 21:46 1549人阅读评论(2) 收藏
1.在任务一开始运行,注入Url时即出现以下错误. InjectorJob: Injecting urlDir: urls InjectorJob: Using class org.apache.go ...
OpenVAS漏洞扫描基础教程之OpenVAS概述及安装及配置OpenVAS服务
OpenVAS漏洞扫描基础教程之OpenVAS概述及安装及配置OpenVAS服务 1. OpenVAS基础知识 OpenVAS(Open Vulnerability Assessment Sys ...
Python基础教程之List对象转
Python基础教程之List对象时间:2014-01-19 来源:服务器之家投稿:root 1.PyListObject对象typedef struct { PyObjec ...

随机推荐

LBS配置
js: <script type="text/javascript" src="http://api.map.baidu.com/api?v=2.0&ak= ...
仿校内textarea输入框字数限制效果
这是一个仿校内textarea回复消息输入框限制字数的效果,具体表现如下: 普通状态是一个输入框,当光标获取焦点时,出现字数记录和回复按钮 PS:上边那个小三角可不是用的图片. 普通状态效果如下: 获 ...
sublime text3 安装package control
20141104日更新的安装代码为 import urllib.request,os,hashlib; h = '7183a2d3e96f11eeadd761d777e62404' + 'e330c6 ...
Python新手学习基础之运算符——比较运算符
比较运算符比较运算符可以使用比较两个值,所有的内建类型都支持比较运算.当用运算符比较两个值时,结果是一个逻辑值,不是True,就是False. 有一点要注意的是,不同的类型的比较方式不一样,数字类型 ...
strcpy and memcpy
1. Inconsist length. char a3[2]; char *a = "Itis " strcpy(a3, a); It is wrong. a3 will b ...
setf
independent flags boolalpha read/write bool elements as alphabetic strings (true and false). showbas ...
storyboard和xib的区别
storyboard只是算是帮你布局,流程什么的,xib的另一种形势,比xib功能多,但是和分享完全没有半点关系你暂时可以理解为高级xib
数据库范式（1NF 2NF 3NF BCNF）详解
数据库的设计范式是数据库设计所需要满足的规范,满足这些规范的数据库是简洁的.结构明晰的,同时,不会发生插入(insert).删除(delete)和更新(update)操作异常.反之则是乱七八糟,不仅给 ...
WPF与输入法冲突研究之二：汉字输入法会导致WPF程序的崩溃！
如果是输入非汉字的数据信息,可以添加一下内容: xmlns:input="clr-namespace:System.Windows.Input;assembly=PresentationCo ...
<经验杂谈>C#/.Net字符串操作方法小结
字符串操作是C#中最基本的.最常见的.也是用的最多的,以下我总结了几种常见的方法 1.把字符串按照分隔符转换成 List /// <summary> /// 把字符串按照分隔符转换成 L ...

【Nutch2.2.1基础教程之3】Nutch2.2.1配置文件

【Nutch2.2.1基础教程之3】Nutch2.2.1配置文件的更多相关文章

随机推荐

热门专题