1. Detecting Languages During Indexing

  在索引的时候,solr可以使用langid UpdateRequestProcessor来识别语言,然后映射文本到特定语言的字段.solr支持这个功能的两个实现:

  1. Tika的语言解析功能:http://tika.apache.org/0.10/detection.html
  2. LangDetect语言解析:http://code.google.com/p/language-detection/

  可以从 http://blog.mikemccandless.com/2011/10/accuracy-and-performance-of-googles.html中看到它们之间的对比.一般情况下,LangDetect支持更多的语言,具有更高的性能.

  参考http://wiki.apache.org/solr/LanguageDetection获取更多的关于langid UpdateRequestProcessor信息.

1.1 Configuring Language Detection

  可以在solrconfig.xml中配置langid UpdateRequestProcessor.两个实现具有相同的参数,最少,你需要指定语言识别的字段和字段的结果语言编码.

1.2 Configuring Tika Language Detection

  这里是solrconfig.xml 中 Tika langid UpdateRequestProcessor的最小的配置.

<processor
class="org.apache.solr.update.processor.TikaLanguageIdentifierUpdateProcessorFactory">
<lst name="defaults">
<str name="langid.fl">title,subject,text,keywords</str>
<str name="langid.langField">language_s</str>
</lst>
</processor>

1.3 Configuring LangDetect Language Detection

  这里是solrconfig.xml中最小的LangDetect langid配置.

<processor
class="org.apache.solr.update.processor.LangDetectLanguageIdentifierUpdateProcessorFac
tory">
<lst name="defaults">
<str name="langid.fl">title,subject,text,keywords</str>
<str name="langid.langField">language_s</str>
</lst>
</processor>

1.4 langid Parameters

  正如上面所提到的,两个langid  UpdateRequestProcessor的实现具有相同的参数:

参数 类型 默认 必填 描述
langid Boolean true no 开启/关闭语言解析
langid.fl string none yes 逗号或者空格分隔的字段列表.用于语言探测解析.
langid.langField string none yes 对返回的语言代码指定字段
langid.langsField multivalued string none no 对返回的语言代码指定字段.如果使用langid.map.individual,每一个解析的语言都被添加到这个字段.
langid.overwrite Boolean false no 指定langField和langsField字段的内容是否被重写.如果它们包含值的话.
langid.lcmap string none false 空格分隔的列表,指定冒号分隔的语言代码用于语言解析.举例,你可以能使用这个映射中文,日文,韩文到一个cjk字段,并且映射美国英语和英国英语到一个en代码.可以使用langid.lcmap=ja:cjk zh:cjk ko:cjk
. This affects both the values put into the  en_GB:en en_US:en.这使这两个值放入到langField和langsField字段中.
langid.threshold float 0.5 no Specifies a threshold value between 0 and 1 that the language
identification score must reach before  accepts it. With longer langid
text fields, a high threshold such at 0.8 will give good results. For
shorter text fields, you may need to lower the threshold for language
identification, though you will be risking somewhat lower quality
results. We recommend experimenting with your data to tune your
results.
langid.whitelist string none no Specifies a list of allowed language identification codes. Use this in
combination with  to ensure that you only index langid.map
documents into fields that are in your schema.
langid.map Boolean false no Enables field name mapping. If true, Solr will map field names for all
fields listed in  . langid.fl
langid.map.fl string none no A comma-separated list of fields for  that is different langid.map
than the fields specified in  . langid.fl
langid.map.keepOrig Boolean false no If true, Solr will copy the field during the field name mapping process,
leaving the original field in place.
langid..map.individual Boolean false no If true, Solr will detect and map languages for each field individually
langid.map.individual.fl stromh none no 逗号分割的字段列表,使用 langid.map.individual.不同于langid.fl中指定的字段.
langid.fallbackFields string none no If no language is detected that meets the  score langid.threshold
, or if the detected language is not on the  , this langid.whitelist
field specifies language codes to be used as fallback values. If no
appropriate fallback languages are found, Solr will use the language
code specified in  .
langid.fallback string none no Specifies a language code to use if no language is detected or
specified in  .
langid.map.lcmap string determined by
langid.lcmap
no A space-separated list specifying colon delimited language code
mappings to use when mapping field names. For example, you might
use this to make Chinese, Japanese, and Korean language fields use
a common  suffix, and map both American and British English *_cjk
fields to a single  by using  *_en langid.map.lcmap=ja:cjk
. zh:cjk ko:cjk en_GB:en en_US:en
langid.map.pattern Java
regular
expression
none no By default, fields are mapped as <field>_<language>. To change this
pattern, you can specify a Java regular expression in this parameter.
langid.map.replace Java replace none no By default, fields are mapped as <field>_<language>. To change this
pattern, you can specify a Java replace in this parameter.
langid.enforceSchema Boolean true no If false, the  processor does not validate field names against langid
your schema. This may be useful if you plan to rename or delete
fields later in the UpdateChain

1.6.7 Detecting Languages During Indexing的更多相关文章

  1. 1.6 Indexing and Basic Data Operations--目录

    1.6.1 什么是 Indexing 1.6.2 Uploading Data with Index Handlers 1.6.3 Uploading Data with Solr Cell usin ...

  2. 1.5.8 语言分析器(Analyzer)

    语言分析器(Analyzer) 这部分包含了分词器(tokenizer)和过滤器(filter)关于字符转换和使用指定语言的相关信息.对于欧洲语言来说,tokenizer是相当直接的,Tokens被空 ...

  3. Importing/Indexing database (MySQL or SQL Server) in Solr using Data Import Handler--转载

    原文地址:https://gist.github.com/maxivak/3e3ee1fca32f3949f052 Install Solr download and install Solr fro ...

  4. Solr 6.7学习笔记(03)-- 样例配置文件 solrconfig.xml

    位于:${solr.home}\example\techproducts\solr\techproducts\conf\solrconfig.xml <?xml version="1. ...

  5. Solr基础知识二(导入数据)

    上一篇讲述了solr的安装启动过程,这一篇讲述如何导入数据到solr里. 一.准备数据 1.1 学生相关表 创建学生表.学生专业关联表.专业表.学生行业关联表.行业表.基础信息表,并创建一条小白的信息 ...

  6. Go Programming Language

    [Go Programming Language] 1.go run %filename 可以直接编译并运行一个文件,期间不会产生临时文件.例如 main.go. go run main.go 2.P ...

  7. Indexing Sensor Data

    In particular embodiments, a method includes, from an indexer in a sensor network, accessing a set o ...

  8. ESSENTIALS OF PROGRAMMING LANGUAGES (THIRD EDITION) :编程语言的本质 —— (一)

    # Foreword> # 序 This book brings you face-to-face with the most fundamental idea in computer prog ...

  9. 论文阅读(Xiang Bai——【CVPR2012】Detecting Texts of Arbitrary Orientations in Natural Images)

    Xiang Bai--[CVPR2012]Detecting Texts of Arbitrary Orientations in Natural Images 目录 作者和相关链接 方法概括 方法细 ...

随机推荐

  1. Linux下Python获取IP地址

    <lnmp一键安装包>中需要获取ip地址,有2种情况:如果服务器只有私网地址没有公网地址,这个时候获取的IP(即私网地址)不能用来判断服务器的位置,于是取其网关地址用来判断服务器在国内还是 ...

  2. Spark RDD概念学习系列之RDD的缓存(八)

      RDD的缓存 RDD的缓存和RDD的checkpoint的区别 缓存是在计算结束后,直接将计算结果通过用户定义的存储级别(存储级别定义了缓存存储的介质,现在支持内存.本地文件系统和Tachyon) ...

  3. 做 fzu oj 1106 题目学到的

    题目如下 这道题的意识就是给一个数问是否可以又阶乘之和构成,而难点主要是在于如果是7的话就是1!+3!,并不是单纯的从1的阶乘开始加,而是没顺序的,所以这题就得用到递归. (大概就是函数自己调用函数自 ...

  4. LeetCode 刷题记录(二)

    写在前面:因为要准备面试,开始了在[LeetCode]上刷题的历程.LeetCode上一共有大约150道题目,本文记录我在<http://oj.leetcode.com>上AC的所有题目, ...

  5. URAL 2070 Interesting Numbers (找规律)

    题意:在[L, R]之间求:x是个素数,因子个数是素数,同时满足两个条件,或者同时不满足两个条件的数的个数. 析:很明显所有的素数,因数都是2,是素数,所以我们只要算不是素数但因子是素数的数目就好,然 ...

  6. 基于jQuery的视频和音频播放器jPlayer

    jPlayer见网络上资料很少,官方英文资料太坑爹TAT,于是就写一个手记给大家参考下.据我观察,jPlayer的原理主要是用到HTML5,在不支持HTML5的浏览器上使用SWF.做到全兼容,这一点很 ...

  7. JavaScript 不重复的随机数

    在 JavaScript 中,一般产生的随机数会重复,但是有时我们需要不重复的随机数,如何实现?本文给于解决方法,需要的朋友可以参考下     在 JavaScript 中,一般产生的随机数会重复,但 ...

  8. BeanFactory和ApplicationContext的作用和区别

    BeanFactory和ApplicationContext的作用和区别 作用: 1. BeanFactory负责读取bean配置文档,管理bean的加载,实例化,维护bean之间的依赖关系,负责be ...

  9. Mysql用户密码设置修改和权限分配

    我的mysql安装在c:\mysql 一.更改密码 第一种方式: 1.更改之前root没有密码的情况 c:\mysql\bin>mysqladmin -u root password " ...

  10. Ext_两种处理服务器端返回值的方式

    1.Form表单提交返回值处理 //提交基本信息表单  f.form.submit({      clientValidation:true,      //表单提交后台处理地址      url:' ...