lucene源码分析(7)Analyzer分析

1.Analyzer的使用

Analyzer使用在IndexWriter的构造方法

 /**

   * Constructs a new IndexWriter per the settings given in <code>conf</code>.

   * If you want to make "live" changes to this writer instance, use

   * {@link #getConfig()}.

   *

   * <p>

   * <b>NOTE:</b> after ths writer is created, the given configuration instance

   * cannot be passed to another writer.

   *

   * @param d

   *          the index directory. The index is either created or appended

   *          according <code>conf.getOpenMode()</code>.

   * @param conf

   *          the configuration settings according to which IndexWriter should

   *          be initialized.

   * @throws IOException

   *           if the directory cannot be read/written to, or if it does not

   *           exist and <code>conf.getOpenMode()</code> is

   *           <code>OpenMode.APPEND</code> or if there is any other low-level

   *           IO error

   */

  public IndexWriter(Directory d, IndexWriterConfig conf) throws IOException {

    enableTestPoints = isEnableTestPoints();

    conf.setIndexWriter(this); // prevent reuse by other instances

    config = conf;

    infoStream = config.getInfoStream();

    softDeletesEnabled = config.getSoftDeletesField() != null;

    // obtain the write.lock. If the user configured a timeout,

    // we wrap with a sleeper and this might take some time.

    writeLock = d.obtainLock(WRITE_LOCK_NAME);

    boolean success = false;

    try {

      directoryOrig = d;

      directory = new LockValidatingDirectoryWrapper(d, writeLock);

      analyzer = config.getAnalyzer();

      mergeScheduler = config.getMergeScheduler();

      mergeScheduler.setInfoStream(infoStream);

      codec = config.getCodec();

      OpenMode mode = config.getOpenMode();

      final boolean indexExists;

      final boolean create;

      if (mode == OpenMode.CREATE) {

        indexExists = DirectoryReader.indexExists(directory);

        create = true;

      } else if (mode == OpenMode.APPEND) {

        indexExists = true;

        create = false;

      } else {

        // CREATE_OR_APPEND - create only if an index does not exist

        indexExists = DirectoryReader.indexExists(directory);

        create = !indexExists;

      }

      // If index is too old, reading the segments will throw

      // IndexFormatTooOldException.

      String[] files = directory.listAll();

      // Set up our initial SegmentInfos:

      IndexCommit commit = config.getIndexCommit();

      // Set up our initial SegmentInfos:

      StandardDirectoryReader reader;

      if (commit == null) {

        reader = null;

      } else {

        reader = commit.getReader();

      }

      if (create) {

        if (config.getIndexCommit() != null) {

          // We cannot both open from a commit point and create:

          if (mode == OpenMode.CREATE) {

            throw new IllegalArgumentException("cannot use IndexWriterConfig.setIndexCommit() with OpenMode.CREATE");

          } else {

            throw new IllegalArgumentException("cannot use IndexWriterConfig.setIndexCommit() when index has no commit");

          }

        }

        // Try to read first.  This is to allow create

        // against an index that's currently open for

        // searching.  In this case we write the next

        // segments_N file with no segments:

        final SegmentInfos sis = new SegmentInfos(Version.LATEST.major);

        if (indexExists) {

          final SegmentInfos previous = SegmentInfos.readLatestCommit(directory);

          sis.updateGenerationVersionAndCounter(previous);

        }

        segmentInfos = sis;

        rollbackSegments = segmentInfos.createBackupSegmentInfos();

        // Record that we have a change (zero out all

        // segments) pending:

        changed();

      } else if (reader != null) {

        // Init from an existing already opened NRT or non-NRT reader:

        if (reader.directory() != commit.getDirectory()) {

          throw new IllegalArgumentException("IndexCommit's reader must have the same directory as the IndexCommit");

        }

        if (reader.directory() != directoryOrig) {

          throw new IllegalArgumentException("IndexCommit's reader must have the same directory passed to IndexWriter");

        }

        if (reader.segmentInfos.getLastGeneration() == 0) {

          // TODO: maybe we could allow this?  It's tricky...

          throw new IllegalArgumentException("index must already have an initial commit to open from reader");

        }

        // Must clone because we don't want the incoming NRT reader to "see" any changes this writer now makes:

        segmentInfos = reader.segmentInfos.clone();

        SegmentInfos lastCommit;

        try {

          lastCommit = SegmentInfos.readCommit(directoryOrig, segmentInfos.getSegmentsFileName());

        } catch (IOException ioe) {

          throw new IllegalArgumentException("the provided reader is stale: its prior commit file \"" + segmentInfos.getSegmentsFileName() + "\" is missing from index");

        }

        if (reader.writer != null) {

          // The old writer better be closed (we have the write lock now!):

          assert reader.writer.closed;

          // In case the old writer wrote further segments (which we are now dropping),

          // update SIS metadata so we remain write-once:

          segmentInfos.updateGenerationVersionAndCounter(reader.writer.segmentInfos);

          lastCommit.updateGenerationVersionAndCounter(reader.writer.segmentInfos);

        }

        rollbackSegments = lastCommit.createBackupSegmentInfos();

      } else {

        // Init from either the latest commit point, or an explicit prior commit point:

        String lastSegmentsFile = SegmentInfos.getLastCommitSegmentsFileName(files);

        if (lastSegmentsFile == null) {

          throw new IndexNotFoundException("no segments* file found in " + directory + ": files: " + Arrays.toString(files));

        }

        // Do not use SegmentInfos.read(Directory) since the spooky

        // retrying it does is not necessary here (we hold the write lock):

        segmentInfos = SegmentInfos.readCommit(directoryOrig, lastSegmentsFile);

        if (commit != null) {

          // Swap out all segments, but, keep metadata in

          // SegmentInfos, like version & generation, to

          // preserve write-once.  This is important if

          // readers are open against the future commit

          // points.

          if (commit.getDirectory() != directoryOrig) {

            throw new IllegalArgumentException("IndexCommit's directory doesn't match my directory, expected=" + directoryOrig + ", got=" + commit.getDirectory());

          }

          SegmentInfos oldInfos = SegmentInfos.readCommit(directoryOrig, commit.getSegmentsFileName());

          segmentInfos.replace(oldInfos);

          changed();

          if (infoStream.isEnabled("IW")) {

            infoStream.message("IW", "init: loaded commit \"" + commit.getSegmentsFileName() + "\"");

          }

        }

        rollbackSegments = segmentInfos.createBackupSegmentInfos();

      }

      commitUserData = new HashMap<>(segmentInfos.getUserData()).entrySet();

      pendingNumDocs.set(segmentInfos.totalMaxDoc());

      // start with previous field numbers, but new FieldInfos

      // NOTE: this is correct even for an NRT reader because we'll pull FieldInfos even for the un-committed segments:

      globalFieldNumberMap = getFieldNumberMap();

      validateIndexSort();

      config.getFlushPolicy().init(config);

      bufferedUpdatesStream = new BufferedUpdatesStream(infoStream);

      docWriter = new DocumentsWriter(flushNotifications, segmentInfos.getIndexCreatedVersionMajor(), pendingNumDocs,

          enableTestPoints, this::newSegmentName,

          config, directoryOrig, directory, globalFieldNumberMap);

      readerPool = new ReaderPool(directory, directoryOrig, segmentInfos, globalFieldNumberMap,

          bufferedUpdatesStream::getCompletedDelGen, infoStream, conf.getSoftDeletesField(), reader);

      if (config.getReaderPooling()) {

        readerPool.enableReaderPooling();

      }

      // Default deleter (for backwards compatibility) is

      // KeepOnlyLastCommitDeleter:

      // Sync'd is silly here, but IFD asserts we sync'd on the IW instance:

      synchronized(this) {

        deleter = new IndexFileDeleter(files, directoryOrig, directory,

                                       config.getIndexDeletionPolicy(),

                                       segmentInfos, infoStream, this,

                                       indexExists, reader != null);

        // We incRef all files when we return an NRT reader from IW, so all files must exist even in the NRT case:

        assert create || filesExist(segmentInfos);

      }

      if (deleter.startingCommitDeleted) {

        // Deletion policy deleted the "head" commit point.

        // We have to mark ourself as changed so that if we

        // are closed w/o any further changes we write a new

        // segments_N file.

        changed();

      }

      if (reader != null) {

        // We always assume we are carrying over incoming changes when opening from reader:

        segmentInfos.changed();

        changed();

      }

      if (infoStream.isEnabled("IW")) {

        infoStream.message("IW", "init: create=" + create + " reader=" + reader);

        messageState();

      }

      success = true;

    } finally {

      if (!success) {

        if (infoStream.isEnabled("IW")) {

          infoStream.message("IW", "init: hit exception on init; releasing write lock");

        }

        IOUtils.closeWhileHandlingException(writeLock);

        writeLock = null;

      }

    }

  }

2.Analyzer的定义

/**

 * An Analyzer builds TokenStreams, which analyze text.  It thus represents a

 * policy for extracting index terms from text.

 * <p>

 * In order to define what analysis is done, subclasses must define their

 * {@link TokenStreamComponents TokenStreamComponents} in {@link #createComponents(String)}.

 * The components are then reused in each call to {@link #tokenStream(String, Reader)}.

 * <p>

 * Simple example:

 * <pre class="prettyprint">

 * Analyzer analyzer = new Analyzer() {

 *  {@literal @Override}

 *   protected TokenStreamComponents createComponents(String fieldName) {

 *     Tokenizer source = new FooTokenizer(reader);

 *     TokenStream filter = new FooFilter(source);

 *     filter = new BarFilter(filter);

 *     return new TokenStreamComponents(source, filter);

 *   }

 *   {@literal @Override}

 *   protected TokenStream normalize(TokenStream in) {

 *     // Assuming FooFilter is about normalization and BarFilter is about

 *     // stemming, only FooFilter should be applied

 *     return new FooFilter(in);

 *   }

 * };

 * </pre>

 * For more examples, see the {@link org.apache.lucene.analysis Analysis package documentation}.

 * <p>

 * For some concrete implementations bundled with Lucene, look in the analysis modules:

 * <ul>

 *   <li><a href="{@docRoot}/../analyzers-common/overview-summary.html">Common</a>:

 *       Analyzers for indexing content in different languages and domains.

 *   <li><a href="{@docRoot}/../analyzers-icu/overview-summary.html">ICU</a>:

 *       Exposes functionality from ICU to Apache Lucene.

 *   <li><a href="{@docRoot}/../analyzers-kuromoji/overview-summary.html">Kuromoji</a>:

 *       Morphological analyzer for Japanese text.

 *   <li><a href="{@docRoot}/../analyzers-morfologik/overview-summary.html">Morfologik</a>:

 *       Dictionary-driven lemmatization for the Polish language.

 *   <li><a href="{@docRoot}/../analyzers-phonetic/overview-summary.html">Phonetic</a>:

 *       Analysis for indexing phonetic signatures (for sounds-alike search).

 *   <li><a href="{@docRoot}/../analyzers-smartcn/overview-summary.html">Smart Chinese</a>:

 *       Analyzer for Simplified Chinese, which indexes words.

 *   <li><a href="{@docRoot}/../analyzers-stempel/overview-summary.html">Stempel</a>:

 *       Algorithmic Stemmer for the Polish Language.

 * </ul>

 */

可以看出，Analyzer针对不同的语言给出了不同的方式

其中，common抽象出所有Analyzer类，如下图所示

lucene源码分析(7)Analyzer分析的更多相关文章

Lucene 源码分析之倒排索引（三）
上文找到了 collect(-) 方法,其形参就是匹配的文档 Id,根据代码上下文,其中 doc 是由 iterator.nextDoc() 获得的,那 DefaultBulkScorer.itera ...
一个lucene源码分析的博客
ITpub上的一个lucene源码分析的博客,写的比较全面:http://blog.itpub.net/28624388/cid-93356-list-1/
lucene源码分析的一些资料
针对lucene6.1较新的分析:http://46aae4d1e2371e4aa769798941cef698.devproxy.yunshipei.com/conansonic/article/d ...
ArrayList源码和多线程安全问题分析
1.ArrayList源码和多线程安全问题分析在分析ArrayList线程安全问题之前,我们线对此类的源码进行分析,找出可能出现线程安全问题的地方,然后代码进行验证和分析. 1.1 数据结构 Arr ...
Okhttp3源码解析(3)-Call分析(整体流程)
### 前言前面我们讲了 [Okhttp的基本用法](https://www.jianshu.com/p/8e404d9c160f) [Okhttp3源码解析(1)-OkHttpClient分析]( ...
Okhttp3源码解析(2)-Request分析
### 前言前面我们讲了 [Okhttp的基本用法](https://www.jianshu.com/p/8e404d9c160f) [Okhttp3源码解析(1)-OkHttpClient分析]( ...
Spring mvc之源码 handlerMapping和handlerAdapter分析
Spring mvc之源码 handlerMapping和handlerAdapter分析本篇并不是具体分析Spring mvc,所以好多细节都是一笔带过,主要是带大家梳理一下整个Spring mv ...
HashMap的源码学习以及性能分析
HashMap的源码学习以及性能分析一).Map接口的实现类 HashTable.HashMap.LinkedHashMap.TreeMap 二).HashMap和HashTable的区别 1).H ...
ThreadLocal源码及相关问题分析
前言在高并发的环境下,当我们使用一个公共的变量时如果不加锁会出现并发问题,例如SimpleDateFormat,但是加锁的话会影响性能,对于这种情况我们可以使用ThreadLocal.ThreadL ...
物联网防火墙himqtt源码之MQTT协议分析
物联网防火墙himqtt源码之MQTT协议分析 himqtt是首款完整源码的高性能MQTT物联网防火墙 - MQTT Application FireWall,C语言编写,采用epoll模式支持数十万 ...

随机推荐

mdadm详细使用手册
1. 文档信息当前版本 1.2 创建人朱荣泽创建时间 2011.01.07 修改历史版本号时间内容 1.0 2011.01.07 创建<mdadm详细使用手册>1.0文档 1. ...
Delphi XE5 图解为Android应用制作签名
http://redboy136.blog.163.com/blog/static/107188432201381872820132 Delphi XE5 图解为Android应用制作签名 2013- ...
Python学习-37.Python中的正则表达式
作为一门现代语言,正则表达式是必不可缺的,在Python中,正则表达式位于re模块. import re 这里不说正则表达式怎样去匹配,例如\d代表数字,^代表开头(也代表非,例如^a-z则不匹配任何 ...
Javascript设计模式理论与实战：观察者模式
观察者模式主要应用于对象之间一对多的依赖关系,当一个对象发生改变时,多个对该对象有依赖的其他对象也会跟着做出相应改变,这就非常适合用观察者模式来实现.使用观察者模式可以根据需要增加或删除对象,解决一对 ...
ASP.NET Core真实管道详解[1]
ASP.NET Core管道虽然在结构组成上显得非常简单,但是在具体实现上却涉及到太多的对象,所以我们在 <ASP.NET Core管道深度剖析[共4篇]> 中围绕着一个经过极度简化的模拟 ...
C#存储过程调用的三个方法
//带参数的SQL语句 private void sql_param() { SqlConnection conn = new SqlConnection("server=WIN-OUD59 ...
附加属性来控制控件中，要扩展模块的visibility
可解决: 文本框控件中的按钮,DataGridColumnHeader中加入Filter控件... cs文件中的附加属性 + 样式文件中的 template+控件 -> visibility ...
Swift 如何像 C语言那样接收入口参数？
我们都知道在 Swift 语言当中不再有 main 函数了,可能了解过 C语言或者 Java 语言的同学对这一点赶到深深的不适.总之,取而代之的是 main.swift. int main(int a ...
falcon nodata 小坑一枚
按照官方文档配置完一切正常,唯独 nodata, 明明有正常的数据,但是为什么 nodata 会认为是没收到呢困扰许久,直到看了数据库中的数据才恍然大悟 falcon_portal库中的 hosts ...
Elasticsearch系列(五)----JAVA客户端之TransportClient操作详解
Elasticsearch JAVA操作有三种客户端: 1.TransportClient 2.JestClient 3.RestClient 还有种是2.3中有的NodeClient,在5.5.1中 ...

lucene源码分析(7)Analyzer分析

lucene源码分析(7)Analyzer分析的更多相关文章

随机推荐

热门专题