1、背景

我们知道在sql中是可以实现 group by 字段a,字段b，那么这种效果在elasticsearch中该如何实现呢？此处我们记录在elasticsearch中的3种方式来实现这个效果。

2、实现多字段聚合的思路

图片来源：https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-terms-aggregation.html

从上图中，我们可以知道，可以通过3种方式来实现多字段的聚合操作。

3、需求

根据省(province)和性别(sex)来进行聚合，然后根据聚合后的每个桶的数据，在根据每个桶中的最大年龄(age)来进行倒序排序。

4、数据准备

4.1 创建索引

PUT /index_person

{

  "settings": {

    "number_of_shards": 1

  },

  "mappings": {

    "properties": {

      "id": {

        "type": "long"

      },

      "name": {

        "type": "keyword"

      },

      "province": {

        "type": "keyword"

      },

      "sex": {

        "type": "keyword"

      },

      "age": {

        "type": "integer"

      },

      "address": {

        "type": "text",

        "analyzer": "ik_max_word",

        "fields": {

          "keyword": {

            "type": "keyword",

            "ignore_above": 256

          }

        }

      }

    }

  }

}

4.2 准备数据

PUT /_bulk

{"create":{"_index":"index_person","_id":1}}

{"id":1,"name":"张三","sex":"男","age":20,"province":"湖北","address":"湖北省黄冈市罗田县匡河镇"}

{"create":{"_index":"index_person","_id":2}}

{"id":2,"name":"李四","sex":"男","age":19,"province":"江苏","address":"江苏省南京市"}

{"create":{"_index":"index_person","_id":3}}

{"id":3,"name":"王武","sex":"女","age":25,"province":"湖北","address":"湖北省武汉市江汉区"}

{"create":{"_index":"index_person","_id":4}}

{"id":4,"name":"赵六","sex":"女","age":30,"province":"北京","address":"北京市东城区"}

{"create":{"_index":"index_person","_id":5}}

{"id":5,"name":"钱七","sex":"女","age":16,"province":"北京","address":"北京市西城区"}

{"create":{"_index":"index_person","_id":6}}

{"id":6,"name":"王八","sex":"女","age":45,"province":"北京","address":"北京市朝阳区"}

5、实现方式

5.1 multi_terms实现

5.1.1 dsl

GET /index_person/_search

{

  "size": 0,

  "aggs": {

    "agg_province_sex": {

      "multi_terms": {

        "size": 10,

        "shard_size": 25,

        "order":{

          "max_age": "desc"

        },

        "terms": [

          {

            "field": "province",

            "missing": "defaultProvince"

          },

          {

            "field": "sex"

          }

        ]

      },

      "aggs": {

        "max_age": {

          "max": {

            "field": "age"

          }

        }

      }

    }

  }

}

5.1.2 java 代码

    @Test

    @DisplayName("多term聚合-根据省和性别聚合，然后根据最大年龄倒序")

    public void agg01() throws IOException {

        SearchRequest searchRequest = new SearchRequest.Builder()

                .size(0)

                .index("index_person")

                .aggregations("agg_province_sex", agg ->

                        agg.multiTerms(multiTerms ->

                                        multiTerms.terms(term -> term.field("province"))

                                                .terms(term -> term.field("sex"))

                                                .order(new NamedValue<>("max_age", SortOrder.Desc))

                                )

                                .aggregations("max_age", ageAgg ->

                                        ageAgg.max(max -> max.field("age")))

                )

                .build();

        System.out.println(searchRequest);

        SearchResponse<Object> response = client.search(searchRequest, Object.class);

        System.out.println(response);

    }

5.1.3 运行结果

5.2 script实现

5.2.1 dsl

GET /index_person/_search

{

  "size": 0,

  "runtime_mappings": {

    "runtime_province_sex": {

      "type": "keyword",

      "script": """

          String province = doc['province'].value;

          String sex = doc['sex'].value;

          emit(province + '|' + sex);

      """

    }

  },

  "aggs": {

    "agg_province_sex": {

      "terms": {

        "field": "runtime_province_sex",

        "size": 10,

        "shard_size": 25,

        "order": {

          "max_age": "desc"

        }

      },

      "aggs": {

        "max_age": {

          "max": {

            "field": "age"

          }

        }

      }

    }

  }

}

5.2.2 java代码

@Test

    @DisplayName("多term聚合-根据省和性别聚合，然后根据最大年龄倒序")

    public void agg02() throws IOException {

        SearchRequest searchRequest = new SearchRequest.Builder()

                .size(0)

                .index("index_person")

                .runtimeMappings("runtime_province_sex", field -> {

                    field.type(RuntimeFieldType.Keyword);

                    field.script(script -> script.inline(new InlineScript.Builder()

                            .lang(ScriptLanguage.Painless)

                            .source("String province = doc['province'].value;\n" +

                                    "          String sex = doc['sex'].value;\n" +

                                    "          emit(province + '|' + sex);")

                            .build()));

                    return field;

                })

                .aggregations("agg_province_sex", agg ->

                        agg.terms(terms ->

                                        terms.field("runtime_province_sex")

                                                .size(10)

                                                .shardSize(25)

                                                .order(new NamedValue<>("max_age", SortOrder.Desc))

                                )

                                .aggregations("max_age", minAgg ->

                                        minAgg.max(max -> max.field("age")))

                )

                .build();

        System.out.println(searchRequest);

        SearchResponse<Object> response = client.search(searchRequest, Object.class);

        System.out.println(response);

    }

5.2.3 运行结果

5.3 通过copyto实现

我本地测试过，通过copyto没实现，此处故先不考虑

5.5 通过pipeline来实现

实现思路：

创建mapping时，多创建一个字段pipeline_province_sex，该字段的值由创建数据时指定pipeline来生产。

5.4.1 创建mapping

PUT /index_person

{

  "settings": {

    "number_of_shards": 1

  },

  "mappings": {

    "properties": {

      "id": {

        "type": "long"

      },

      "name": {

        "type": "keyword"

      },

      "province": {

        "type": "keyword"

      },

      "sex": {

        "type": "keyword"

      },

      "age": {

        "type": "integer"

      },

      "pipeline_province_sex":{

        "type": "keyword"

      },

      "address": {

        "type": "text",

        "analyzer": "ik_max_word",

        "fields": {

          "keyword": {

            "type": "keyword",

            "ignore_above": 256

          }

        }

      }

    }

  }

}

此处指定了一个字段pipeline_province_sex，该字段的值会由pipeline来处理。

5.4.2 创建pipeline

PUT _ingest/pipeline/pipeline_index_person_provice_sex

{

  "description": "将provice和sex的值拼接起来",

  "processors": [

    {

      "set": {

        "field": "pipeline_province_sex",

        "value": ["{{province}}", "{{sex}}"]

      },

      "join": {

        "field": "pipeline_province_sex",

        "separator": "|"

      }

    }

  ]

}

5.4.3 插入数据

PUT /_bulk?pipeline=pipeline_index_person_provice_sex

{"create":{"_index":"index_person","_id":1}}

{"id":1,"name":"张三","sex":"男","age":20,"province":"湖北","address":"湖北省黄冈市罗田县匡河镇"}

{"create":{"_index":"index_person","_id":2}}

{"id":2,"name":"李四","sex":"男","age":19,"province":"江苏","address":"江苏省南京市"}

{"create":{"_index":"index_person","_id":3}}

{"id":3,"name":"王武","sex":"女","age":25,"province":"湖北","address":"湖北省武汉市江汉区"}

{"create":{"_index":"index_person","_id":4}}

{"id":4,"name":"赵六","sex":"女","age":30,"province":"北京","address":"北京市东城区"}

{"create":{"_index":"index_person","_id":5}}

{"id":5,"name":"钱七","sex":"女","age":16,"province":"北京","address":"北京市西城区"}

{"create":{"_index":"index_person","_id":6}}

{"id":6,"name":"王八","sex":"女","age":45,"province":"北京","address":"北京市朝阳区"}

注意： 此处的插入需要指定上一步的pipeline

PUT /_bulk?pipeline=pipeline_index_person_provice_sex

5.4.4 聚合dsl

GET /index_person/_search

{

  "size": 0,

  "aggs": {

    "agg_province_sex": {

      "terms": {

        "field": "pipeline_province_sex",

        "size": 10,

        "shard_size": 25,

        "order": {

          "max_age": "desc"

        }

      },

      "aggs": {

        "max_age": {

          "max": {

            "field": "age"

          }

        }

      }

    }

  }

}

5.4.5 运行结果

6、实现代码

https://gitee.com/huan1993/spring-cloud-parent/blob/master/es/es8-api/src/main/java/com/huan/es8/aggregations/bucket/MultiTermsAggs.java

7、参考文档

https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-terms-aggregation.html

elasticsearch多字段聚合实现方式的更多相关文章

elasticsearch 多字段聚合或者对字段子串聚合
以下是字段子串聚合,截取 'your_field' 前八位进行聚合的 Script script = new Script("doc['your_field'].getValue().sub ...
Elastic Stack之ElasticSearch分布式集群yum方式搭建
Elastic Stack之ElasticSearch分布式集群yum方式搭建作者:尹正杰版权声明:原创作品,谢绝转载!否则将追究法律责任. 一.搜索引擎及Lucene基本概念 1>.什么 ...
ElasticSearch6.0 高级应用之多字段聚合Aggregation（二）
ElasticSearch6.0 多字段聚合网上完整的资料很少 ,所以作者经过查阅资料,编写了聚合高级使用例子例子是根据电商搜索实际场景模拟出来的希望给大家带来帮助! 下面我们开始吧! 1. 创建 ...
跟我一起学extjs5(17--Grid金额字段单位MVVM方式的选择)
跟我一起学extjs5(17--Grid金额字段单位MVVM方式的选择) 这一节来完毕Grid中的金额字段的金额单位的转换.转换旰使用MVVM特性,整体上和控制菜单的几种模式类似.首先 ...
Dynamics CRM 通过Odata创建及更新记录各类型字段的赋值方式
CRM中通过Odata方式去创建或者更新记录时,各种类型的字段的赋值方式各不相同,这里转载一篇博文很详细的列出了各类型字段赋值方式,以供后期如有遗忘再次查询使用. http://luoyong0201 ...
Elastic Stack之ElasticSearch分布式集群二进制方式部署
Elastic Stack之ElasticSearch分布式集群二进制方式部署作者:尹正杰版权声明:原创作品,谢绝转载!否则将追究法律责任. 想必大家都知道ELK其实就是Elasticsearc ...
修改MySQL数据库中表和表中字段的编码方式的方法
今天向MySQL数据库中的一张表添加含有中文的数据,可是老是出异常,检查程序并没有发现错误,无奈呀,后来重新检查这张表发现表的编码方式为latin1并且原想可以插入中文的字段的编码方式也是latin1 ...
[Elasticsearch] 多字段搜索 (六) - 自定义_all字段，跨域查询及精确值字段
自定义_all字段在元数据:_all字段中,我们解释了特殊的_all字段会将其它所有字段中的值作为一个大字符串进行索引.尽管将所有字段的值作为一个字段进行索引并不是非常灵活.如果有一个自定义的_al ...
ElasticSearch 6.2 Mapping参数说明及text类型字段聚合查询配置
背景: 由于本人使用的是6.0以上的版本es,在使用发现很多中文博客对于mapping参数的说明已过时.ES6.0以后有很多参数变化. 现我根据官网总结mapping最新的参数,希望能对大家有用处. ...
【转】elasticsearch中字段类型默认显示{ "foo": { "type": "text", "fields": { "keyword": {"type": "keyword", "ignore_above": 256} }
官方原文链接:https://www.elastic.co/cn/blog/strings-are-dead-long-live-strings 转载原文连接:https://segmentfault ...

随机推荐

这三大特性，让 G1 取代了 CMS！
大家好,我是树哥. 之前我们聊过 CMS 回收器,但那时候我们说 CMS 回收器已经落伍了,现在应该是用 G1 回收器的时候了.那么 G1 回收器到底有什么魔力,它比 CMS 回收器相比强在哪里呢?今 ...
第四篇:理解vue代码
解释以下代码: 实现输入框中能够打字的功能 <el-input v-model="input" placeholder="在这打字"></el ...
变废为宝: 使用废旧手机实现实时监控方案(RTSP/RTMP方案)
随着手机淘汰的速度越来越快,大多数手机功能性能很强劲就不再使用了,以大牛直播SDK现有方案为例,本文探讨下,如何用废旧手机实现实时监控方案(把手机当摄像头做监控之用): 本方案需要准备一个手机作为采集 ...
KingbaseESV8R6 snapshot too old的配置和测试
背景书接上文,我们很好的理解了xmin和xid的区别.我们继续上文<KingbaseESV8R6不同隔离级下xmin的区别>来讨论 snapshot too old 的功能. 当king ...
zookeeper_mac安装总结
Zookeeper mac安装总结 1. 执行 brew install zookeeper 可能遇到报错 Error: The following directories are not writa ...
day05-线程的应用04
7.线程的应用03 7.4坦克大战5.0版增加功能: 我方坦克在发射的子弹消亡之后,才能发射新的子弹==>拓展:发射多颗子弹怎么办,控制一次最多只能发射5颗子弹让敌人坦克发射的子弹消亡之后, ...
阿色全息脑图，及制作软件AHMM
阿色全息脑图 AHMM 全息脑图是按照大系统观原理开发的新型思维工具,用于升维思考. 让您以系统的观点看待世界,专注系统的结构信息--全息,抓住事物的本质,透过表象和数据发现规律. 世间每项事物都是一 ...
异步编程promise
异步编程发展异步编程经历了 callback.promise.async/await.generator四个阶段,其中promise和async/await使用最为频繁,而generator因为语法 ...
Elasticsearch：设置Elastic账户安全
获取Docker容器名称和ID
docker ps --format "{{.Names}}" docker ps -q

elasticsearch多字段聚合实现方式