spark 集成elasticsearch

pyspark读写elasticsearch依赖elasticsearch-hadoop包，需要首先在这里下载，版本号可以通过自行修改url解决。

"""

write data to elastic search

https://starsift.com/2018/01/18/integrating-pyspark-and-elasticsearch/

"""

from __future__ import print_function

import os

import json

from pyspark import SparkContext

from pyspark import SparkConf

os.environ["PYSPARK_PYTHON"] = "/usr/bin/python3"

os.environ["PYSPARK_DRIVER_PYTHON"] = "/usr/bin/python3"

# pay attention here, jars could be added at here

os.environ['PYSPARK_SUBMIT_ARGS'] = \

    '--jars /home/buxizhizhoum/2-Learning/pyspark_tutorial/jars/elasticsearch-hadoop-6.4.2/dist/elasticsearch-spark-20_2.11-6.4.2.jar ' \

    '--packages org.apache.spark:spark-streaming-kafka-0-8_2.11:2.3.1 ' \

    'pyspark-shell'

conf = SparkConf().setAppName("write_es").setMaster("local[2]")

sc = SparkContext(conf=conf)

# config refer: https://www.elastic.co/guide/en/elasticsearch/hadoop/current/configuration.html

es_write_conf = {

    # specify the node that we are sending data to (this should be the master)

    "es.nodes": 'localhost',

    # specify the port in case it is not the default port

    "es.port": '',

    # specify a resource in the form 'index/doc-type'

    "es.resource": 'testindex/testdoc',

    # is the input JSON?

    "es.input.json": "yes",

    # is there a field in the mapping that should be used to specify the ES document ID

    "es.mapping.id": "doc_id"

}

if __name__ == "__main__":

    data = [

        {'': '', 'doc_id': 1},

        {'': '', 'doc_id': 2},

        {'': '', 'doc_id': 3},

        {'': '', 'doc_id': 4},

        {'': '', 'doc_id': 5},

        {'': '', 'doc_id': 6},

        {'': '', 'doc_id': 7},

        {'': '', 'doc_id': 8},

        {'': '', 'doc_id': 9},

        {'': '', 'doc_id': 10}

    ]

    rdd = sc.parallelize(data)

    rdd = rdd.map(lambda x: (x["doc_id"], json.dumps(x)))

    rdd.saveAsNewAPIHadoopFile(

        path='-',

        outputFormatClass="org.elasticsearch.hadoop.mr.EsOutputFormat",

        keyClass="org.apache.hadoop.io.NullWritable",

        valueClass="org.elasticsearch.hadoop.mr.LinkedMapWritable",

        # critically, we must specify our `es_write_conf`

        conf=es_write_conf)

refer: https://starsift.com/2018/01/18/integrating-pyspark-and-elasticsearch/

spark 集成elasticsearch的更多相关文章

Spark：利用Eclipse构建Spark集成开发环境
前一篇文章“Apache Spark学习:将Spark部署到Hadoop 2.2.0上”介绍了如何使用Maven编译生成可直接运行在Hadoop 2.2.0上的Spark jar包,而本文则在此基础上 ...
使用spark访问elasticsearch的数据
使用spark访问elasticsearch的数据,前提是spark能访问hive,hive能访问es http://blog.csdn.net/ggz631047367/article/detail ...
spark集成hive遭遇mysql check失败的问题
问题: spark集成hive,启动spark-shell或者spark-sql的时候,报错: INFO MetaStoreDirectSql: MySQL check failed, assumin ...
springboot集成elasticsearch
在基础阶段学习ES一般是首先是安装ES后借助 Kibana 来进行CURD 了解ES的使用: 在进阶阶段可以需要学习ES的底层原理,如何通过Version来实现乐观锁保证ES不出问题等核心原理: 第 ...
Spark 整合ElasticSearch
Spark 整合ElasticSearch 因为做资料搜索用到了ElasticSearch,最近又了解一下 Spark ML,先来演示一个Spark 读取/写入 ElasticSearch 简单示例. ...
使用Logstash同步数据至Elasticsearch，Spring Boot中集成Elasticsearch实现搜索
安装logstash.同步数据至ElasticSearch 为什么使用logstash来同步,CSDN上有一篇文章简要的分析了以下几种同步工具的优缺点:https://blog.csdn.net/la ...
Spark搭档Elasticsearch
Spark与elasticsearch结合使用是一种常用的场景,小编在这里整理了一些Spark与ES结合使用的方法.一. write data to elasticsearch利用elasticsea ...
SpringBoot 集成 Elasticsearch
前面在 ubuntu 完成安装 elasticsearch,现在我们SpringBoot将集成elasticsearch. 1.创建SpringBoot项目我们这边直接引入NoSql中Spring ...
数据湖应用解析：Spark on Elasticsearch一致性问题
摘要:脏数据对数据计算的正确性带来了很严重的影响.因此,我们需要探索一种方法,能够实现Spark写入Elasticsearch数据的可靠性与正确性. 概述 Spark与Elasticsearch(es ...

随机推荐

搭建好lamp，部署owncloud。
先上传两个文件上传到root目录上传到opt目录下 #mkdir /opt/dvd1 #mount /dev/sr0 /opt/dvd1 #cd /etc/yum.repos.d/ #vi dvd ...
C/C++ 与 Python 的通信
作者:Jerry Jho链接:https://www.zhihu.com/question/23003213/answer/56121859来源:知乎著作权归作者所有.商业转载请联系作者获得授权,非商 ...
CentOS7.3编译hadoop2.7.3源码
在使用hive或者是kylin时,可以选择文件的压缩格式,但是这个需要有hadoop native库的支持,默认情况下,hadoop官方发布的二进制包中是不包含native库的,所以无法使用一些压缩相 ...
Linux使用NFS服务实现远程共享
首先安装 apt install -y nfs-kernel-server nfs-common 编辑配置文件 vim /etc/exports 添加内容: /root/test *(rw,sync, ...
表单（同步提交）和AJAX（异步提交）示范
表单提交(同步提交) HTML文件: PHP文件: 这样就能接收到HTML里输入的内容,注意: FORM表头method为POST,PHP文件获取的方法就是$_POST,method为GET,PHP的 ...
C Mysql API连接Mysql
最近都在查看MYsql C API文档,也遇到了很多问题,下面来简单的做一个总结. mysql多线程问题 mysql多线程处理不好,经常会发生coredump,见使用Mysql出core一文. 单线程 ...
asp.net网站中增删文件夹会导致Session或cache等等丢失
因为这会导致网站资源本身重新加载. 如果要改变文件和文件夹,一般应该是对 app_data 下进行操作.
day15(模块引用笔记)
import spam文件名是spam.py,模块名则是spam# 首次导入模块发生?件事# 1. 会产生一个模块的名称空间# 2. 执行文件spam.py,将执行过程中产生的名字都放到模块的名称空间 ...
Java中的Map List Set等集合类
一.概述二 set map list的区别三. Collections类和Collection接口四. List接口,有序可重复的集合五. Set接口,代表无序,不可重复的集合六. Map接 ...
c# Color 颜色设置
#region 笔刷颜色 public SolidColorBrush col(byte a, byte r, byte g, byte b) { #region 颜色说明 //Color.FromA ...

spark 集成elasticsearch

spark 集成elasticsearch的更多相关文章

随机推荐

热门专题