flume中sink到hdfs，文件系统频繁产生文件和出现乱码，文件滚动配置不起作用？

　　问题描述

　解决办法

　　先把这个hdfs目录下的数据删除。并修改配置文件flume-conf.properties，重新采集。

# Licensed to the Apache Software Foundation (ASF) under one

# or more contributor license agreements.  See the NOTICE file

# distributed with this work for additional information

# regarding copyright ownership.  The ASF licenses this file

# to you under the Apache License, Version 2.0 (the

# "License"); you may not use this file except in compliance

# with the License.  You may obtain a copy of the License at

#

#  http://www.apache.org/licenses/LICENSE-2.0

#

# Unless required by applicable law or agreed to in writing,

# software distributed under the License is distributed on an

# "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY

# KIND, either express or implied.  See the License for the

# specific language governing permissions and limitations

# under the License.

# The configuration file needs to define the sources,

# the channels and the sinks.

# Sources, channels and sinks are defined per agent,

# in this case called 'agent'

agent1.sources = spool-source1

agent1.sinks = hdfs-sink1

agent1.channels = ch1

#Define and configure an Spool directory source

agent1.sources.spool-source1.channels=ch1

agent1.sources.spool-source1.type=spooldir

agent1.sources.spool-source1.spoolDir=/home/hadoop/data/flume/sqooldir//--/

agent1.sources.spool-source1.ignorePattern=event(_\d{}\-d{}\-d{}\_d{}\_d{})?\.log(\.COMPLETED)?

agent1.sources.spool-source1.deserializer.maxLineLength=

#Configure channel

agent1.channels.ch1.type = file

agent1.channels.ch1.checkpointDir = /home/hadoop/data/flume/checkpointDir

agent1.channels.ch1.dataDirs = /home/hadoop/data/flume/dataDirs

#Define and configure a hdfs sink

agent1.sinks.hdfs-sink1.channel = ch1

agent1.sinks.hdfs-sink1.type = hdfs

agent1.sinks.hdfs-sink1.hdfs.path = hdfs://master:9000/flume/%Y%m%d

agent1.sinks.hdfs-sink1.hdfs.useLocalTimeStamp = true

agent1.sinks.hdfs-sink1.hdfs.rollInterval =

agent1.sinks.hdfs-sink1.hdfs.rollSize =

agent1.sinks.hdfs-sink1.hdfs.rollCount =

agent1.sinks.hdfs-sink1.hdfs.minBlockReplicas=

agent1.sinks.hdfs-sink1.hdfs.idleTimeout=

#agent1.sinks.hdfs-sink1.hdfs.codeC = snappy

agent1.sinks.hdfs-sink1.hdfs.fileType=DataStream

#agent1.sinks.hdfs-sink1.hdfs.writeFormat=Text

# For each one of the sources, the type is defined

#agent.sources.seqGenSrc.type = seq

# The channel can be defined as follows.

#agent.sources.seqGenSrc.channels = memoryChannel

# Each sink's type must be defined

#agent.sinks.loggerSink.type = logger

#Specify the channel the sink should use

#agent.sinks.loggerSink.channel = memoryChannel

# Each channel's type is defined.

#agent.channels.memoryChannel.type = memory

# Other config values specific to each type of channel(sink or source)

# can be defined as well

# In this case, it specifies the capacity of the memory channel

#agent.channels.memoryChannel.capacity =

　　教大家一招：大家在这些如flume的配置文件，最好还是去看官网，学会扩展，别只局限于别人的博客的文档，当然可以作为参考。关键还是来源于官方！

[hadoop@master sqooldir]$ $HADOOP_HOME/bin/hadoop fs -rm -r /flume/

// :: WARN util.NativeCodeLoader: Unable to load native-hadoop library for your platform... using builtin-java classes where applicable

// :: INFO fs.TrashPolicyDefault: Namenode trash configuration: Deletion interval =  minutes, Emptier interval =  minutes.

Deleted /flume/

[hadoop@master sqooldir]$

　　重新开启flume

[hadoop@master flume]$ pwd

/home/hadoop/app/flume

[hadoop@master flume]$ bin/flume-ng agent -n agent1 -f conf/flume-conf.properties

　　如果你的问题，还有副本数的问题，自行去解决。将$HADOOP_HOME/etc/hadoop/下的hdfs-site.xml的属性（master、slave1和slave2都要修改）

<property>

                <name>dfs.replication</name>

                <value></value>

                <description>Set to  for pseudo-distributed mode,Set to  for distributed mode,Set to  for distributed mode.</description>

 </property>

　　记得重启hadoop集群。

flume中sink到hdfs，文件系统频繁产生文件和出现乱码，文件滚动配置不起作用？的更多相关文章

flume中sink到hdfs，文件系统频繁产生文件，文件滚动配置不起作用？
在测试hdfs的sink,发现sink端的文件滚动配置项起不到任何作用,配置如下: a1.sinks.k1.type=hdfs a1.sinks.k1.channel=c1 a1.sinks.k1.h ...
HDFS文件系统上传时序图 PB级文件存储时序图
自己设计的时序图. 来自为知笔记(Wiz)
大数据学习笔记之Hadoop（二）：HDFS文件系统
文章目录一 HDFS概念 1.1 概念 1.2 组成 1.3 HDFS 文件块大小二 HFDS命令行操作三 HDFS客户端操作 3.1 eclipse环境准备 3.1.1 jar包准备 3.2 ...
Flume中的HDFS Sink配置参数说明【转】
转:http://lxw1234.com/archives/2015/10/527.htm 关键字:flume.hdfs.sink.配置参数 Flume中的HDFS Sink应该是非常常用的,其中的配 ...
Flume监听文件目录sink至hdfs配置
一:flume介绍 Flume是一个分布式.可靠.和高可用的海量日志聚合的系统,支持在系统中定制各类数据发送方,用于收集数据:同时,Flume提供对数据进行简单处理,并写到各种数据接受方(可定制)的能 ...
Flume实时监控目录sink到hdfs，再用sparkStreaming监控hdfs的这个目录，对数据进行计算
目标:Flume实时监控目录sink到hdfs,再用sparkStreaming监控hdfs的这个目录,对数据进行计算 1.flume的配置,配置spoolDirSource_hdfsSink.pro ...
在Spark shell中基于HDFS文件系统进行wordcount交互式分析
Spark是一个分布式内存计算框架,可部署在YARN或者MESOS管理的分布式系统中(Fully Distributed),也可以以Pseudo Distributed方式部署在单个机器上面,还可以以 ...
我理解中的Hadoop HDFS分布式文件系统
一,什么是分布式文件系统,分布式文件系统能干什么在学习一个文件系统时,首先我先想到的是,学习它能为我们提供什么样的服务,它的价值在哪里,为什么要去学它.以这样的方式去理解它之后在日后的深入学习中才能 ...
将存储在本地的大量分散的小文件，合并并保存在hdfs文件系统中
import java.io.BufferedInputStream; import java.io.File; import java.io.FileInputStream; import java ...

随机推荐

相机拍照友盟检测crash是为什么？
友盟报错如下* setObjectForKey: object cannot be nil (key: UIImagePickerControllerOriginalImage)(null)(( 0 ...
【转载】eclipse中批量修改Java类文件中引入的package包路径
原博客地址:http://my.oschina.net/leeoo/blog/37852 当复制其他工程中的包到新工程的目录中时,由于包路径不同,出现红叉,下面的类要一个一个修改包路径,类文件太多的话 ...
Service和Servlet的区别
1. 整体概念 Servlet是Java对于Web开发而产生的一项技术,可以说Servlet技术是Java专有的,它是服务器端的技术,客户端通常是浏览器,Servlet提供了请求/响应模式,是JAVA ...
第一性原理：First principle thinking是什么？
作者:沧海桑田链接:https://www.zhihu.com/question/40550274/answer/225236964来源:知乎著作权归作者所有.商业转载请联系作者获得授权,非商业转载请 ...
运维派企业面试题1 监控MySQL主从同步是否异常
Linux运维必会的实战编程笔试题(19题) 企业面试题1:(生产实战案例):监控MySQL主从同步是否异常,如果异常,则发送短信或者邮件给管理员.提示:如果没主从同步环境,可以用下面文本放到文件里读 ...
caffe(5) 其他常用层及参数
本文讲解一些其它的常用层,包括:softmax_loss层,Inner Product层,accuracy层,reshape层和dropout层及其它们的参数配置. 1.softmax-loss so ...
Intel NUC迷你机2019年底迎来i9 8核心16线程
Intel处理器这两年全年提速,虽然10nm新工艺受阻,但核心数在全面增加,从发烧到桌面到低功耗莫不如此,如今连NUC迷你机也要全新进化了,一年多之后就会迎来8核心16线程,而且也划入i9序列. 根据 ...
Linux-批量添加用户stu01..stu03,并设置固定的密码123456 (要求不能使用循环for while)
最终目标: useradd stu01;echo 123456|passwd --stdin stu01 useradd stu02;echo 123456|passwd --stdin stu02 ...
OpenJDK源码研究笔记(六)--观察者模式工具类(Observer和Observable)和应用示例
本文主要讲解OpenJDK观察者模式的2个工具类,java.util.Observer观察者接口,java.util.Observable被观察者基类. 然后,给出了一个常见的观察者应用示例. Obs ...
gcc/g++命令参数笔记
1. gcc -E source_file.c -E,只执行到预编译.直接输出预编译结果. 2. gcc -S source_file.c -S,只执行到源代码到汇编代码的转换,输出汇编代码. 3. ...

flume中sink到hdfs，文件系统频繁产生文件和出现乱码，文件滚动配置不起作用？

flume中sink到hdfs，文件系统频繁产生文件和出现乱码，文件滚动配置不起作用？的更多相关文章

随机推荐

热门专题