three supported reliability levels: * End-to-end * Store on failure * Best effort
https://github.com/cloudera/flume/blob/master/flume-docs/src/docs/UserGuide/Introduction
| === Reliability | |
| Reliability, the ability to continue delivering events in the face of | |
| failures without losing data, is a vital feature of Flume. Large | |
| distributed systems can and do suffer partial failures in many ways - | |
| physical hardware can fail, resources such as network bandwidth or | |
| memory can become scarce, or software can crash or run slowly. Flume | |
| emphasizes fault-tolerance as a core design principle and keeps | |
| running and collecting data even when many components have failed. | |
| Flume can guarantee that all data received by an agent node will | |
| eventually make it to the collector at the end of its flow as long as | |
| the agent node keeps running. That is, data can be *reliably* | |
| delivered to its eventual destination. | |
| However, reliable delivery can be very resource intensive and is often | |
| a stronger guarantee than some data sources require. Therefore, Flume | |
| allows the user to specify, on a per-flow basis, the level of | |
| reliability required. There are three supported reliability levels: | |
| * End-to-end | |
| * Store on failure | |
| * Best effort | |
| .A Note About Reliability | |
| ****************** | |
| Although Flume is extremely tolerant to machine, network, and software | |
| failures, there is never any such thing as '100% reliability'. If all | |
| the machines in a Flume installation were irrevocably destroyed in | |
| some terrible data center incident, all copies of Flume's data would | |
| be lost and there would be no way to recover them. Therefore all of | |
| Flume's reliability levels make guarantees about data delivery 'until | |
| some maximum number of failures have occurred'. Flume's failure modes | |
| - in terms of what can fail and what will keep running if they do - | |
| are described in detail later in this guide. | |
| ****************** | |
| The *end-to-end* reliability level guarantees that once Flume accepts | |
| an event, that event will make it to the endpoint - as long as the | |
| agent that accepted the event remains live long enough. The first | |
| thing the agent does in this setting is write the event to disk in a | |
| ''write-ahead log'' (WAL) so that, if the agent crashes and restarts, | |
| knowledge of the event is not lost. After the event has successfully | |
| made its way to the end of its flow, an acknowledgment is sent back to | |
| the originating agent so that it knows it no longer needs to store the | |
| event on disk. This reliability level can withstand any number of | |
| failures downstream of the initial agent. | |
| The *store on failure* reliability level causes nodes to only require | |
| an acknowledgement from the node one hop downstream. If the sending | |
| node detects a failure, it will store data on its local disk until the | |
| downstream node is repaired, or an alternate downstream destination | |
| can be selected. While this is effective, data can be lost if a | |
| compound or silent failure occurs. | |
| The *best-effort* reliability level sends data to the next hop with no | |
| attempts to confirm or retry delivery. If nodes fail, any data that | |
| they were in the process of transmitting or receiving can be | |
| lost. This is the weakest reliability level, but also the most | |
| lightweight. |
| === Reliability | |
| Reliability, the ability to continue delivering events in the face of | |
| failures without losing data, is a vital feature of Flume. Large | |
| distributed systems can and do suffer partial failures in many ways - | |
| physical hardware can fail, resources such as network bandwidth or | |
| memory can become scarce, or software can crash or run slowly. Flume | |
| emphasizes fault-tolerance as a core design principle and keeps | |
| running and collecting data even when many components have failed. | |
| Flume can guarantee that all data received by an agent node will | |
| eventually make it to the collector at the end of its flow as long as | |
| the agent node keeps running. That is, data can be *reliably* | |
| delivered to its eventual destination. | |
| However, reliable delivery can be very resource intensive and is often | |
| a stronger guarantee than some data sources require. Therefore, Flume | |
| allows the user to specify, on a per-flow basis, the level of | |
| reliability required. There are three supported reliability levels: | |
| * End-to-end | |
| * Store on failure | |
| * Best effort | |
| .A Note About Reliability | |
| ****************** | |
| Although Flume is extremely tolerant to machine, network, and software | |
| failures, there is never any such thing as '100% reliability'. If all | |
| the machines in a Flume installation were irrevocably destroyed in | |
| some terrible data center incident, all copies of Flume's data would | |
| be lost and there would be no way to recover them. Therefore all of | |
| Flume's reliability levels make guarantees about data delivery 'until | |
| some maximum number of failures have occurred'. Flume's failure modes | |
| - in terms of what can fail and what will keep running if they do - | |
| are described in detail later in this guide. | |
| ****************** | |
| The *end-to-end* reliability level guarantees that once Flume accepts | |
| an event, that event will make it to the endpoint - as long as the | |
| agent that accepted the event remains live long enough. The first | |
| thing the agent does in this setting is write the event to disk in a | |
| ''write-ahead log'' (WAL) so that, if the agent crashes and restarts, | |
| knowledge of the event is not lost. After the event has successfully | |
| made its way to the end of its flow, an acknowledgment is sent back to | |
| the originating agent so that it knows it no longer needs to store the | |
| event on disk. This reliability level can withstand any number of | |
| failures downstream of the initial agent. | |
| The *store on failure* reliability level causes nodes to only require | |
| an acknowledgement from the node one hop downstream. If the sending | |
| node detects a failure, it will store data on its local disk until the | |
| downstream node is repaired, or an alternate downstream destination | |
| can be selected. While this is effective, data can be lost if a | |
| compound or silent failure occurs. | |
| The *best-effort* reliability level sends data to the next hop with no | |
| attempts to confirm or retry delivery. If nodes fail, any data that | |
| they were in the process of transmitting or receiving can be | |
| lost. This is the weakest reliability level, but also the most | |
| lightweight. |
three supported reliability levels: * End-to-end * Store on failure * Best effort的更多相关文章
- SignalR Supported Platforms -摘自网络
SignalR is supported under a variety of server and client configurations. In addition, each transpor ...
- store操作
store.remove(rs); store.sync({ success: function (e, opt) { this.store.commitChanges(); }, failure: ...
- extjs 解决使用store.sync()方法更新item有时不触发后台action的问题
问题描述: extjs 解决使用store.sync()方法更新item有时不触发后台action,不出发后台action的原因是item的字段值没有变化 解决方法: item.setDirty(tr ...
- PMP用语集
AC actual cost 实际成本 ACWP actual cost of work performed 已完工作实际成本 BAC budget at completion 完工预算 BCWP b ...
- Solaris10安装配置LDAP(iPlanet Directory Server )
Solaris10安装光盘自带了iPlanet Directory Server安装包,系统管理员可以利用iPlanet Directory Server在Solaris系统创建一个LDAP Serv ...
- memory ordering 内存排序
Memory ordering - Wikipedia https://en.wikipedia.org/wiki/Memory_ordering https://zh.wikipedia.org/w ...
- Web测试介绍2一 安全测试
安全测试是在IT软件产品的生命周期中,特别是产品开发基本完成到发布阶段,对产品进行检验以验证产品符合安全需求定义和产品质量标准的过程. 主要安全需求包括: (i) 认证 Authent ...
- Python 3.6.0的sqlite3模块无法执行VACUUM语句
Python 3.6.0的sqlite3模块存在一个bug(见issue 29003),无法执行VACUUM语句. 一执行就出现异常: Traceback (most recent call last ...
- Flume1.5.0的安装、部署、简单应用(含伪分布式、与hadoop2.2.0、hbase0.96的案例)
目录: 一.什么是Flume? 1)flume的特点 2)flume的可靠性 3)flume的可恢复性 4)flume 的 一些核心概念 二.flume的官方网站在哪里? 三.在哪里下载? 四.如何安 ...
随机推荐
- XML布局文件于Java代码使用问题
2013-9-21 问题一.不同的XML文件中相同类型的控件id相同,那么将这些不同的布局xml组合在一个大的布局中,如何解决相同id问题 ? 解决办法: 不同的布局文件XML要组合成一个新的大布局, ...
- wxpython example
#!/usr/bin/env python #---------------------------------------------------------------------------- ...
- Ruby自动化测试(操作符的坑)
事情是这样的: times++ @ddr = DDR::DDR.new() 执行到这里的时候,总是报错:'+@' undefied method.刚开始的时候以为是机器在重启过程中一些不稳定函数调用或 ...
- BZOJ2243 [SDOI2011]染色(树链剖分+线段树合并)
题目链接 BZOJ2243 树链剖分 $+$ 线段树 线段树每个节点维护$lc$, $rc$, $s$ $lc$代表该区间的最左端的颜色,$rc$代表该区间的最右端的颜色 $s$代表该区间的所有连续颜 ...
- [翻译] NumSharp的数组切片功能 [:]
原文地址:https://medium.com/scisharp/slicing-in-numsharp-e56c46826630 翻译初稿(英文水平有限,请多包涵): 由于Numsharp新推出了数 ...
- Database | SQL
Basic of MySQL 创建数据库: mysql> create database xxj; Query OK, row affected (0.00 sec) 列举数据库: mysql& ...
- 一些yuv视频下载地址
因为测试需要下载一些yuv视频地址,现存一个可以下载yuv视频的地址以备后用 http://trace.eas.asu.edu/yuv/index.html ftp://ftp.ldv.e-techn ...
- rocketMq---------相关命令
搭建就不详细说了,cent7.x的系统,openJdk8,maven3.x,gradle4.10.2, git 1.8.3.1 直接下载相关的二进制压缩包,解压即用,方便. 下面看常用的管理命令 ro ...
- Eclipse工程中Java Build Path中的JDK版本和Java Compiler Compiler compliance level的区别(转)
在这里记录一下在eclipse中比较容易搞混淆和设置错误的地方.如下图所示的功能: 最精准的解释如下: Build Path是运行时环境 Compiler是编译时环境 假设,你的代码用到泛型,Bu ...
- Android Service实现双向通信(一)
首先,大概来总结一下与Service的通信方式有很多种: 通过BroadCastReceiver:这种方式是最简单的,只能用来交换简单的数据: 通过Messager:这种方式是通过一个传递一个Mess ...