prometheus 监控项

此处记录prometheus监控项，exporter为 node_exporter

vim rules.yml

groups:

- name: node

  rules:

  - alert: server_status

    expr: up{job="node"} == 0

    for: 15s

    labels:

      severity: 'critical'

    annotations:

      summary: " node_exporter is down"

- name: cluster

  rules:

  - alert: CPU

    expr: (1-rate(node_cpu_seconds_total{mode="idle"}[1m]))*100 > 90

    for: 5s

    labels:

      severity: 'warning'

    annotations:

      summary: " cpu利用率超过 90%，{{ .Labels.name }}当前值: {{ $value }}%"

#  - alert: LOAD1

#    expr: node_load5 > Logical_CPU_core_total*0.3 or node_load1 > Logical_CPU_core_total*0.4 or node_load15 >  Logical_CPU_core_total*0.2

#    for: 5s

#    labels:

#      severity: 'critical'

#    annotations:

#      summary: " load过高 当前值为 {{ $value }}"

  - alert: LOAD1

    expr: node_load1 > Logical_CPU_core_total*3

    for: 5s

    labels:

      severity: 'warning'

    annotations:

      summary: " load1>cpu*3 当前值为 {{ $value }}"

  - alert: LOAD5

    expr:  node_load5 > Logical_CPU_core_total*2

    for: 5s

    labels:

      severity: 'warning'

    annotations:

      summary: " load5>cpu*2 当前值为 {{ $value }}"

  - alert: LOAD15

    expr: node_load15 >  Logical_CPU_core_total*2

    for: 5s

    labels:

      severity: 'warning'

    annotations:

      summary: " load15>cpu*2 当前值为 {{ $value }}"

  - alert: space_root

    expr: (1-node_filesystem_avail_bytes{fstype=~"xfs|ext4",mountpoint="/"}/node_filesystem_size_bytes{fstype=~"xfs|ext4",mountpoint="/"})*100 > 80

    for: 5s

    labels:

      severity: 'critical'

    annotations:

      summary: " /下空间使用率大于80%  当前值为{{ $value }}% "

  - alert: space_data

    expr: (1-node_filesystem_avail_bytes{fstype=~"xfs|ext4",mountpoint="/data"}/node_filesystem_size_bytes{fstype=~"xfs|ext4",mountpoint="/data"})*100 > 80

    for: 5s

    labels:

      severity: 'critical'

    annotations:

      summary: " /data空间使用率大于80% 当前值为{{ $value }}% "

  - alert: upload_rate

    expr: rate(node_network_transmit_bytes_total{device="eth0"}[1m])/1048576 > 10

    for: 5s

    labels:

      severity: 'warning'

    annotations:

      summary: " 上传速率大于10M 当前值为{{ $value }}M"

  - alert: download_rate

    expr: rate(node_network_receive_bytes_total{device="eth0"}[1m])/1048576 > 10

    for: 5s

    labels:

      severity: 'warning'

    annotations:

      summary: " 下载速率大于10M 当前值为{{ $value }}M "

  - alert: inode_size

    expr: (1-node_filesystem_files_free{fstype=~"xfs|ext4",mountpoint="/"}/node_filesystem_files{fstype=~"xfs|ext4",mountpoint="/"})*100 > 50

    for: 5s

    labels:

      severity: 'critical'

    annotations:

      summary: " /下inode使用率大于50% 当前值为{{ $value }}% "

  - alert: Memory_usage

    expr: (1-(node_memory_MemAvailable_bytes)/node_memory_MemTotal_bytes)*100 > 80

    for: 5s

    labels:

      severity: 'warning'

    annotations:

      summary: "内存使用率大于80% 当前值为{{ $value }}% "

  - alert: iowait

    expr: (avg by (instance) (rate(node_cpu_seconds_total{mode="iowait"}[5m])) * 100) > 50

    for: 5s

    labels:

      severity: 'critical'

    annotations:

      summary: "cpu iowait大于50% 当前值为{{ $value }}% "

  - alert: procs_zombie

    expr: procs_zombie > 20

    for: 5s

    labels:

      severity: 'critical'

    annotations:

      summary: " procs_zombie 大于20 当前值为{{ $value }} "

  - alert: logined_users

    expr: logined_users_total > 25

    for: 5s

    labels:

      severity: 'critical'

    annotations:

      summary: "logined_users 大于25 当前值为{{ $value }} "

prometheus 监控项的更多相关文章

prometheus 监控ElasticSearch核心指标
ES监控方案本文主要讲述使用 Prometheus监控ES,梳理核心监控指标并构建 Dashboard ,当集群有异常或者节点发生故障时,可以根据性能图表以高效率的方式进行问题诊断,再对核心指标筛选 ...
Prometheus Operator自定义监控项
Prometheus Operator默认的监控指标并不能完全满足实际的监控需求,这时候就需要我们自己根据业务添加自定义监控.添加一个自定义监控的步骤如下: 1.创建一个ServiceMonitor对 ...
prometheus node-exporter增加新的自定义监控项
项目中collector中新增加自己所需监控项即可定义启动node-exporter是传入的参数 var ( phpEndPoint = kingpin.Flag("collector.p ...
prometheus监控系统
关于Prometheus Prometheus是一套开源的监控系统,它将所有信息都存储为时间序列数据:因此实现一种Profiling监控方式,实时分析系统运行的状态.执行时间.调用次数等,以找到系统的 ...
Prometheus监控⼊⻔简介
文档目录: • prometheus是什么?• prometheus能为我们带来些什么• prometheus对于运维的要求• prometheus多图效果展示 1) Prometheus是什么pro ...
Prometheus监控学习笔记之Prometheus不完全避坑指南
0x00 概述 Prometheus 是一个开源监控系统,它本身已经成为了云原生中指标监控的事实标准,几乎所有 k8s 的核心组件以及其它云原生系统都以 Prometheus 的指标格式输出自己的运行 ...
Prometheus监控学习笔记之360基于Prometheus的在线服务监控实践
0x00 初衷最近参与的几个项目,无一例外对监控都有极强的要求,需要对项目中各组件进行详细监控,如服务端API的请求次数.响应时间.到达率.接口错误率.分布式存储中的集群IOPS.节点在线情况.偏移 ...
Grafana+Zabbix+Prometheus 监控系统
环境说明软件版本操作系统 IP地址 Grafana 5.4.3-1 Centos7.5 192.168.18.231 Prometheus 2.6.1 Centos7.5 192.168.18. ...
Kubernetes容器集群管理环境 - Prometheus监控篇
一.Prometheus介绍之前已经详细介绍了Kubernetes集群部署篇,今天这里重点说下Kubernetes监控方案-Prometheus+Grafana.Prometheus(普罗米修斯)是一 ...

随机推荐

mysql的最左索引匹配原则
最近复习数据库,主要看的是mysql.很多东西忘得一干二净.看到某乎上有个答案非常给力,就记录一下,以后方便查看. 链接:https://www.zhihu.com/question/36996520 ...
[转帖]Chrome 错误代码：ERR_UNSAFE_PORT
Chrome 错误代码:ERR_UNSAFE_PORT 2018年07月18日 09:07:50 孤舟听雨阅读数 182 https://blog.csdn.net/u013043762/artic ...
在子类中，若要调用父类中被覆盖的方法，可以使用super关键字
在子类中,若要调用父类中被覆盖的方法,可以使用super关键字. package text; class Parent { int x; public Parent() { ...
ubuntu 安装 TensorFlow、opencv3 的 tips
安装tensorflow: 创建tensorflow虚拟环境 conda create -n tensorflow python=2.7 输入命令查看可用版本的tensorflow-gpu cond ...
List<HashMap<String,String>> list, 根据hashmap中的某个键的值排序
来源https://blog.51cto.com/zhaodan/1725249 //可以使用Collections.sort(List list, Comparator c)来实现这里举例hash ...
datetime的timedelta对象
datetime.timedelta对象代表两个时间之间的时间差,两个date或datetime对象相减就可以返回一个timedelta对象. 如果有人问你昨天是几号,这个很容易就回答出来了.但是如果 ...
求x到y的最少计算次数（BFS）
时间限制:1秒空间限制:262144K 给定两个-100到100的整数x和y,对x只能进行加1,减1,乘2操作,问最少对x进行几次操作能得到y? 例如:a=3,b=11: 可以通过3*2*2-1,3 ...
js自执行函数
5.1对于函数表达式,在后面加括号即可以让函数立即执行:例如下面这个函数,至于为什么加了括号就可以立即执行,我们可以这么理解,就是像fn1():这样写的话,函数可以立即执行是没问题的,我们在经常会用 ...
EditPlus配置Java编译器
一.环境说明系统: windows 7 64位 editplus version: 4.3 二.设置步骤打开工具中的配置用户工具: 找到用户工具User tools,点击组名Group Name ...
sql：union 与union的使用和区别
SQL UNION 操作符 UNION 操作符用于合并两个或多个 SELECT 语句的结果集. 请注意,UNION 内部的 SELECT 语句必须拥有相同数量的列.列也必须拥有相似的数据类型.同时,每 ...

prometheus 监控项

prometheus 监控项的更多相关文章

随机推荐

热门专题