Data Management Technology(5) -- Recovery
Recovery
Types of Failures
Wrong data entry
- Prevent by having constraints in the database
- Fix with data cleaning
Disk crashes
- Prevent by using redundancy (RAID, archive)
- Fix by using archives
Fire, theft, bankruptcy…
- Buy insurance, change profession…
System failures: most frequent (e.g. power)
- Use recovery
System Failures
Each transaction has internal state
When system crashes, internal state is lost
- Don’t know which parts executed and which didn’t
Remedy: use a log
- A file that records every single action of the transaction
Transactions
Assumption: the database is composed of elements
Usually 1 element = 1 block
Can be smaller (=1 record) or larger (=1 relation)
Assumption: each transaction reads/writes some elements
Correctness Principle
There exists a notion of correctness for the database
- Explicit constraints (e.g. foreign keys)
- Implicit conditions (e.g. sum of sales = sum of invoices)
Correctness principle: if a transaction starts in a correct database state, it ends in a correct database state
Consequence: we only need to guarantee that transactions are atomic, and the database will be correct forever
Primitive Operations of Transactions
INPUT(X)
- read element X to memory buffer
READ(X,t)
- copy element X to transaction local variable t
WRITE(X,t)
- copy transaction local variable t to element X
OUTPUT(X)
- write element X to disk
The Log
An append-only file containing log records
Note: multiple transactions run concurrently, log records are interleaved
After a system crash, use log to:
- Redo some transaction that didn’t commit
- Undo other transactions that didn’t commit
Undo Logging
Log records
transaction T has begun
T has committed
T has aborted
<T,X,v> T has updated element X, and its old value was v
Undo-Logging Rules
U1: If T modifies X, then <T,X,v> must be written to disk before X is written to disk
U2: If T commits, then must be written to disk only after all changes by T are written to disk
Hence: OUTPUTs are done early
Recovery with Undo Log
After system’s crash, run recovery manager
Idea 1. Decide for each transaction T whether it is completed or not
Idea 2. Undo all modifications by incompleted transactions
Recovery manager:
Read log from the end; cases:
- : mark T as completed
- : mark T as completed
- <T,X,v>: if T is not completed
then write X=v to disk
else ignore - : ignore
阅读方向,从下向上
Note: all undo commands are idempotent, If we perform them a second time, no harm is done
stop reading the log:
- We cannot stop until we reach the beginning of the log file
- This is impractical
- Better idea: use checkpointing
Checkpointing
Checkpoint the database periodically
- Stop accepting new transactions
- Wait until all curent transactions complete
- Flush log to disk
- Write a log record, flush
- Resume transactions
Redo Logging
Log records
<T,X,v>= T has updated element X, and its new value is v
R1: If T modifies X, then both <T,X,v> and must be written to disk before X is written to disk
Hence: OUTPUTs are done late
After system’s crash, run recovery manager
Step 1. Decide for each transaction T whether it is completed or not
Step 2. Read log from the beginning, redo all updates of committed transactions
Undo/Redo Logging
Log records, only one change
<T,X,u,v>= T has updated element X, its old value was u, and its new value is v
Recovery with Undo/Redo Log
After system’s crash, run recovery manager
Redo all committed transaction, top-down
Undo all uncommitted transactions, bottom-up
总结
日志undo先写日志(从下向上读)redo先写磁盘(从上到下读)
冲突可串行 & 两阶锁
两个事物使用同一个资源并有一个是写就是冲突的,简单讲就是在冲突可串行并发操作的前驱图中是没有环路的,前驱图无环就是冲突可串行的。
每个事物在使用资源的时候都是先统一取再统一放的,也就是其图示先增后减,斜率不会出现其他变动。
E.g.
Consider the following schedule:
T1 STARTS
T1 reads item B
T1 writes item B with old value 11, new value 12
T2 STARTS
T2 reads item B
T2 writes item B with old value 12, new value 13
T3 STARTS
T3 reads item A
T3 writes item A with old value 29, new value 30
T2 reads item A
T2 writes item A with old value 30, new value 31
T2 COMMITS
T1 reads item D
T1 writes item D with old value 44, new value 45
T3 COMMITS
T1 COMMITS
(a) What serial schedule is this equivalent to? If none, then explain why.
The serializability graph for the above schedule is: T1T2 T3. Any order that complies with the
topological order of the graph like T1 T3 T2 is an equivalent serial schedule for our schedule
(b) Is this schedule consistent with two phase locking? Explain why.
If we assume that all
transactions get the locks exactly before the operation and release them
afterwards, it is not consistent with two phase locking. This is
because T1 releases its lock on B after its second operation while
acquiring a lock on D at its last two operations. By removing the last
two operations of T1 the schedule becomes 2PL.
If we assume that the
transactions get all the locks they need at the beginning of the
transaction, and release them after the finish the operation, this
schedule will be 2PL. The minimum operations that could be added to the
schedule will be “T1 reads item A”. In this case, T1 has to acquire the
lock on A again after releasing its lock on A after its first
write.(这段话太深奥了,我用百度翻译都没看懂。。。)
Data Management Technology(5) -- Recovery的更多相关文章
- Data Management Technology(1) -- Introduction
1.Database concepts (1)Data & Information Information Is any kind of event that affects the stat ...
- Data Management Technology(3) -- SQL
SQL is a very-high-level language, in which the programmer is able to avoid specifying a lot of data ...
- Data Management Technology(2) -- Data Model
1.Data Model Model Is the abstraction of real world Reveal the essence of objects, help people to lo ...
- Data Management Technology(4) -- 关系数据库理论
规范化问题的提出 在规范化理论出现以前,层次和网状数据库的设计只是遵循其模型本身固有的原则,而无具体的理论依据可言,因而带有盲目性,可能在以后的运行和使用中发生许多预想不到的问题. 在关系数据库系统中 ...
- [Windows Azure] Data Management and Business Analytics
http://www.windowsazure.com/en-us/develop/net/fundamentals/cloud-storage/ Managing and analyzing dat ...
- Intel Active Management Technology
http://en.wikipedia.org/wiki/Intel_Active_Management_Technology Intel Active Management Technology F ...
- MySQL vs. MongoDB: Choosing a Data Management Solution
原文地址:http://www.javacodegeeks.com/2015/07/mysql-vs-mongodb.html 1. Introduction It would be fair to ...
- 场景3 Data Management
场景3 Data Management 数据管理 性能优化 OLTP OLAP 物化视图 :表的快照 传输表空间 :异构平台的数据迁移 星型转换 :事实表 OLTP : 在线事务处理 1. trans ...
- Data Management and Data Management Tools
Data Management ObjectivesBy the end o this module, you should understand the fundamentals of data m ...
随机推荐
- Hive時間函數-年份相加減
Hive時間函數-年份相加減 目前為止搜了很多资料,都没有找到Hive关于时间 年份,月份的处理信息,所以就自己想办法截取啦 本来是用了概数,一年365天去取几年前的日期,后来测试的发现不够精准,然后 ...
- Spring Boot常用注解和原理整理
一.启动注解 @SpringBootApplication @Target(ElementType.TYPE) @Retention(RetentionPolicy.RUNTIME) @Documen ...
- Redis 常用知识
Redis 介绍 Redis 一个开源的使用ANSI C语言编写.遵守BSD协议,内存中的数据结构存储系统,它可以用作Key-Value数据库.缓存和消息中间件,并提供多种语API. 支持多种数据结构 ...
- useradd命令详解(转)
1.作用 useradd或adduser命令用来建立用户帐号和创建用户的起始目录,使用权限是超级用户. 2.格式 useradd [-d home] [-s shell] [-c comment] [ ...
- layui table表格 表头与内容列错位问题(只有纵向滚动条的情况)
版本2.4.5 问题展示: 存在问题:正好错位一个纵向滚动条的宽度 思路: 仔细观察th元素及th包裹的子元素div 如下图 发现th宽度莫名的就多了5px 我就纳闷了 解决方案:到table.js ...
- FCC---CSS Flexbox: Add Flex Superpowers to the Tweet Embed
To the right is the tweet embed that will be used as the practical example. Some of the elements wou ...
- 一图了解 CODING 2.0:企业级持续交付解决方案
近日,CODING 在 KubeCon 2019 上海站上正式推出了 DevOps 的一站式解决方案:CODING 2.0. CODING 2.0 进行了产品.产品理念.功能.首页的升级,对用户服务进 ...
- 使用docker部署filebeat和logstash
想用filebeat读取项目的日志,然后发送logstash.logstash官网有相关的教程,但是docker部署的教程都太简洁了.自己折腾了半天,踩了不少坑,总算是将logstash和filebe ...
- screen工具的安装与使用
yum install screen 安装screen screen -S <作业名称> 创建新的页 screen -ls 查询已经存在的页面 screen -r < ...
- SQL语句性能调整原则
一.问题的提出 在应用系统开发初期,由于开发数据库数据比较少,对于查询SQL语句,复杂视图的的编写等体会不出SQL语句各种写法的性能优劣,但是如果将应用系统提交实际应用后,随着数据库中数据的增加,系统 ...