Google全球分布式数据库:Spanner
2012年的OSDI上google发布了Spanner数据库。个人认为Spanner对于版本控制,事务外部一致性的处理,使用TrueTime + Timestamp进行全球备份同步的实现都比较值得一看。个人认为对于其中时序逻辑的理解对在大范围内(通常是全国到全球)部署分布式DB以确保复制同步有重要意义。
key point:
external consistency -> txn sequence
truetime + timestamp, sync & multi-version
global deployment
2PC 2PL
3 basic txns(RW, RO, snapshot)
Spanner: Globally-Distributed Database
Implementation
Different environment: universe
test development production......
Hierarchy
- universe: global
The universe master and the placement driver are currently singletons.
- zone: manage deployment unit; logical & physical isolation
zone master & location proxy
- spanserver
- tablet
Spanserver
software stack
1 leader, server replica, in different data centers
all have:
tablet
$$
(key:string, timestamp:int64) → string
$$Colossus: a distributed filesystem like GFS
Paxos state machine: to support replication, for consistently replicated bag of mappings, replicas set: Paxos group
Each state machine stores its metadata and log in its corresponding tablet. Paxos implementation supports long-lived leaders with time-based leader leases.
Writes must initiate the Paxos protocol at the leader; reads access state directly from the underlying tablet at any replica that is sufficiently up-to-date.
Paxos: implementation pipelined, write in-order
leader uniquely has:
- lock table: the state for two-phase locking
- transaction manager: for distributed transactions, across Paxos group
Directories and Placement
based on k/v map, bucketing abstraction called a directory, which is a set of contiguous keys that share a common prefix.
tablet: different with bigtable, spanner tablet is a container that may encapsulate multiple partitions of the row space
Movedir: background, not a single txn, register fact and uses a transaction to atomically move small data!(actually the fragment, not a big dir)
Data model
- schematized semi-relational tables
- a query language
- generalpurpose transactions
Spanner’s data model is not purely relational, in that rows must have names.
hierarchies: in database schemas via the INTERLEAVE IN: get locality relationships.
TrueTime
API:
- now: return interval[earliest, latest]
- after
- before
underlying time references: GPS and atomic clocks
Concurrency Control
two-phase commit generates a Paxos write for the prepare phase that has no corresponding Spanner client write.
transactions:
- read-write: (including Standalone writes)
- read-only: without locking, any replica that is sufficiently up-to-date
- snapshot-reads: read in the past, no locking, any replica that is sufficiently up-to-date
Paxos leader lease:
timed leases: to make leadership long-lived, for lease votes
lease interval: [discover quorum of votes, no longer has votes]
Smax: the maximum timestamp used by a leader.
two-phase commit: a protocol maintain consistency - unsuccess: rollback
- prepare phase
- commit phase
RW txn:
buffered before written
wound-wait :avoid deadlock
both two have writing lock,
- non-coordinator participant leader
- coordinator leader: skip prepare phase
RO txn:
execution flow:
- assign a timestamp sread
- execute the transaction’s reads as snapshot reads at sread.
simply select sread = TT.now().latest
single Paxos group
Define LastTS() to be the timestamp of the last committed write at a Paxos group.
multiple Paxos groups
Schema-Change Transactions
Discussion
Paxos Truetime consistency
strong consistency cross data centers
data model: not pure relational(can use sql )
tablets are replicated, concurrtency corrtdiantion by Pxaos
txns with multiple Paxos groups --- 2PC coordination
leader
what's the actually difference compared with the classical distributed database?????
consistent versions of the data
the only reading data
the spirit kernel: the timestamp & version control
time mechenism
global-time consistency: timestamp no uncertainty
commit time: interval
there are two txns, to distinguish one happened actually before another
Participant leader -> Transaction manager -> Paxos group
three basic r/w ops, make the external consistency, global timestamp for sync across regions and certain txns sequences
Concurrency control : timestamp management to do
timestamp -> multi-version -> snapshot
almost all the work in spanner around the sequence of timestamp!
condition: multiple data centers
target: external consistency ~= linearizability
Two phase locking:
- growing phase: acquire lock
- shrinking phase: release lock
- 2PC: distributed system, global manage
- 2PL: one node, multi-txns, resource acquire and manage,
TrueTime: local clock -> global clock, which is essentially important for global distributed system because of sync needs.
uncertainty interval[earliest, latest]: try to make it as small as possible(increase accuracy) -> less lock -> increase efficiency
Thus, Timestamps + TrueTime can build a global accessible time service for all the application around the world.
external-consistency invariant: s1 < s2
Google全球分布式数据库:Spanner的更多相关文章
- 全球分布式数据库:Google Spanner(论文翻译)
本文由厦门大学计算机系教师林子雨翻译,翻译质量很高,本人只对极少数翻译得不太恰当的地方进行了修改. [摘要]:Spanner 是谷歌公司研发的.可扩展的.多版本.全球分布式.同步复制数据库.它是第一个 ...
- 全球级的分布式数据库 Google Spanner原理
开发四年只会写业务代码,分布式高并发都不会还做程序员?->>> Google Spanner简介 Spanner 是Google的全球级的分布式数据库 (Globally-Di ...
- 分布式数据库Google Spanner原理分析
Spanner 是Google的全球级的分布式数据库 (Globally-Distributed Database) .Spanner的扩展性达到了令人咋舌的全球级,可以扩展到数百万的机器,数已百计的 ...
- 怎样打造一个分布式数据库——rocksDB, raft, mvcc,本质上是为了解决跨数据中心的复制
摘自:http://www.infoq.com/cn/articles/how-to-build-a-distributed-database?utm_campaign=rightbar_v2& ...
- 这次,听人大教授讲讲分布式数据库的多级一致性|TDSQL 关键技术突破
近年来,凭借高可扩展.高可用等技术特性,分布式数据库正在成为金融行业数字化转型的重要支撑.分布式数据库如何在不同的金融级应用场景下,在确保数据一致性的前提下,同时保障系统的高性能和高可扩展性,是分布式 ...
- 云时代的分布式数据库:阿里分布式数据库服务DRDS
发表于2015-07-15 21:47| 10943次阅读| 来源<程序员>杂志| 27 条评论| 作者王晶昱 <程序员>杂志数据库DRDS分布式沈询 摘要:伴随着系统性能.成 ...
- 从NoSQL到NewSQL,谈交易型分布式数据库建设要点
在上一篇文章<从架构特点到功能缺陷,重新认识分析型分布式数据库>中,我们完成了对不同"分布式数据库"的横向分析,本文Ivan将讲述拆解的第二部分,会结合NoSQL与Ne ...
- 跨时代的分布式数据库 – 阿里云DRDS详解(转)
原文章地址:https://www.csdn.net/article/a/2015-08-28/15827676 跨时代的分布式数据库 – 阿里云DRDS详解 发表于2015-08-28 18:39| ...
- SDP(6):分布式数据库运算环境- Cassandra-Engine
现代信息系统应该是避不开大数据处理的.作为一个通用的系统集成工具也必须具备大数据存储和读取能力.cassandra是一种分布式的数据库,具备了分布式数据库高可用性(high-availability) ...
- 开源分布式数据库SequoiaDB在去哪儿网的实践
编者注: 中国的数据库行业也迎来了一波新的热点事件.分布式数据库这块新消息不断,也让大家开始关注中国的分布式数据库.首先是短短一周内,Pingcap和SequoiaDB巨杉数据库陆续宣布了C轮的数千万 ...
随机推荐
- Vue - 父子级的相互调用
父级调用子级 父级: <script> this.$refs.child.load(); 或 this.$refs.one.load(); </script> 子级: < ...
- Python Code_04InputFunction
代码部分 # coding:utf-8 # author : 写bug的盼盼 # development time : 2021/8/28 6:55 present = input('你想要什么?') ...
- Shell-循环-for-while
- Oracle 不同字符集复合索引长度验证
Oracle 不同字符集复合索引长度验证 背景 前段时间同事找到一个参数, 可以解决Oracle的char和byte 模式存储超长的问题. 很大程度上解决了研发修改SQL的工作量. 但是发现在某些字符 ...
- MYSQL varchar和nvarchar一些学习
MYSQL varchar和nvarchar一些学习 背景 先试用 utfmb3的格式进行一下简单验证 注意脚本都是一样的. create database zhaobsh ; use zhaobsh ...
- [转帖]win10多网卡指定ip走某个网卡的方案
https://zhuanlan.zhihu.com/p/571614314 我的电脑上有两个网卡,一个网卡A(网线),一个是网卡B(WIFI). 需求:网卡A和网卡B是不同的网络,网卡A已经把338 ...
- [转帖]重置 VCSA 6.7 root密码和SSO密码
问题描述 1.用root用户登录 VMware vCenter Server Appliance虚拟机失败,无法登录 2.vCenter Server Appliance 6.7 U1的root帐户错 ...
- [转帖]linux磁盘IO读写性能优化
在LINUX系统中,如果有大量读请求,默认的请求队列或许应付不过来,我们可以 动态调整请求队列数来提高效率,默认的请求队列数存放在/sys/block/xvda/queue/nr_requests 文 ...
- Python学习之三: 编译二进制
Python学习之三: 编译二进制 摘要 每次使用python 执行py文件其实是比较麻烦的 主要是还得安装python的虚拟机,以及安装对应的pip包. 感觉比较繁杂 理论上最快捷的方式是编译成 二 ...
- Nacos集群启动注意事项
简介 Nacos是阿里巴巴开源的一套服务注册发现的应用 使用简单灵活, 是spring Cloud Alibaba的组成部分 现在拆分微服务的部署情况下,极大的需求nacos服务作为支撑 单点情况下存 ...