MongoDB读数据策略
MongoDB是开源文档型NoSQL数据库,它的数据模型灵活,具有高扩展性、高可用性、易用性等特点,能够存储半结构化的数据,并且有丰富的查询语言和索引类型,当前MongoDB已广泛的用在各企业的核心业务系统中。MongoDB也是db-engines排名最高的非关系型数据库。
图片来源:db-engines
在MongoDB读取数据主要是受
当 oplog 同步到大多数节点时,对应节点的 snapshot 才会标记为 commmited,用户读取时,从最新的 commited 状态的 snapshot 读取数据,就能保证读到的数据一定已经同步到的大多数节点。那如何判断oplog 已经同步到大多数节点?
对于primary来说,当secondary 节点的oplog发生变化时,会通过命令将 oplog 进度立即通知给 primary,同时节点间的心跳消息里也会包含最新 oplog 的信息。这样primary 节点能很快知道数据是否已经同步到大多数节点的,并更新 snapshot 的状态。比如当t2已经写入到大多数据节点时,snapshot1、snapshot2都可以更新为 commited 状态。
对于secondary 节点来说,在拉取 oplog 时,primary 节点会将“最新的数据已同步到大多数节点的”的信息返回给 secondary 节点,然后secondary 节点通过这个oplog时间戳来更新自身的 snapshot 状态。
readpreference 读偏好设置 MongoDB 读控制策略除了readconcern策略外,还有readpreference 。 它主要控制数据库客户端驱动从哪个节点读取数据。 这个特性可以方便 地 实现读写分离、就近读取等策略。
readpreference 是由三部分组成,分别是mode、maxStalenessSeconds
、tag set
,其中mode支持五种类型,分别是:primary、primaryPreferred、Secondary、secondaryPreferred、nearest,我们先看几种模式的具体含义。
◀
几种模式介绍
▶
总结 通过上文介绍,我们知道MongoDB读数据策略,有readconcern和readpreference两个重要的概念。 其中readconcern是读数据时的数据一致性级别,它决定了决定读取数据时读到什么样的数据。 通常结合可用性和性能,会将readconcern设置为majority。 而readpreference决定读哪个节点的数据,主要用于实现读写分离上。 另外,MongoDB还提供了其他的配置选项,如写数据策略(writeconcern)这将在后面的文章中介绍。
作者介绍 司马辽太杰是 NineData 工程师。NineData 向企业和个人提供高效、安全的数据库SQL开发、数据库备份、数据复制/迁移/集成、数据对比等能力的产品,它是 开箱即用的 SaaS服务,可以快速提升企业SQL开发效率,保障企业数据安全。 近期,NineData 即将会 支持 MongoDB 、Redis等NoSQL数据库 。 NineData 官网地址: https://ninedata.cloud。 往期内容推荐:
2023,不一样的数据库
2022云数据库技术年度盘点
程序员必备的数据库知识2——Join 算法
程序员必备的数据库知识1——数据存储结构
NineData,领先的多云数据管理平台
read concern
(读策略)、
read preference
(读偏好设置 )两个参数控制,其中
readconcern
决定在读取副本集和分片集数据时的一致性和隔离性,而
readpreference
决定客户端驱动读取哪个数据节点的数据。它们的配合使用,可以提高MongoDB 集群的性能,以及在数据一致性和读性能上做平衡。
readconcern 一致性读策略
Readconcern 主要解决脏读问题,
从3.2版本后开始支持
。比如PSA集群,用户从 MongoDB 的 primary 上读取数据后,这条数据并没有同步从数节点,然后 primary 就故障了。此时不同的
Readconcern值
,MongoDB 返回数据的处理方式是不同的。
Readconcern
有几个不同的参数,分别是local、available、majority、linearizable、snapshot ,数据库在这些参数下的一致性是由弱到强递增的。
◀
几种模式介绍
▶
- Local
- Available
- Majority
Regardless of the read concern level, the most recent data on a node may not reflect the most recent version of the data in the system。
- linearizable
- snapshot
| 最新 oplog 时间戳 | snapshot | 状态 |
| t0 | snapshot0 | committed |
| t1 | snapshot1 | uncommitted |
| t2 | snapshot2 | uncommitted |
| t3 | snapshot3 | uncommitted |
readpreference 读偏好设置 MongoDB 读控制策略除了readconcern策略外,还有readpreference 。 它主要控制数据库客户端驱动从哪个节点读取数据。 这个特性可以方便 地 实现读写分离、就近读取等策略。
- primary
- primaryPreferred
- secondary
- secondaryPreferred
- nearest
primary
模式兼容,只能在其他四种模式下使用。
当选择了使用该参数控制读取数据,客户端会通过比较从节点和主节点的最后一次写时间来估计从节点的过期程度。客户端会把连接指向小于等于maxStalenessSeconds的从节点。
另外,需要注意
maxStalenessSeconds
最小值是90秒,如果小于该值将报错。
You must specify a◀ 标签集 ▶ 如果一个复制集中的成员有tag,就可以通过下面的办法读取到带有具体标签的成员上。 例如,如果某个节点有这样的成员标签:maxStalenessSecondsvalue of 90 seconds or longer: specifying a smallermaxStalenessSecondsvalue will raise an error.
{ "region": "South", "datacenter": "A" }
那么以下tag set可以将读操作指到上述成员(或具有相同标记的其他成员):
◀ 访问案例 ▶ 总结上面的内容,可以通过下面三种方式去定义不同的readpreference策略。[ { "region": "South", "datacenter": "A" }, { } ] // Find members with both tag values. If none are found, read from any eligible member.[ { "region": "South" }, { "datacenter": "A" }, { } ] // Find members with the specified region tag. Only if not found, then find members with the specified datacenter tag. If none are found, read from any eligible member.[ { "datacenter": "A" }, { "region": "South" }, { } ] // Find members with the specified datacenter tag. Only if not found, then find members with the specified region tag. If none are found, read from any eligible member.[ { "region": "South" }, { } ] // Find members with the specified region tag value. If none are found, read from any eligible member.[ { "datacenter": "A" }, { } ] // Find members with the specified datacenter tag value. If none are found, read from any eligible member.[ { } ] // Find any eligible member.
复制集访问方式:mongodb://db0.test.com,db1.test.com,db2.test.com/?replicaSet=myRepl&readPreference=secondaryPreferred&maxStalenessSeconds=150分片集群方式:mongodb://mongos1.test.com,mongos2.test.com/?readPreference=secondaryPreferred&maxStalenessSeconds=150带tag的定式:mongodb://mongos1.test.com/?readPreference=secondaryPreferred&readPreferenceTags=dc:ny,rack:r1&readPreferenceTags=dc:ny&readPreferenceTags=xxx
总结 通过上文介绍,我们知道MongoDB读数据策略,有readconcern和readpreference两个重要的概念。 其中readconcern是读数据时的数据一致性级别,它决定了决定读取数据时读到什么样的数据。 通常结合可用性和性能,会将readconcern设置为majority。 而readpreference决定读哪个节点的数据,主要用于实现读写分离上。 另外,MongoDB还提供了其他的配置选项,如写数据策略(writeconcern)这将在后面的文章中介绍。
作者介绍 司马辽太杰是 NineData 工程师。NineData 向企业和个人提供高效、安全的数据库SQL开发、数据库备份、数据复制/迁移/集成、数据对比等能力的产品,它是 开箱即用的 SaaS服务,可以快速提升企业SQL开发效率,保障企业数据安全。 近期,NineData 即将会 支持 MongoDB 、Redis等NoSQL数据库 。 NineData 官网地址: https://ninedata.cloud。 往期内容推荐:
2023,不一样的数据库
2022云数据库技术年度盘点
程序员必备的数据库知识2——Join 算法
程序员必备的数据库知识1——数据存储结构
NineData,领先的多云数据管理平台