恼人的自旋锁续
1前言
昨天分享了一篇关于自旋锁的经典案例,升级上云之后性能衰减,原因是在 13 版本里面由于开启了 old_snapshot_threshold 参数。昨日文章发出去之后,我发现阅读的人数挺多,甚至把一些读者给"吓到了"。
233 灿总像个骨科医生一样..时不时的发个 pg 故障....吓得我都不敢用pg了.
大佬又在吓唬我们初学者,还敢不敢用?
本着严谨负责的态度,我今天完成了剩下的一些测试。
2现象
首先我在最新的 15.2 版本里面测试了一下👇 左边是未开启的性能,右边是开启后的,同样是 50 并发(昨天 13 版本里 50 并发 CPU 就接近饱和了),可以看到开启 old_snapshot_threshold 参数与否,基本没有太大的性能损耗。
这立马引起了我的兴趣,PostgreSQL 是如何解决这个问题的?我的第一反应是会不会和昨天说的关于 14 里面针对 GetSnapshotData 的优化有关?众所周知,14里面一个较大的提升是在海量连接情况下,性能较之前的版本得到了巨大的提升。
该特性涉及到以下几个 patch(我在运维经验谈里面详细聊过这块的优化,此处不过多赘述)
snapshot scalability: Don’t compute global horizons while building snapshots. snapshot scalability: Move PGXACT->xmin back to PGPROC. snapshot scalability: Move PGXACT->vacuumFlags to ProcGlobal->vacuumFlags. snapshot scalability: Move subxact info to ProcGlobal, remove PGXACT. snapshot scalability: Introduce dense array of in-progress xids. snapshot scalability: cache snapshots using a xact completion counter. Fix race condition in snapshot caching when 2PC is used.
简单搜了一下,果然有这块的记录!
To avoid regressing performance when old_snapshot_threshold is set (as that requires an accurate horizon to be computed), heap_page_prune_opt() doesn't unconditionally call TransactionIdLimitedForOldSnapshots() anymore. Both the computation of the limited horizon, and the triggering of errors (with SetOldSnapshotThresholdTimestamp()) is now only done when necessary to remove tuples.
为了避免在设置 old_snapshot_threshold 时性能下降(因为这需要计算准确的范围),heap_page_prune_opt() 不再无条件地调用 TransactionIdLimitedForOldSnapshots()。有限范围的计算和错误触发(使用 SetOldSnapshotThresholdTimestamp())现在仅在需要删除元组时进行。
Note: This contains a workaround in heap_page_prune_opt() to keep the snapshot_too_old tests working. While that workaround is ugly, the tests currently are not meaningful, and it seems best to address them separately.
注意:这里包含了一个变通方法,可以在 heap_page_prune_opt()中保持 snapshot_too_old 测试正常工作。虽然这种变通方法很难看,但是目前的测试并没有什么意义,似乎最好是分开处理它们。
简而言之是在做页剪枝的时候不再无条件调用 TransactionIdLimitedForOldSnapshots() 该函数了。于是我又下了一个 14 版本,测试结果也类似,开启与否,性能无明显衰减。
3小结
在 14 之后,海量连接情况下的吞吐量得到了极大提升,主要优化的逻辑在于GetSnapshotData() 这个函数上,同时也解决了当开启了 old_snapshot_threshold 参数之后性能急剧衰减的问题,页剪枝的时候不再无条件调用 TransactionIdLimitedForOldSnapshots()。
所以严谨一点的说法是:
13 及之前的版本不建议开启 old_snapshot_threshold 参数,总计四大坑,相较于另外几个无伤大雅的问题,性能衰减就显得尤为刺眼了 14 及以后的版本可以开启 old_snapshot_threshold 参数,性能衰减不明显
以前版本 GetSnapshotData() 这个函数经常会成为性能瓶颈,因此最佳实践是通过 pgbouncer/pgagroal/Odyssey 等连接池工具将连接数进行收敛,我们可能对 pgbouncer 挺熟悉,其实 Odyssey 的性能挺 nice,并且 pgpool 只能用到单核(或者 haproxy + multi pgbouncer?)。
Advanced multi-threaded PostgreSQL connection pooler and request router
同时在 14 里面还引入 idle_session_timeout ,定时查杀超过指定时间的空闲会话,假如是 14 以前的版本,可以使用 pg_timeout 插件。