PostgreSQL 19 preview - bufmgr(缓冲区管理器) 优化
PostgreSQL 19 preview - bufmgr(缓冲区管理器) 优化
Andres Freund 一次提交了5个bufmgr相关的patch:
1、 https://github.com/postgres/postgres/commit/819dc118c0f6cd1fc08f0e807702b4bc0b0d8733
2、 https://github.com/postgres/postgres/commit/3c2b97b29ee33d5619779fd10e06eee07d4700da
3、 https://github.com/postgres/postgres/commit/2a2e1b470b9b4e9c7d76bc227db02a2efbb57473
4、 https://github.com/postgres/postgres/commit/3baae90013df58a8d124afa79df07075b8ebea09
5、 https://github.com/postgres/postgres/commit/5e89985928795f243dc287210c2aa016dfd00bfe
Improve ReadRecentBuffer() scalability
While testing a new potential use for ReadRecentBuffer(), Andres
reported that it scales badly when called concurrently for the same
buffer by many backends. Instead of a naive (but wrong) coding with
PinBuffer(), it used the spinlock, so that it could be careful to pin
only if the buffer was valid and holding the expected block, to avoid
breaking invariants in eg GetVictimBuffer(). Unfortunately that made it
less scalable than PinBuffer(), which uses compare-exchange instead. We can fix that by giving PinBuffer() a new skip_if_not_valid mode that
doesn't pin invalid buffers. It might occasionally skip when it
shouldn't due to the unlocked read of the header flags, but that's
unlikely and perfectly acceptable for an opportunistic optimisation
routine, and it can only succeed when it really should due to the
compare-exchange loop.
Note that this fixes ReadRecentBuffer()'s failure to bump the usage
count. While this could be seen as a bug, there currently aren't cases
affected by this in core, so it doesn't seem worth backpatching that portion.
Author: Thomas Munro <[email protected]>
Reported-by: Andres Freund <[email protected]>
Reviewed-by: Andres Freund <[email protected]>
Reviewed-by: Matthias van de Meent <[email protected]>
---
bufmgr: Introduce FlushUnlockedBuffer
There were several copies of code locking a buffer, flushing its contents, and
unlocking the buffer. It seems worth centralizing that into a helper function.
Reviewed-by: Matthias van de Meent <[email protected]>
---
bufmgr: Fix signedness of mask variable in BufferSync()
BM_PERMANENT is defined as 1U<<31, which is a negative number when interpreted
as a signed integer. Unfortunately the mask variable in BufferSync() was
signed. This has been wrong for a long time, but failed to fail, due to
integer conversion rules.
However, in an upcoming patch the width of the state variable will be
increased, with the wrong signedness leading to never flushing permanent
buffers - luckily caught in a test.
It seems better to fix this separately, instead of doing so as part of a
large, otherwise mechanical, patch.
Reviewed-by: Matthias van de Meent <[email protected]>
---
bufmgr: fewer calls to BufferDescriptorGetContentLock
We're planning to merge buffer content locks into BufferDesc.state. To reduce
the size of that patch, centralize calls to BufferDescriptorGetContentLock().
The biggest part of the change is in assertions, by introducing
BufferIsLockedByMe[InMode]() (and removing BufferIsExclusiveLocked()). This
seems like an improvement even without aforementioned plans.
Additionally replace some direct calls to LWLockAcquire() with calls to
LockBuffer().
Reviewed-by: Matthias van de Meent <[email protected]>
---
bufmgr: Don't lock buffer header in StrategyGetBuffer()
Previously StrategyGetBuffer() acquired the buffer header spinlock for every
buffer, whether it was reusable or not. If reusable, it'd be returned, with
the lock held, to GetVictimBuffer(), which then would pin the buffer with
PinBuffer_Locked(). That's somewhat violating the spirit of the guidelines for
holding spinlocks (i.e. that they are only held for a few lines of consecutive
code) and necessitates using PinBuffer_Locked(), which scales worse than
PinBuffer() due to holding the spinlock. This alone makes it worth changing
the code.
However, the main reason to change this is that a future commit will make
PinBuffer_Locked() slower (due to making UnlockBufHdr() slower), to gain
scalability for the much more common case of pinning a pre-existing buffer. By
pinning the buffer with a single atomic operation, iff the buffer is reusable,
we avoid any potential regression for miss-heavy workloads. There strictly are
fewer atomic operations for each potential buffer after this change.
The price for this improvement is that freelist.c needs two CAS loops and
needs to be able to set up the resource accounting for pinned buffers. The
latter is achieved by exposing a new functionfor that purpose from bufmgr.c,
that seems better than exposing the entire private refcount infrastructure.
The improvement seems worth the complexity.
Reviewed-by: Robert Haas <[email protected]>
Reviewed-by: Matthias van de Meent <[email protected]>
来看看AI的解读:
这5个 patch 都聚焦于 PostgreSQL 19 的缓冲区管理器(bufmgr),主要围绕性能提升、代码结构优化、未来扩展性等方面。下面逐条解读:
1. Improve ReadRecentBuffer() scalability
核心内容:
优化 ReadRecentBuffer() 的并发扩展性,减少高并发场景下的锁竞争。 新增 PinBuffer() 的 skip_if_not_valid 模式,使其在 buffer 无效时跳过 pin 操作,减少不必要的 spinlock。 修复了 ReadRecentBuffer() 没有提升 usage count 的问题(有助于 buffer 替换算法的准确性,但目前影响不大)。
目的:
明显提升高并发读场景下的性能,减少锁瓶颈,为未来更复杂的 buffer 管理策略打好基础。
2. Introduce FlushUnlockedBuffer
核心内容:
引入 FlushUnlockedBuffer() 辅助函数,将原本分散的“锁定-刷新-解锁”流程封装成一个统一的 helper。 用于减少重复代码、统一行为,提升代码可维护性。
目的:
提高代码复用性,减少维护成本,并为后续功能扩展或行为变更提供更清晰的入口。
3. Fix signedness of mask variable in BufferSync()
核心内容:
修正 BufferSync() 中 mask 变量的有符号/无符号类型错误。 BM_PERMANENT 用无符号位移,原 mask 为 signed int,可能导致逻辑错误。
目的:
保证未来扩展(如状态位宽增加)时行为正确,防止永久 buffer 没有被正确处理(flush)。 是为后续大补丁做准备的“前置修复”,确保类型安全和代码健壮性。
4. Fewer calls to BufferDescriptorGetContentLock
核心内容:
集中化对 BufferDescriptorGetContentLock() 的调用,减少分散的直接调用。 新增 BufferIsLockedByMeInMode 辅助函数,替换原有 BufferIsExclusiveLocked()。 部分 LWLockAcquire() 调用也被 LockBuffer() 取代。
目的:
为将 buffer content lock 合并到 BufferDesc.state 做铺垫。 让锁管理更集中,便于未来迁移和扩展,减少代码耦合。 提升代码可读性和一致性,为更高效的锁粒度和锁管理做准备。
5. Don't lock buffer header in StrategyGetBuffer()
核心内容:
StrategyGetBuffer() 不再对所有 buffer 都先加锁,只在 reusable buffer 时才进行 pin 操作。 这样避免了持锁时间过长和锁粒度过细的问题。 新增 bufmgr.c 的接口,方便 freelist.c 进行 pinned buffer 的资源统计。
目的:
提升 buffer 策略获取(如缓存淘汰、分配)时的性能,减少锁竞争。 为后续 PinBuffer_Locked() 变慢但更可扩展做铺垫,保证高 miss 场景的性能不回退。 是为未来更复杂的 buffer 状态管理和资源统计机制做准备。
总结与未来展望
这5个 patch 都有明确的性能优化、可维护性提升和代码结构演进目标。
更重要的是,它们为 PostgreSQL 未来在缓冲区管理器领域的更高并发扩展性、更灵活的锁管理、更强的资源统计与控制能力等新功能做了基础性铺垫。
比如合并锁状态、优化 CAS 原子操作、集中化辅助函数和接口,都是为后续内核升级和新特性(如更智能的 buffer pool、更多元的缓存策略、云原生场景)做准备。