PostgreSQL码农集散地

PostgreSQL 18 preview - 改进buffer manager API

PostgreSQL 18 preview - 改进buffer manager API适配未来read_stream.c 的改进

PostgreSQL 18 改进buffer manager API 以适配未来read_stream.c 的改进.

https://git.postgresql.org/gitweb/?p=postgresql.git;a=commit;h=01261fb07888642efa98ba8d4117654bfd2a413d

Improve buffer manager API for backend pin limits. master  
author Thomas Munro <[email protected]>   
Fri, 14 Mar 2025 02:10:43 +0000 (15:10 +1300)  
committer Thomas Munro <[email protected]>   
Fri, 14 Mar 2025 04:13:09 +0000 (17:13 +1300)  
commit 01261fb07888642efa98ba8d4117654bfd2a413d  
tree a7fb87d483db2ba42f2f2226f3283982fc6d48d8 tree  
parent 7c99dc587a010a0c40d72a0e435111ca7a371c02 commit | diff  
Improve buffer manager API for backend pin limits.  

Previously the support functions assumed that the caller needed one pin  
to make progress, and could optionally use some more, allowing enough  
for every connection to do the same.  Add a couple more functionsfor
callers that want to know:  

* what the maximum possible number could be, irrespective of currently  
  held pins, for space planning purposes  

* how many additional pins they could acquire right now, without the  
  special case allowing one pin, for callers that already hold pins and  
  could already make progress even if no extra pins are available  

The pin limit logic began in commit 31966b15.  This refactoring is  
better suited to read_stream.c, which will be adjusted to respect the  
remaining limit as it changes over time in a follow-up commit.  It also  
computes MaxProportionalPins up front, to avoid performing divisions  
whenever a caller needs to check the balance.  

Reviewed-by: Andres Freund <[email protected]> (earlier versions)  
Discussion: https://postgr.es/m/CA%2BhUKGK_%3D4CVmMHvsHjOVrK6t4F%3DLBpFzsrr3R%2BaJYN8kcTfWg%40mail.gmail.com  

https://git.postgresql.org/gitweb/?p=postgresql.git;a=commit;h=31966b15

bufmgr: Introduce infrastructure for faster relation extension  
author Andres Freund <[email protected]>   
Wed, 5 Apr 2023 23:21:09 +0000 (16:21 -0700)  
committer Andres Freund <[email protected]>   
Wed, 5 Apr 2023 23:21:09 +0000 (16:21 -0700)  
commit 31966b151e6ab7a6284deab6e8fe5faddaf2ae4c  
tree 76deeba4e15702f9596a6d935a8bd185554b8b45 tree  
parent 8eda7314652703a2ae30d6c4a69c378f6813a7f2 commit | diff  
bufmgr: Introduce infrastructure for faster relation extension  

The primary bottlenecks for relation extension are:  

1) The extension lock is held while acquiring a victim buffer for the new  
   page. Acquiring a victim buffer can require writing out the old page  
   contents including possibly needing to flush WAL.  

2) When extending via ReadBuffer() et al, we write a zero page during the  
   extension, and then later write out the actual page contents. This can  
   nearly double the write rate.  

3) The existing bulk relation extension infrastructure in hio.c just amortized  
   the cost of acquiring the relation extension lock, but none of the other  
   costs.  

Unfortunately 1) cannot currently be addressed in a central manner as the  
callers to ReadBuffer() need to acquire the extension lock. To address that,  
this this commit moves the responsibility for acquiring the extension lock  
into bufmgr.c functions. That allows to acquire the relation extension lock  
for just the required time. This will also allow us to improve relation  
extension further, without changing callers.  

The reason we write all-zeroes pages during relation extension is that we hope  
to get ENOSPC errors earlier that way (largely works, except for CoW  
filesystems). It is easier to handle out-of-space errors gracefully if the  
page doesn't yet contain actual tuples. This commit addresses 2), by using the  
recently introduced smgrzeroextend(), which extends the relation, without  
dirtying the kernel page cache for all the extended pages.  

To address 3), this commit introduces a function to extend a relation by  
multiple blocks at a time.  

There are three new exposed functions: ExtendBufferedRel() for extending the  
relation by a single block, ExtendBufferedRelBy() to extend a relation by  
multiple blocks at once, and ExtendBufferedRelTo() for extending a relation up  
to a certain size.  

To avoid duplicating code between ReadBuffer(P_NEW) and the new functions,  
ReadBuffer(P_NEW) now implements relation extension with  
ExtendBufferedRel(), using a flag to tell ExtendBufferedRel() that the  
relation lock is already held.  

Note that this commit does not yet lead to a meaningful performance or  
scalability improvement - for that uses of ReadBuffer(P_NEW) will need to be  
converted to ExtendBuffered*(), which will be done in subsequent commits.  

Reviewed-by: Heikki Linnakangas <[email protected]>  
Reviewed-by: Melanie Plageman <[email protected]>  
Discussion: https://postgr.es/m/[email protected]  

AI 解读patch

2023 年的 31966b15 这个patch是对PostgreSQL的bufmgr(缓冲区管理器)的改进,旨在加速表或索引(relation)的扩展操作。  以下是详细解读:

核心目标:

  • 提高表/索引扩展速度:  当前表/索引扩展的主要瓶颈在于锁的持有时间和不必要的写入操作。

主要瓶颈(解决的问题):

  1. 扩展锁持有时间过长:  在为新页获取victim buffer时,扩展锁被持有。获取victim buffer可能需要写出旧页的内容,包括刷新WAL,这会阻塞其他操作。
  2. 重复写入:  使用ReadBuffer()等函数扩展表/索引时,会先写入一个全零页,然后再写入实际内容。这导致写入量几乎翻倍。
  3. hio.c中的批量扩展机制不完善:hio.c中的现有批量扩展机制只分摊了获取扩展锁的成本,没有解决其他瓶颈。

解决方案:

  1. 将扩展锁管理移至bufmgr.c:  将获取扩展锁的职责转移到bufmgr.c的函数中。这样可以更精确地控制锁的持有时间,只在必要时持有锁。  这为未来的进一步优化奠定了基础,而无需修改调用者。
  2. 使用smgrzeroextend():  使用新引入的smgrzeroextend()函数来扩展表/索引。smgrzeroextend()可以在不污染内核页缓存的情况下扩展表/索引,避免了写入全零页的问题。  这解决了重复写入的问题,并利用了文件系统层面的优化。
  3. 引入批量扩展函数:  引入了一个一次扩展多个块的函数,以减少锁的获取次数和函数调用开销。

新引入的函数:

  • ExtendBufferedRel():扩展表/索引一个块。
  • ExtendBufferedRelBy():扩展表/索引多个块。
  • ExtendBufferedRelTo():扩展表/索引到指定大小。

重构ReadBuffer(P_NEW):

  • ReadBuffer(P_NEW)现在使用ExtendBufferedRel()来实现表/索引扩展,并使用一个标志来指示扩展锁是否已经被持有。  这避免了代码重复。

重要说明:

  • 当前性能提升有限:  这个patch本身并没有带来显著的性能或可扩展性提升。  要实现真正的性能提升,需要将ReadBuffer(P_NEW)的调用转换为使用ExtendBuffered*()函数。  后续的commit会完成这个转换。

总结:

这个patch为PostgreSQL的表/索引扩展机制引入了新的基础设施,通过更精细的锁管理、避免重复写入和引入批量扩展函数,为未来的性能优化奠定了基础。  它并没有直接解决所有性能问题,而是为后续的优化工作铺平了道路。  关键在于将ReadBuffer(P_NEW)的调用替换为新的ExtendBuffered*()函数。


2025 PostgreSQL 18 这个patch改进了缓冲区管理器(buffer manager)的API,以更好地处理后端进程的pin(固定)数量限制。

背景:

  • 之前的pin数量限制逻辑(始于commit 31966b15)假设调用者需要至少一个pin才能继续执行,并且可以选择使用更多pin,允许每个连接都这样做。
  • 这种假设在某些情况下不够灵活,特别是对于已经持有pin并且即使没有额外的pin也能继续执行的调用者。

问题:

  • 之前的API只提供了一种方式来查询可用的pin数量,没有区分调用者是否已经持有pin,以及调用者是否需要至少一个pin才能继续执行。
  • 缺乏对最大可能pin数量的查询,这对于空间规划(space planning)目的很有用。

解决方案:

这个patch引入了两个新的函数,以提供更精细的pin数量查询:

  1. 查询最大可能pin数量:  允许调用者查询最大可能的pin数量,而无需考虑当前持有的pin数量。这对于空间规划非常有用,例如,确定需要分配多少缓冲区。
  2. 查询额外可用pin数量(不包含特殊情况):  允许已经持有pin的调用者查询可以额外获取的pin数量,而不包括允许获取一个pin的特殊情况。这对于已经持有pin并且即使没有额外pin也能继续执行的调用者很有用。

具体改进:

  • 引入新函数:  引入了新的API函数,用于查询最大可能pin数量和额外可用pin数量(不包含特殊情况)。(具体函数名未在patch描述中给出,需要查看实际代码)
  • 预先计算MaxProportionalPins:  预先计算MaxProportionalPins,以避免在每次调用者需要检查余额时都执行除法运算。这提高了性能。
  • 更适合read_stream.c:  这次重构更适合read_stream.c,后续的commit会调整read_stream.c以尊重剩余的pin数量限制,并随着时间推移进行调整。

目的:

  • 更精细的pin数量控制:  允许调用者更精细地控制pin数量,避免过度使用pin,并提高资源利用率。
  • 优化read_stream.c:  为read_stream.c的优化奠定基础,使其能够更好地处理pin数量限制。
  • 提高性能:  通过预先计算MaxProportionalPins,提高性能。

总结:

这个patch通过引入新的API函数和优化现有逻辑,改进了缓冲区管理器的API,以更好地处理后端进程的pin数量限制。  它允许调用者更精细地控制pin数量,并为read_stream.c的优化奠定了基础。  关键在于引入了查询最大可能pin数量和额外可用pin数量(不包含特殊情况)的API函数。