PostgreSQL码农集散地

刚刚, PostgreSQL 18发布史诗级特性

刚刚, PostgreSQL 18发布史诗级特性

重磅消息, 刚刚PostgreSQL 18支持了异步IO框架, 这个功能的目的是通过将IO请求提前, 让负责IO的worker去异步完成IO操作, 减少IO等待, 提升整体的系统性能. 并且在AIO中还将支持DIO, 提升IO吞吐并降低CPU消耗.

目前这个功能框架已提交, 通过io_method来配置使用何种IO, 正在开发相应的IO method(并且给开发者预留了扩展方法), IO子模块正陆续接入中. 估计后面会有一大堆相关的提交, 不得不说这是一个史诗级特性.

目前涉及的几个补丁如下:

https://git.postgresql.org/gitweb/?p=postgresql.git;a=commit;h=02844012b304ba80d1c48d51f6fe10bb622490cc

aio: Basic subsystem initialization  
author Andres Freund <[email protected]>   
Mon, 17 Mar 2025 22:51:33 +0000 (18:51 -0400)  
committer Andres Freund <[email protected]>   
Mon, 17 Mar 2025 22:51:33 +0000 (18:51 -0400)  
commit 02844012b304ba80d1c48d51f6fe10bb622490cc  
tree c7753eb6c900a00ebdaa2311b87aefbb21d9f588 tree  
parent 65db3963ae7154b8f01e4d73dc6b1ffd81c70e1e commit | diff  
aio: Basic subsystem initialization  

This commit just does the minimal wiring up of the AIO subsystem, added in the  
next commit, to the rest of the system. The next commit contains more details  
about motivation and architecture.  

This commit is kept separate to make it easier to review, separating the  
changes across the tree, from the implementation of the new subsystem.  

We discussed squashing this commit with the main commit before merging AIO,  
but there has been a mild preference for keeping it separate.  

Reviewed-by: Heikki Linnakangas <[email protected]>  
Reviewed-by: Noah Misch <[email protected]>  
Discussion: https://postgr.es/m/uvrtrknj4kdytuboidbhwclo4gxhswwcpgadptsjvjqcluzmah%40brqs62irg4dt  

https://git.postgresql.org/gitweb/?p=postgresql.git;a=commit;h=da7226993fd4b73d8b40abb7167d124eada97f2e

aio: Add core asynchronous I/O infrastructure  
author Andres Freund <[email protected]>   
Mon, 17 Mar 2025 22:51:33 +0000 (18:51 -0400)  
committer Andres Freund <[email protected]>   
Mon, 17 Mar 2025 22:51:33 +0000 (18:51 -0400)  
commit da7226993fd4b73d8b40abb7167d124eada97f2e  
tree 6dfb9949c552c6a6aa6c5511e77a2477ccb9641b tree  
parent 02844012b304ba80d1c48d51f6fe10bb622490cc commit | diff  
aio: Add core asynchronous I/O infrastructure  

The main motivations to use AIO in PostgreSQL are:  

a) Reduce the time spent waiting for IO by issuing IO sufficiently early.  

   In a few places we have approximated this using posix_fadvise() based  
   prefetching, but that is fairly limited (no completion feedback, double the  
   syscalls, only works with buffered IO, only works on some OSs).  

b) Allow to use Direct-I/O (DIO).  

   DIO can offload most of the work for IO to hardware and thus increase  
   throughput / decrease CPU utilization, as well as reduce latency.  While we  
   have gained the ability to configure DIO in d4e71df6, it is not yet usable  
for real world workloads, as every IO is executed synchronously.  

For portability, the new AIO infrastructure allows to implement AIO using  
different methods. The choice of the AIO method is controlled by the new  
io_method GUC. As of this commit, the only implemented method is "sync",  
i.e. AIO is not actually executed asynchronously. The "sync" method exists to  
allow to bypass most of the new code initially.  

Subsequent commits will introduce additional IO methods, including a  
cross-platform method implemented using worker processes and a linux specific  
method using io_uring.  

To allow different parts of postgres to use AIO, the core AIO infrastructure  
does not need to know what kind of files it is operating on. The necessary  
behavioral differences for different files are abstracted as "AIO  
Targets"
. One example target would be smgr. For boring portability reasons,  
all targets currently need to be added to an array in aio_target.c.  This  
commit does not implement any AIO targets, just the infrastructure for
them. The smgr target will be added in a later commit.  

Completion (and other events) of IOs for one type of file (i.e. one AIO  
target) need to be reacted to differently, based on the IO operation and the  
callsite. This is made possible by callbacks that can be registered on  
IOs. E.g. an smgr read into a local buffer does not need to update the  
corresponding BufferDesc (as there is none), but a read into shared buffers  
does.  This commit does not contain any callbacks, they will be added in
subsequent commits.  

For now the AIO infrastructure only understands READV and WRITEV operations,  
but it is expected that more operations will be added. E.g. fsync/fdatasync,  
flush_range and network operations like send/recv.  

As of this commit, nothing uses the AIO infrastructure. Later commits will add  
an smgr target, md.c and bufmgr.c callbacks and then finally use AIO for
read_stream.c IO, which, in one fell swoop, will convert all read stream users  
to AIO.  

The goal is to use AIO in many more places. There are patches to use AIO for
checkpointer and bgwriter that are reasonably close to being ready. There also  
are prototypes to use it for WAL, relation extension, backend writes and many  
more. Those prototypes were important to ensure the design of the AIO  
subsystem is not too limiting (e.g. WAL writes need to happen in critical  
sections, which influenced a lot of the design).  

A future commit will add an AIO README explaining the AIO architecture and how  
to use the AIO subsystem. The README is added later, as it references details  
only added in later commits.  

Many many more people than the folks named below have contributed with  
feedback, work on semi-independent patches etc. E.g. various folks have  
contributed patches to use the read stream infrastructure (added by Thomas in
b5a9b18cd0b) in more places. Similarly, a *lot* of folks have contributed to  
the CI infrastructure, which I had started to work on to make adding AIO  
feasible.  

Some of the work by contributors has gone into the "v1" prototype of AIO,  
which heavily influenced the current design of the AIO subsystem. None of the  
code from that directly survives, but without the prototype, the current  
version of the AIO infrastructure would not exist.  

Similarly, the reviewers below have not necessarily looked at the current  
design or the whole infrastructure, but have provided very valuable input. I  
am to blame for problems, not they.  

Author: Andres Freund <[email protected]>  
Co-authored-by: Thomas Munro <[email protected]>  
Co-authored-by: Nazir Bilal Yavuz <[email protected]>  
Co-authored-by: Melanie Plageman <[email protected]>  
Reviewed-by: Heikki Linnakangas <[email protected]>  
Reviewed-by: Noah Misch <[email protected]>  
Reviewed-by: Jakub Wartak <[email protected]>  
Reviewed-by: Melanie Plageman <[email protected]>  
Reviewed-by: Robert Haas <[email protected]>  
Reviewed-by: Dmitry Dolgov <[email protected]>  
Reviewed-by: Antonin Houska <[email protected]>  
Discussion: https://postgr.es/m/uvrtrknj4kdytuboidbhwclo4gxhswwcpgadptsjvjqcluzmah%40brqs62irg4dt  
Discussion: https://postgr.es/m/[email protected]  
Discussion: https://postgr.es/m/stj36ea6yyhoxtqkhpieia2z4krnam7qyetc57rfezgk4zgapf@gcnactj4z56m  

https://git.postgresql.org/gitweb/?p=postgresql.git;a=commit;h=55b454d0e14084c841a034073abbf1a0ea937a45

aio: Infrastructure for io_method=worker  
author Andres Freund <[email protected]>   
Tue, 18 Mar 2025 14:52:33 +0000 (10:52 -0400)  
committer Andres Freund <[email protected]>   
Tue, 18 Mar 2025 15:54:01 +0000 (11:54 -0400)  
commit 55b454d0e14084c841a034073abbf1a0ea937a45  
tree 4bdd85a6acb02123b35bfeb1c33c1c071be30ba1 tree  
parent 549ea06e4217aca10d3a73dc09cf5018c51bc23a commit | diff  
aio: Infrastructure for io_method=worker  

This commit contains the basic, system-wide, infrastructure for
io_method=worker. It does not yet actually execute IO, this commit just  
provides the infrastructure for running IO workers, kept separate for easier  
review.  

The number of IO workers can be adjusted with a PGC_SIGHUP GUC. Eventually  
we'd like to make the number of workers dynamically scale up/down based on the  
current "IO load".  

To allow the number of IO workers to be increased without a restart, we need  
to reserve PGPROC entries for the workers unconditionally. This has been  
judged to be worth the cost. If it turns out to be problematic, we can  
introduce a PGC_POSTMASTER GUC to control the maximum number.  

As io workers might be needed during shutdown, e.g. for AIO during the  
shutdown checkpoint, a new PMState phase is added. IO workers are shut down  
after the shutdown checkpoint has been performed and walsender/archiver have  
shut down, but before the checkpointer itself shuts down. See also  
87a6690cc69.  

Updates PGSTAT_FILE_FORMAT_ID due to the addition of a new BackendType.  

Reviewed-by: Noah Misch <[email protected]>  
Co-authored-by: Thomas Munro <[email protected]>  
Co-authored-by: Andres Freund <[email protected]>  
Discussion: https://postgr.es/m/uvrtrknj4kdytuboidbhwclo4gxhswwcpgadptsjvjqcluzmah%40brqs62irg4dt  
Discussion: https://postgr.es/m/[email protected]  
Discussion: https://postgr.es/m/stj36ea6yyhoxtqkhpieia2z4krnam7qyetc57rfezgk4zgapf@gcnactj4z56m  

https://git.postgresql.org/gitweb/?p=postgresql.git;a=commit;h=247ce06b883d7b3a40d08312dc03dfb37fbff212

aio: Add io_method=worker  
author Andres Freund <[email protected]>   
Tue, 18 Mar 2025 14:52:33 +0000 (10:52 -0400)  
committer Andres Freund <[email protected]>   
Tue, 18 Mar 2025 15:54:01 +0000 (11:54 -0400)  
commit 247ce06b883d7b3a40d08312dc03dfb37fbff212  
tree 6e45be6994fa574c077c32ab316eb9579d67519e tree  
parent 55b454d0e14084c841a034073abbf1a0ea937a45 commit | diff  
aio: Add io_method=worker  

The previous commit introduced the infrastructure to start io_workers. This  
commit actually makes the workers execute IOs.  

IO workers consume IOs from a shared memory submission queue, run traditional  
synchronous system calls, and perform the shared completion handling  
immediately.  Client code submits most requests by pushing IOs into the  
submission queue, and waits (if necessary) using condition variables.  Some  
IOs cannot be performed in another process due to lack of infrastructure for
reopening the file, and must processed synchronously by the client code when  
submitted.  

For now the default io_method is changed to "worker". We should re-evaluate  
that around beta1, we might want to be careful and set the default to "sync"
for 18.  

Reviewed-by: Noah Misch <[email protected]>  
Co-authored-by: Thomas Munro <[email protected]>  
Co-authored-by: Andres Freund <[email protected]>  
Discussion: https://postgr.es/m/uvrtrknj4kdytuboidbhwclo4gxhswwcpgadptsjvjqcluzmah%40brqs62irg4dt  
Discussion: https://postgr.es/m/[email protected]  
Discussion: https://postgr.es/m/stj36ea6yyhoxtqkhpieia2z4krnam7qyetc57rfezgk4zgapf@gcnactj4z56m  

AI解读

这个补丁为 PostgreSQL 添加了核心的异步 I/O (AIO) 基础设施。

主要目的:

  • 减少 I/O 等待时间: 通过提前发出 I/O 请求,减少程序等待 I/O 完成的时间。 之前使用 posix_fadvise() 进行预取,但这种方法有局限性(没有完成反馈、双倍的系统调用、仅适用于buffer I/O、仅适用于某些操作系统)。
  • 允许使用直接 I/O (DIO): DIO 可以将大部分 I/O 工作offload到硬件,从而提高吞吐量、降低 CPU 利用率并减少延迟。  虽然之前已经可以配置 DIO,但由于每个 I/O 都是同步执行的,因此还不能在实际工作负载中使用。

实现方式:

  • 可移植性: 为了保证可移植性,新的 AIO 基础设施允许使用不同的方法来实现 AIO。  AIO 方法的选择由新的 GUC 参数 io_method 控制。
  • "sync" 方法:  目前唯一实现的方法是 "sync",即 AIO 实际上不是异步执行的。  "sync" 方法的存在是为了允许最初绕过大部分新代码。
  • 后续方法: 后续的提交将引入其他 I/O 方法,包括使用工作进程实现的跨平台方法和使用 io_uring 的 Linux 特定方法。
  • AIO Targets: 为了允许 PostgreSQL 的不同部分使用 AIO,核心 AIO 基础设施不需要知道它正在操作的文件类型。  不同文件的必要行为差异被抽象为 "AIO Targets"。  例如,smgr 就是一个目标。  所有目标目前都需要添加到 aio_target.c 中的一个数组中。
  • 回调函数: 基于 I/O 操作和调用点,需要以不同的方式响应一种文件类型(即一个 AIO 目标)的 I/O 完成(和其他事件)。  这是通过可以在 I/O 上注册的回调函数实现的。  例如,读取到本地缓冲区的 smgr 不需要更新相应的 BufferDesc(因为没有),但读取到共享缓冲区的则需要。
  • 支持的操作: 目前,AIO 基础设施只理解 READV 和 WRITEV 操作,但预计会添加更多操作。  例如,fsync/fdatasync、flush_range 和网络操作(如 send/recv)。

当前状态:

  • 未使用:  目前没有任何地方使用 AIO 基础设施。
  • 后续计划: 后续的提交将添加 smgr 目标、md.c 和 bufmgr.c 回调,然后最终将 AIO 用于 read_stream.c I/O,这将使所有读取流用户都转换为 AIO。
  • 未来目标: 目标是在更多地方使用 AIO。  已经有一些补丁接近完成,用于 checkpointer 和 bgwriter。  还有一些原型用于 WAL、关系扩展、后端写入等等。  这些原型对于确保 AIO 子系统的设计没有太大的限制非常重要(例如,WAL 写入需要在关键部分进行,这影响了许多设计)。
  • 文档: 未来将添加一个 AIO README,解释 AIO 架构以及如何使用 AIO 子系统。  README 将稍后添加,因为它引用了仅在后续提交中添加的细节。

贡献者:

许多人通过反馈、独立补丁等方式做出了贡献。  例如,许多人贡献了补丁,以在更多地方使用读取流基础设施。  同样,许多人为 CI 基础设施做出了贡献,这使得添加 AIO 成为可能。

总结:

这个补丁是 PostgreSQL 中异步 I/O 的基础。它定义了 AIO 的架构,并提供了一个框架,允许在 PostgreSQL 的不同部分使用 AIO。虽然这个补丁本身并没有启用 AIO,但它为后续的提交奠定了基础,这些提交将添加 AIO 目标、回调函数,并最终启用 AIO 用于各种 I/O 操作。  最终目标是提高 PostgreSQL 的 I/O 性能,减少延迟,并降低 CPU 利用率。


这个补丁为 io_method=worker 实现了基础的、系统范围的基础设施。 重要的是,它还没有真正执行 I/O 操作。 这个补丁只是提供了运行 I/O worker 的基础设施,为了方便审查,将其单独提交。

主要内容:

  • I/O Worker 数量配置:  I/O worker 的数量可以通过一个 PGC_SIGHUP GUC 参数进行调整。 最终目标是根据当前的 "I/O 负载" 动态地增加/减少 worker 的数量。
  • 预留 PGPROC 条目: 为了允许在不重启的情况下增加 I/O worker 的数量,需要无条件地为 worker 预留 PGPROC 条目。 开发者认为这样做是值得的。 如果出现问题,可以引入一个 PGC_POSTMASTER GUC 参数来控制最大数量。
  • 新的 PMState 阶段:  由于在关闭期间可能需要 I/O worker,例如在关闭检查点期间进行 AIO,因此添加了一个新的 PMState 阶段。 I/O worker 在执行完关闭检查点、walsender/archiver 关闭后,但在检查点进程本身关闭之前关闭。  这与 87a6690cc69 相关。
  • 更新 PGSTAT_FILE_FORMAT_ID: 由于添加了一个新的 BackendType,因此更新了 PGSTAT_FILE_FORMAT_ID。

总结:

这个补丁是实现 io_method=worker 的关键一步。 它设置了运行 I/O worker 的框架,包括配置 worker 数量、预留资源以及在 PostgreSQL 关闭期间正确处理 worker。  虽然它本身不执行任何 I/O 操作,但它为后续的提交奠定了基础,这些提交将实际使用这些 worker 来执行异步 I/O。  该补丁还考虑了在不重启服务器的情况下动态调整 worker 数量的需求,以及在服务器关闭期间对 I/O worker 的依赖。


这个补丁实现了 io_method=worker,让 worker 真正开始执行 I/O 操作。

关键点:

  • I/O Worker 工作模式: I/O worker 从共享内存提交队列中获取 I/O 请求,执行传统的同步系统调用,并立即执行共享的完成处理。
  • 提交队列和条件变量: 客户端代码通过将 I/O 请求推送到提交队列来提交大多数请求,并使用条件变量等待(如果需要)。
  • 同步处理: 由于缺乏重新打开文件的基础设施,某些 I/O 请求无法在另一个进程中执行,因此必须由客户端代码在提交时同步处理。
  • 默认 io_method:  目前,默认的 io_method 已更改为 "worker"。 开发者计划在 beta1 版本前后重新评估这一点,并可能为了谨慎起见,在 PostgreSQL 18 中将默认值设置为 "sync"。

总结:

这个补丁是 io_method=worker 功能的实现,它利用之前补丁中建立的基础设施,让 I/O worker 进程开始实际执行 I/O 操作。  I/O 请求通过共享内存队列传递给 worker,worker 执行同步系统调用,然后处理完成事件。  该补丁还处理了某些 I/O 请求必须同步执行的情况。  最后,该补丁将默认的 io_method 设置为 "worker",但开发者计划在发布前重新评估这个决定。  总的来说,这个补丁标志着 PostgreSQL 异步 I/O 功能向前迈出了重要一步。