刚刚, PostgreSQL 18发布史诗级特性
刚刚, PostgreSQL 18发布史诗级特性
重磅消息, 刚刚PostgreSQL 18支持了异步IO框架, 这个功能的目的是通过将IO请求提前, 让负责IO的worker去异步完成IO操作, 减少IO等待, 提升整体的系统性能. 并且在AIO中还将支持DIO, 提升IO吞吐并降低CPU消耗.
目前这个功能框架已提交, 通过io_method来配置使用何种IO, 正在开发相应的IO method(并且给开发者预留了扩展方法), IO子模块正陆续接入中. 估计后面会有一大堆相关的提交, 不得不说这是一个史诗级特性.
目前涉及的几个补丁如下:
https://git.postgresql.org/gitweb/?p=postgresql.git;a=commit;h=02844012b304ba80d1c48d51f6fe10bb622490cc
aio: Basic subsystem initialization
author Andres Freund <[email protected]>
Mon, 17 Mar 2025 22:51:33 +0000 (18:51 -0400)
committer Andres Freund <[email protected]>
Mon, 17 Mar 2025 22:51:33 +0000 (18:51 -0400)
commit 02844012b304ba80d1c48d51f6fe10bb622490cc
tree c7753eb6c900a00ebdaa2311b87aefbb21d9f588 tree
parent 65db3963ae7154b8f01e4d73dc6b1ffd81c70e1e commit | diff
aio: Basic subsystem initialization
This commit just does the minimal wiring up of the AIO subsystem, added in the
next commit, to the rest of the system. The next commit contains more details
about motivation and architecture.
This commit is kept separate to make it easier to review, separating the
changes across the tree, from the implementation of the new subsystem.
We discussed squashing this commit with the main commit before merging AIO,
but there has been a mild preference for keeping it separate.
Reviewed-by: Heikki Linnakangas <[email protected]>
Reviewed-by: Noah Misch <[email protected]>
Discussion: https://postgr.es/m/uvrtrknj4kdytuboidbhwclo4gxhswwcpgadptsjvjqcluzmah%40brqs62irg4dt
https://git.postgresql.org/gitweb/?p=postgresql.git;a=commit;h=da7226993fd4b73d8b40abb7167d124eada97f2e
aio: Add core asynchronous I/O infrastructure
author Andres Freund <[email protected]>
Mon, 17 Mar 2025 22:51:33 +0000 (18:51 -0400)
committer Andres Freund <[email protected]>
Mon, 17 Mar 2025 22:51:33 +0000 (18:51 -0400)
commit da7226993fd4b73d8b40abb7167d124eada97f2e
tree 6dfb9949c552c6a6aa6c5511e77a2477ccb9641b tree
parent 02844012b304ba80d1c48d51f6fe10bb622490cc commit | diff
aio: Add core asynchronous I/O infrastructure
The main motivations to use AIO in PostgreSQL are:
a) Reduce the time spent waiting for IO by issuing IO sufficiently early.
In a few places we have approximated this using posix_fadvise() based
prefetching, but that is fairly limited (no completion feedback, double the
syscalls, only works with buffered IO, only works on some OSs).
b) Allow to use Direct-I/O (DIO).
DIO can offload most of the work for IO to hardware and thus increase
throughput / decrease CPU utilization, as well as reduce latency. While we
have gained the ability to configure DIO in d4e71df6, it is not yet usable
for real world workloads, as every IO is executed synchronously.
For portability, the new AIO infrastructure allows to implement AIO using
different methods. The choice of the AIO method is controlled by the new
io_method GUC. As of this commit, the only implemented method is "sync",
i.e. AIO is not actually executed asynchronously. The "sync" method exists to
allow to bypass most of the new code initially.
Subsequent commits will introduce additional IO methods, including a
cross-platform method implemented using worker processes and a linux specific
method using io_uring.
To allow different parts of postgres to use AIO, the core AIO infrastructure
does not need to know what kind of files it is operating on. The necessary
behavioral differences for different files are abstracted as "AIO
Targets". One example target would be smgr. For boring portability reasons,
all targets currently need to be added to an array in aio_target.c. This
commit does not implement any AIO targets, just the infrastructure for
them. The smgr target will be added in a later commit.
Completion (and other events) of IOs for one type of file (i.e. one AIO
target) need to be reacted to differently, based on the IO operation and the
callsite. This is made possible by callbacks that can be registered on
IOs. E.g. an smgr read into a local buffer does not need to update the
corresponding BufferDesc (as there is none), but a read into shared buffers
does. This commit does not contain any callbacks, they will be added in
subsequent commits.
For now the AIO infrastructure only understands READV and WRITEV operations,
but it is expected that more operations will be added. E.g. fsync/fdatasync,
flush_range and network operations like send/recv.
As of this commit, nothing uses the AIO infrastructure. Later commits will add
an smgr target, md.c and bufmgr.c callbacks and then finally use AIO for
read_stream.c IO, which, in one fell swoop, will convert all read stream users
to AIO.
The goal is to use AIO in many more places. There are patches to use AIO for
checkpointer and bgwriter that are reasonably close to being ready. There also
are prototypes to use it for WAL, relation extension, backend writes and many
more. Those prototypes were important to ensure the design of the AIO
subsystem is not too limiting (e.g. WAL writes need to happen in critical
sections, which influenced a lot of the design).
A future commit will add an AIO README explaining the AIO architecture and how
to use the AIO subsystem. The README is added later, as it references details
only added in later commits.
Many many more people than the folks named below have contributed with
feedback, work on semi-independent patches etc. E.g. various folks have
contributed patches to use the read stream infrastructure (added by Thomas in
b5a9b18cd0b) in more places. Similarly, a *lot* of folks have contributed to
the CI infrastructure, which I had started to work on to make adding AIO
feasible.
Some of the work by contributors has gone into the "v1" prototype of AIO,
which heavily influenced the current design of the AIO subsystem. None of the
code from that directly survives, but without the prototype, the current
version of the AIO infrastructure would not exist.
Similarly, the reviewers below have not necessarily looked at the current
design or the whole infrastructure, but have provided very valuable input. I
am to blame for problems, not they.
Author: Andres Freund <[email protected]>
Co-authored-by: Thomas Munro <[email protected]>
Co-authored-by: Nazir Bilal Yavuz <[email protected]>
Co-authored-by: Melanie Plageman <[email protected]>
Reviewed-by: Heikki Linnakangas <[email protected]>
Reviewed-by: Noah Misch <[email protected]>
Reviewed-by: Jakub Wartak <[email protected]>
Reviewed-by: Melanie Plageman <[email protected]>
Reviewed-by: Robert Haas <[email protected]>
Reviewed-by: Dmitry Dolgov <[email protected]>
Reviewed-by: Antonin Houska <[email protected]>
Discussion: https://postgr.es/m/uvrtrknj4kdytuboidbhwclo4gxhswwcpgadptsjvjqcluzmah%40brqs62irg4dt
Discussion: https://postgr.es/m/[email protected]
Discussion: https://postgr.es/m/stj36ea6yyhoxtqkhpieia2z4krnam7qyetc57rfezgk4zgapf@gcnactj4z56m
https://git.postgresql.org/gitweb/?p=postgresql.git;a=commit;h=55b454d0e14084c841a034073abbf1a0ea937a45
aio: Infrastructure for io_method=worker
author Andres Freund <[email protected]>
Tue, 18 Mar 2025 14:52:33 +0000 (10:52 -0400)
committer Andres Freund <[email protected]>
Tue, 18 Mar 2025 15:54:01 +0000 (11:54 -0400)
commit 55b454d0e14084c841a034073abbf1a0ea937a45
tree 4bdd85a6acb02123b35bfeb1c33c1c071be30ba1 tree
parent 549ea06e4217aca10d3a73dc09cf5018c51bc23a commit | diff
aio: Infrastructure for io_method=worker
This commit contains the basic, system-wide, infrastructure for
io_method=worker. It does not yet actually execute IO, this commit just
provides the infrastructure for running IO workers, kept separate for easier
review.
The number of IO workers can be adjusted with a PGC_SIGHUP GUC. Eventually
we'd like to make the number of workers dynamically scale up/down based on the
current "IO load".
To allow the number of IO workers to be increased without a restart, we need
to reserve PGPROC entries for the workers unconditionally. This has been
judged to be worth the cost. If it turns out to be problematic, we can
introduce a PGC_POSTMASTER GUC to control the maximum number.
As io workers might be needed during shutdown, e.g. for AIO during the
shutdown checkpoint, a new PMState phase is added. IO workers are shut down
after the shutdown checkpoint has been performed and walsender/archiver have
shut down, but before the checkpointer itself shuts down. See also
87a6690cc69.
Updates PGSTAT_FILE_FORMAT_ID due to the addition of a new BackendType.
Reviewed-by: Noah Misch <[email protected]>
Co-authored-by: Thomas Munro <[email protected]>
Co-authored-by: Andres Freund <[email protected]>
Discussion: https://postgr.es/m/uvrtrknj4kdytuboidbhwclo4gxhswwcpgadptsjvjqcluzmah%40brqs62irg4dt
Discussion: https://postgr.es/m/[email protected]
Discussion: https://postgr.es/m/stj36ea6yyhoxtqkhpieia2z4krnam7qyetc57rfezgk4zgapf@gcnactj4z56m
https://git.postgresql.org/gitweb/?p=postgresql.git;a=commit;h=247ce06b883d7b3a40d08312dc03dfb37fbff212
aio: Add io_method=worker
author Andres Freund <[email protected]>
Tue, 18 Mar 2025 14:52:33 +0000 (10:52 -0400)
committer Andres Freund <[email protected]>
Tue, 18 Mar 2025 15:54:01 +0000 (11:54 -0400)
commit 247ce06b883d7b3a40d08312dc03dfb37fbff212
tree 6e45be6994fa574c077c32ab316eb9579d67519e tree
parent 55b454d0e14084c841a034073abbf1a0ea937a45 commit | diff
aio: Add io_method=worker
The previous commit introduced the infrastructure to start io_workers. This
commit actually makes the workers execute IOs.
IO workers consume IOs from a shared memory submission queue, run traditional
synchronous system calls, and perform the shared completion handling
immediately. Client code submits most requests by pushing IOs into the
submission queue, and waits (if necessary) using condition variables. Some
IOs cannot be performed in another process due to lack of infrastructure for
reopening the file, and must processed synchronously by the client code when
submitted.
For now the default io_method is changed to "worker". We should re-evaluate
that around beta1, we might want to be careful and set the default to "sync"
for 18.
Reviewed-by: Noah Misch <[email protected]>
Co-authored-by: Thomas Munro <[email protected]>
Co-authored-by: Andres Freund <[email protected]>
Discussion: https://postgr.es/m/uvrtrknj4kdytuboidbhwclo4gxhswwcpgadptsjvjqcluzmah%40brqs62irg4dt
Discussion: https://postgr.es/m/[email protected]
Discussion: https://postgr.es/m/stj36ea6yyhoxtqkhpieia2z4krnam7qyetc57rfezgk4zgapf@gcnactj4z56m
AI解读
这个补丁为 PostgreSQL 添加了核心的异步 I/O (AIO) 基础设施。
主要目的:
减少 I/O 等待时间: 通过提前发出 I/O 请求,减少程序等待 I/O 完成的时间。 之前使用 posix_fadvise()进行预取,但这种方法有局限性(没有完成反馈、双倍的系统调用、仅适用于buffer I/O、仅适用于某些操作系统)。允许使用直接 I/O (DIO): DIO 可以将大部分 I/O 工作offload到硬件,从而提高吞吐量、降低 CPU 利用率并减少延迟。 虽然之前已经可以配置 DIO,但由于每个 I/O 都是同步执行的,因此还不能在实际工作负载中使用。
实现方式:
可移植性: 为了保证可移植性,新的 AIO 基础设施允许使用不同的方法来实现 AIO。 AIO 方法的选择由新的 GUC 参数 io_method控制。"sync" 方法: 目前唯一实现的方法是 "sync",即 AIO 实际上不是异步执行的。 "sync" 方法的存在是为了允许最初绕过大部分新代码。 后续方法: 后续的提交将引入其他 I/O 方法,包括使用工作进程实现的跨平台方法和使用 io_uring的 Linux 特定方法。AIO Targets: 为了允许 PostgreSQL 的不同部分使用 AIO,核心 AIO 基础设施不需要知道它正在操作的文件类型。 不同文件的必要行为差异被抽象为 "AIO Targets"。 例如, smgr就是一个目标。 所有目标目前都需要添加到aio_target.c中的一个数组中。回调函数: 基于 I/O 操作和调用点,需要以不同的方式响应一种文件类型(即一个 AIO 目标)的 I/O 完成(和其他事件)。 这是通过可以在 I/O 上注册的回调函数实现的。 例如,读取到本地缓冲区的 smgr不需要更新相应的BufferDesc(因为没有),但读取到共享缓冲区的则需要。支持的操作: 目前,AIO 基础设施只理解 READV和WRITEV操作,但预计会添加更多操作。 例如,fsync/fdatasync、flush_range和网络操作(如send/recv)。
当前状态:
未使用: 目前没有任何地方使用 AIO 基础设施。 后续计划: 后续的提交将添加 smgr目标、md.c和bufmgr.c回调,然后最终将 AIO 用于read_stream.cI/O,这将使所有读取流用户都转换为 AIO。未来目标: 目标是在更多地方使用 AIO。 已经有一些补丁接近完成,用于 checkpointer 和 bgwriter。 还有一些原型用于 WAL、关系扩展、后端写入等等。 这些原型对于确保 AIO 子系统的设计没有太大的限制非常重要(例如,WAL 写入需要在关键部分进行,这影响了许多设计)。 文档: 未来将添加一个 AIO README,解释 AIO 架构以及如何使用 AIO 子系统。 README 将稍后添加,因为它引用了仅在后续提交中添加的细节。
贡献者:
许多人通过反馈、独立补丁等方式做出了贡献。 例如,许多人贡献了补丁,以在更多地方使用读取流基础设施。 同样,许多人为 CI 基础设施做出了贡献,这使得添加 AIO 成为可能。
总结:
这个补丁是 PostgreSQL 中异步 I/O 的基础。它定义了 AIO 的架构,并提供了一个框架,允许在 PostgreSQL 的不同部分使用 AIO。虽然这个补丁本身并没有启用 AIO,但它为后续的提交奠定了基础,这些提交将添加 AIO 目标、回调函数,并最终启用 AIO 用于各种 I/O 操作。 最终目标是提高 PostgreSQL 的 I/O 性能,减少延迟,并降低 CPU 利用率。
这个补丁为 io_method=worker 实现了基础的、系统范围的基础设施。 重要的是,它还没有真正执行 I/O 操作。 这个补丁只是提供了运行 I/O worker 的基础设施,为了方便审查,将其单独提交。
主要内容:
I/O Worker 数量配置: I/O worker 的数量可以通过一个 PGC_SIGHUPGUC 参数进行调整。 最终目标是根据当前的 "I/O 负载" 动态地增加/减少 worker 的数量。预留 PGPROC 条目: 为了允许在不重启的情况下增加 I/O worker 的数量,需要无条件地为 worker 预留 PGPROC条目。 开发者认为这样做是值得的。 如果出现问题,可以引入一个PGC_POSTMASTERGUC 参数来控制最大数量。新的 PMState 阶段: 由于在关闭期间可能需要 I/O worker,例如在关闭检查点期间进行 AIO,因此添加了一个新的 PMState阶段。 I/O worker 在执行完关闭检查点、walsender/archiver关闭后,但在检查点进程本身关闭之前关闭。 这与 87a6690cc69 相关。更新 PGSTAT_FILE_FORMAT_ID: 由于添加了一个新的 BackendType,因此更新了PGSTAT_FILE_FORMAT_ID。
总结:
这个补丁是实现 io_method=worker 的关键一步。 它设置了运行 I/O worker 的框架,包括配置 worker 数量、预留资源以及在 PostgreSQL 关闭期间正确处理 worker。 虽然它本身不执行任何 I/O 操作,但它为后续的提交奠定了基础,这些提交将实际使用这些 worker 来执行异步 I/O。 该补丁还考虑了在不重启服务器的情况下动态调整 worker 数量的需求,以及在服务器关闭期间对 I/O worker 的依赖。
这个补丁实现了 io_method=worker,让 worker 真正开始执行 I/O 操作。
关键点:
I/O Worker 工作模式: I/O worker 从共享内存提交队列中获取 I/O 请求,执行传统的同步系统调用,并立即执行共享的完成处理。 提交队列和条件变量: 客户端代码通过将 I/O 请求推送到提交队列来提交大多数请求,并使用条件变量等待(如果需要)。 同步处理: 由于缺乏重新打开文件的基础设施,某些 I/O 请求无法在另一个进程中执行,因此必须由客户端代码在提交时同步处理。 默认 io_method: 目前,默认的 io_method已更改为 "worker"。 开发者计划在 beta1 版本前后重新评估这一点,并可能为了谨慎起见,在 PostgreSQL 18 中将默认值设置为 "sync"。
总结:
这个补丁是 io_method=worker 功能的实现,它利用之前补丁中建立的基础设施,让 I/O worker 进程开始实际执行 I/O 操作。 I/O 请求通过共享内存队列传递给 worker,worker 执行同步系统调用,然后处理完成事件。 该补丁还处理了某些 I/O 请求必须同步执行的情况。 最后,该补丁将默认的 io_method 设置为 "worker",但开发者计划在发布前重新评估这个决定。 总的来说,这个补丁标志着 PostgreSQL 异步 I/O 功能向前迈出了重要一步。