PostgreSQL码农集散地

PostgreSQL 18 preview - read_stream 预读逻辑优化, 提升IO性能

PostgreSQL 18 preview - read_stream 预读逻辑优化, 提升IO性能

read_stream.c 文件负责管理读取流的逻辑,它尝试在合适的时候向内核发出读取建议(read-ahead advice),以优化 I/O 性能。具体来说,当使用缓冲 I/O 并且读取的是连续的块时,内核自身的预读机制应该已经足够有效,因此 read_stream.c 会避免在这种情况下发出额外的读取建议。 然而,之前的实现存在一个问题:它在某些情况下过早地放弃了发出读取建议,导致后续的读取操作可能会遇到本可避免的 I/O 停滞(stall)。具体来说,之前的逻辑只会在随机跳转后的第一次读取时发出建议,且仅限于 io_combine_limit 块的范围。如果后续的读取操作仍然在连续的块范围内,但没有发出读取建议,可能会导致性能下降。 这个 patch 主要是对 read_stream.c 文件中的读取流(read stream)逻辑进行了改进,特别是在处理密集流(dense streams)时的读取建议(read-ahead advice)机制。 https://git.postgresql.org/gitweb/?p=postgresql.git;a=commit;h=7ea8cd15661e3b0da4b57be2f25fdd512951576f
Improve read_stream.c advice for dense streams.  
author Thomas Munro <[email protected]>   
Fri, 14 Mar 2025 21:23:59 +0000 (10:23 +1300)  
committer Thomas Munro <[email protected]>   
Sat, 15 Mar 2025 06:04:54 +0000 (19:04 +1300)  
commit 7ea8cd15661e3b0da4b57be2f25fdd512951576f  
tree 35c13962d82bb8fc75b960e6d948e57eead760af tree  
parent 11bd8318602fc2282a6201f714c15461dc2009c6 commit | diff  
Improve read_stream.c advice for dense streams.  

read_stream.c tries not to issue read-ahead advice when it thinks the  
kernel's own read-ahead should be active, ie when using buffered I/O and  
reading sequential blocks.  It previously gave up too easily, and issued  
advice only for the first read of up to io_combine_limit blocks in a  
larger range of sequential blocks after random jump.  The following read  
could suffer an avoidable I/O stall.  
  
Fix, by continuing to issue advice until the corresponding preadv()  
calls catch up with the start of the region we're currently issuing  
advice for, if ever.  That's when the kernel actually sees the  
sequential pattern.  Advice is now disabled only when the stream is  
entirely sequential as far as we can see in the look-ahead window, or  
in other words, when a sequential region is larger than we can cover  
with the current io_concurrency and io_combine_limit settings.  
  
While refactoring the advice control logic, also get rid of the  
"suppress_advice" argument that was passed around between functions to  
skip useless posix_fadvise() calls immediately followed by preadv().  
read_stream_start_pending_read() can figure that out, so let's  
concentrate knowledge of advice heuristics in fewer places (our goal  
being to make advice-based I/O concurrency a legacy mode soon).  

The problem cases were revealed by Tomas Vondra's extensive regression  
testing with many different disk access patterns using Melanie  
Plageman's streaming Bitmap Heap Scan patch, in a battle against the  
venerable always-issue-advice-and-always-one-block-at-a-time code.  

Reviewed-by: Andres Freund <[email protected]> (earlier version)  
Reported-by: Melanie Plageman <[email protected]>  
Reported-by: Tomas Vondra <[email protected]>  
Reported-by: Andres Freund <[email protected]>  
Tested-by: Melanie Plageman <[email protected]>  
Discussion: https://postgr.es/m/CA%2BhUKGK_%3D4CVmMHvsHjOVrK6t4F%3DLBpFzsrr3R%2BaJYN8kcTfWg%40mail.gmail.com  
Discussion: https://postgr.es/m/CA%2BhUKGJ3HSWciQCz8ekP1Zn7N213RfA4nbuotQawfpq23%2Bw-5Q%40mail.gmail.com  

AI 解读

这个 patch 主要是对 read_stream.c 文件中的读取流(read stream)逻辑进行了改进,特别是在处理密集流(dense streams)时的读取建议(read-ahead advice)机制。以下是该 patch 的详细解读:

问题背景

read_stream.c 文件负责管理读取流的逻辑,它尝试在合适的时候向内核发出读取建议(read-ahead advice),以优化 I/O 性能。具体来说,当使用缓冲 I/O 并且读取的是连续的块时,内核自身的预读机制应该已经足够有效,因此 read_stream.c 会避免在这种情况下发出额外的读取建议。 然而,之前的实现存在一个问题:它在某些情况下过早地放弃了发出读取建议,导致后续的读取操作可能会遇到本可避免的 I/O 停滞(stall)。具体来说,之前的逻辑只会在随机跳转后的第一次读取时发出建议,且仅限于 io_combine_limit 块的范围。如果后续的读取操作仍然在连续的块范围内,但没有发出读取建议,可能会导致性能下降。

解决方案

该 patch 通过以下方式解决了上述问题:
  1. 持续发出读取建议 :
    现在, read_stream.c 会持续发出读取建议,直到对应的 preadv() 调用追赶上当前发出建议的区域的起始位置。这意味着内核能够更好地识别出连续的读取模式,从而更有效地进行预读。
  2. 禁用读取建议的条件 :
    只有当读取流在当前的预读窗口(look-ahead window)内完全连续时,才会禁用读取建议。换句话说,如果连续的读取区域超出了当前的 io_concurrency 和 io_combine_limit 设置所能覆盖的范围,读取建议才会被禁用。
  3. 移除冗余参数 :
    在重构读取建议控制逻辑的过程中,patch 移除了 suppress_advice 参数。这个参数原本用于在 posix_fadvise() 调用后立即跳过无用的 preadv() 调用。现在, read_stream_start_pending_read() 函数能够自行判断是否需要发出读取建议,从而将读取建议的启发式逻辑集中在更少的地方。这一改动也是为了逐步将基于读取建议的 I/O 并发模式变为遗留模式(legacy mode)。

测试与反馈

该 patch 的改进是通过 Tomas Vondra 的广泛回归测试发现的,测试中使用了 Melanie Plageman 的流式位图堆扫描(streaming Bitmap Heap Scan)补丁,模拟了多种不同的磁盘访问模式。这些测试揭示了之前实现中的问题,并验证了该 patch 的有效性。

总结

这个 patch 通过改进 read_stream.c 中的读取建议逻辑,解决了在处理密集流时可能出现的 I/O 停滞问题。它通过持续发出读取建议、优化禁用读取建议的条件,以及移除冗余参数,提升了读取流的性能。这一改进经过了广泛的测试和讨论,并得到了相关开发者的认可。