因收缩LVM分区而遇到的一个BUG
这是我一个好友在国产操作系统中遇到的问题,在我这里发布一下,可以点击阅读原文查看。
环境
openEuler 22.03 LTS SP4
kernel 5.10.0-216.0.0.115.0e2203sp4.aarch64
分区:xfs
问题
在测试环境当中需要调整磁盘分区的大小,将原来的3磁盘分区进行适当的收缩,为新增的第4个分区腾出1T的存储空间。原来的/dev/vda3上使用了LVM对挂载点进行了管理。需要先对LVM进行收缩直至pvresize释放收缩后的磁盘空间。然后用parted对分区进行调整。调整完成后,lsblk发现因为二进制算法的原因,/dev/vda3上大概有接近400G左右的空间没有被LVM所管理。就想把这个部分磁盘空间重新分配给/dev/vda3,并让LVM重新使用起来。前面的步骤很顺利,但是在做xfs_growfs时遇到了一个报错
# xfs_growfs /dev/openeuler/homemeta-data=/dev/loop0 isize=256 agcount=52,agsize=76288719 blks= sectsz=512 attr=2data = bsize=4096 blocks=3905982455, imaxpct=5= sunit=0 swidth=0 blksnaming =version 2 bsize=4096 ascii-ci=0log =internal bsize=4096 blocks=32768, version=2= sectsz=512 sunit=0 blks, lazy-count=0realtime =none extsz=4096 blocks=0, rtextents=0xfs_growfs: XFS_IOC_FSGROWFSDATA xfsctl failed: Invalid argument |
解决方案
根据报错信息分析,应该是xfs_growfs没有精准的识别分区中的剩余空间大小出来,所以就专门找了一下这个报错的解决办法。碰巧在红帽上找到了这个报错的相关说明,发生这个问题的根本原因(及翻译)如下:
In the case where growing a filesystem would leave the last AG too small, the fixup code has an overflow in the calculation of the new size with one fewer ag, because "nagcount" is a 32 bit number. If the new filesystem has > 2^32 blocks in it this causes a problem resulting in an EINVAL return from growfs.如果文件系统增大会导致最后一个 AG 太小,则修复代码会在计算新大小时溢出,因为“nagcount”是一个 32 位数。如果新文件系统中有 > 2^32 个块,则会导致出现问题,从而导致 growfs 返回 EINVAL。 |
红帽官方是修复了这个bug,而且最低修复版本是在kernel 2.6.18上,这个内核版本号对应的操作系统是十几年前的RedHat 5和6。而解决方案发布时间是在2024年8月7日。这个bug的具体修复时间因为红帽的机制,现在不能在bugzilla上面查找到最初的bug上报时间了。因为在当前openEuler 22.03 LTS上又遇到这个问题,初步判断openEuler社区是没有对这个bug进行修复(如果知道有修复的朋友,请告知)
所以在这里记录一下红帽的手动操作方案
手动调节方案
手动调整设备大小,这样就不需要四舍五入了。
1. 确定文件系统的agsize
# xfs_info /dev/VGDATA/data1meta-data=/dev/VGDATA/data1 isize=256 agcount=18, agsize=268435455 blks <--- agsize = 268435455= sectsz=512 attr=2data = bsize=4096 blocks=4831838190, imaxpct=5 <--- Filesystem block size = 4096= sunit=0 swidth=0 blksnaming =version 2 bsize=4096 ascii-ci=0log =internal bsize=4096 blocks=32768, version=2= sectsz=512 sunit=0 blks, lazy-count=1realtime =none extsz=4096 blocks=0, rtextents=0 |
2. 确定我们尝试增大的设备的大小(您需要知道文件系统位于哪个设备上,如果在 LVM 上,请使用 dmsetup info -c 确定次要编号):
# cat /proc/partitions | grep dm-2major minor #blocks name253 2 27917287424 dm-2 <--- Number of blocks = 27917287424 |
3. 根据以下公式判断是否存在溢出:
(Number of Total Disk Blocks * 1024) modulo (agsize * Filesystem block size) = (27917287424 * 1024) modulo (268435455 * 4096) = (28587302322176) modulo (1099511623680) = 106496 bytes overflow |
4. 确定应使用的块设备大小以确保没有溢出:
(overflow / 1024) = Number of Overflow blocks(106496 / 1024) = 104 Overflow blocks(Number of Total Disk Blocks - Overflow blocks) = (27917287424 - 104) = 27917287320Therefore, block device size closest to but less than a multiple of agsize is 27917287320. |
5. 将其转换为文件系统块大小的倍数(检查 xfs_info;此示例中为 4096),这就是增加文件系统的大小:
(Block Size multiple of agsize * device block size) / Filesystem block size = (27917287320 * 1024) / 4096 = 6979321830 filesystem blocks |
6. 将文件系统扩展到这个新的块大小:
# xfs_growfs -D 6979321830 /dev/VGDATA/data1附:红帽官方解决方案说明,有订阅的朋友可以自行阅读原文
Growing XFS filesystem with xfs_growfs fails with "XFS_IOC_FSGROWFSDATA xfsctl failed: Invalid argument" on Red Hat Enterprise Linux 5
(https://access.redhat.com/solutions/57262)
最后
测试环境上的偶然骚操作遇到这么一个坑,最后按照红帽的文档把这个问题解决了。但是我不知道这个问题怎么会在kernel 5.10的系统上遇到,看来我们的操作系统还有很多坑要踩啊,加油吧!!!!