PostgreSQL码农集散地

PolarDB 100 问 | 如何处理 bulkload 编译异常

参考文档点击文末阅读原文打开; 推荐《最好的PostgreSQL学习镜像》;


PolarDB 100 问 | PolarDB 11 编译 pg_bulkload 插件报错

问题: PolarDB 11 编译 pg_bulkload 插件报错

报错如下:

writer_binary.c: In function ‘open_output_file’:  
writer_binary.c:430:14: error: too few arguments to function ‘BasicOpenFilePerm’  
 430 |         fd = BasicOpenFilePerm(fname, O_WRONLY | O_CREAT | O_EXCL | PG_BINARY,  
     |              ^~~~~~~~~~~~~~~~~  
In file included from /home/postgres/tmp_basedir_polardb_pg_1100_bld/include/server/utils/sharedtuplestore.h:17,  
                from /home/postgres/tmp_basedir_polardb_pg_1100_bld/include/server/nodes/execnodes.h:27,  
                from ../include/reader.h:18,  
                from ../include/binary.h:15,  
                from writer_binary.c:19:  
/home/postgres/tmp_basedir_polardb_pg_1100_bld/include/server/storage/fd.h:115:17: note: declared here  
 115 | extern int      BasicOpenFilePerm(const char *fileName, int fileFlags, mode_t fileMode, bool polar_vfs);  
     |                 ^~~~~~~~~~~~~~~~~  
make[1]: *** [<builtin>: writer_binary.o] Error 1  
make[1]: Leaving directory '/tmp/pg_bulkload/lib'
make: *** [Makefile:27: all] Error 2  

复现方法

1、搭建PolarDB开发环境, 在开发环境中通过源码编译安装PolarDB 11.

1.1、拉取一个你熟悉的操作系统的PolarDB开发环境Docker镜像, 例如ubuntu22.04:

docker pull registry.cn-hangzhou.aliyuncs.com/polardb_pg/polardb_pg_devel:ubuntu22.04      

1.2、创建并运行容器

docker run -d -it -P --shm-size=1g --cap-add=SYS_PTRACE --cap-add SYS_ADMIN --privileged=true --name polardb_pg_devel registry.cn-hangzhou.aliyuncs.com/polardb_pg/polardb_pg_devel:ubuntu22.04 bash      

1.3、进入容器

# 进入容器    
docker exec -ti polardb_pg_devel bash      

1.4、克隆PolarDB 11源码,编译部署 PolarDB-PG 实例。

# 例如这里拉取 POLARDB_11_STABLE 分支;     
# PS: 截止2024.9.24 PolarDB开源的最新分支为: POLARDB_15_STABLE      
cd /tmp      
git clone -c core.symlinks=true --depth 1 -b POLARDB_11_STABLE https://github.com/ApsaraDB/PolarDB-for-PostgreSQL      

# 编译PolarDB 11并初始化实例    
cd /tmp/PolarDB-for-PostgreSQL      
./polardb_build.sh --without-fbl --debug=off      

# 验证PolarDB-PG      
psql -c 'SELECT version();'

           version                  
--------------------------------      
PostgreSQL 11.9 (POLARDB 11.9)      
(1 row)    

# 在容器内关闭、启动PolarDB数据库方法如下:      
pg_ctl stop -m fast -D ~/tmp_master_dir_polardb_pg_1100_bld        
pg_ctl start -D ~/tmp_master_dir_polardb_pg_1100_bld        

2、在容器中下载pg_bulkload源码:

cd /tmp    
git clone -c core.symlinks=true --depth 1 https://github.com/ossc-db/pg_bulkload  

3、编译pg_bulkload

cd /tmp/pg_bulkload    
USE_PGXS=1 make install    

报错如下:

In file included from /home/postgres/tmp_basedir_polardb_pg_1100_bld/include/server/storage/pg_shmem.h:30,  
                from recovery.c:29:  
/home/postgres/tmp_basedir_polardb_pg_1100_bld/include/server/utils/hsearch.h:78:9: error: unknown type name ‘MemoryContext’  
  78 |         MemoryContext hcxt;                     /* memory context to use for allocations */  
     |         ^~~~~~~~~~~~~  
make[1]: *** [<builtin>: recovery.o] Error 1  
make[1]: Leaving directory '/tmp/pg_bulkload/bin'
make: *** [Makefile:27: all] Error 2  

看报错应该是缺少包含MemoryContext的头文件引用, 搜索代码, 找到了这个定义在utils/palloc.h中.

/home/postgres/tmp_basedir_polardb_pg_1100_bld/include/server/utils/palloc.h    

... ...  
/*  
 * Type MemoryContextData is declared in nodes/memnodes.h.  Most users  
 * of memory allocation should just treat it as an abstract type, so we  
 * do not provide the struct contents here.  
 */  
typedef struct MemoryContextData *MemoryContext;  
... ...  

修改报错的recovery.c代码:

vi bin/recovery.c  

... ...  
#include "storage/bufpage.h"  
// 新增一个 include :    
#include "utils/palloc.h"  
#include "storage/pg_shmem.h"  
... ...  

再次编译, 出现了新的报错:

cd /tmp/pg_bulkload    
USE_PGXS=1 make install    

# 报错如下:  
writer_binary.c: In function ‘open_output_file’:  
writer_binary.c:430:14: error: too few arguments to function ‘BasicOpenFilePerm’  
 430 |         fd = BasicOpenFilePerm(fname, O_WRONLY | O_CREAT | O_EXCL | PG_BINARY,  
     |              ^~~~~~~~~~~~~~~~~  
In file included from /home/postgres/tmp_basedir_polardb_pg_1100_bld/include/server/utils/sharedtuplestore.h:17,  
                from /home/postgres/tmp_basedir_polardb_pg_1100_bld/include/server/nodes/execnodes.h:27,  
                from ../include/reader.h:18,  
                from ../include/binary.h:15,  
                from writer_binary.c:19:  
/home/postgres/tmp_basedir_polardb_pg_1100_bld/include/server/storage/fd.h:115:17: note: declared here  
 115 | extern int      BasicOpenFilePerm(const char *fileName, int fileFlags, mode_t fileMode, bool polar_vfs);  
     |                 ^~~~~~~~~~~~~~~~~  
make[1]: *** [<builtin>: writer_binary.o] Error 1  
make[1]: Leaving directory '/tmp/pg_bulkload/lib'
make: *** [Makefile:27: all] Error 2  

看起来是pg_bulkload调用的BasicOpenFilePerm被PolarDB修改过, 和PostgreSQL 11的不兼容了.

如何解决这个报错呢?

解决办法

1、搜索PolarDB 11源码中BasicOpenFilePerm的定义

/home/postgres/tmp_basedir_polardb_pg_1100_bld/include/server/storage/fd.h  


/* POLAR: add polar_vfs parameter */  
extern int      BasicOpenFilePerm(const char *fileName, int fileFlags, mode_t fileMode, bool polar_vfs);  

相比PostgreSQL 11多了一个bool polar_vfs参数.

这个polar_vfs到底应该设置为true还是false呢?

简单可以这么来理解: PolarDB 支持共享存储架构, 多个计算节点可以访问同一份数据, 因为各个机器有本地缓存(cache/file meta cache等) 访问共享存储时不能像单机PostgreSQL实例那样使用本地文件系统和本地文件操作的系统调用接口, 而是需要通过一层PolarDB自己定义的pfs, 来保证共享数据的一致性等.

看PolarDB代码的话, 很多接口都适配了polar_vfs, 例如:

/tmp/PolarDB-for-PostgreSQL/src/backend/storage/file/fd.c  


typedef struct vfd  
{  
       int                     fd;                             /* current FD, or VFD_CLOSED if none */  
       unsigned short fdstate;         /* bitflags for VFD's state */  
       ResourceOwner resowner;         /* owner, for automatic cleanup */  
       File            nextFree;               /* link to next free VFD, if in freelist */  
       File            lruMoreRecently;        /* doubly linked recency-of-use list */  
       File            lruLessRecently;  
       off_t           seekPos;                /* current logical file position, or -1 */  
       off_t           fileSize;               /* current size of file (0 if not temporary) */  
       char       *fileName;           /* name of file, or NULL for unused VFD */  
       /* NB: fileName is malloc'
d, and must be free'd when closing the VFD */  
       int                     fileFlags;              /* open(2) flags for (re)opening the file */  
       mode_t          fileMode;               /* mode to pass to open(2) */  

       bool            polar_vfs;                      /* POLAR: Whether to call the vfs interface */  
} Vfd;  

... ...  
fsync_fname(const char *fname, bool isdir, bool polar_vfs)  
durable_rename(const char *oldfile, const char *newfile, int elevel, bool polar_vfs)  
fsync_fname_ext(const char *fname, bool isdir, bool ignore_perm, int elevel, bool polar_vfs)  
fsync_parent_path(const char *fname, int elevel, bool polar_vfs)  
BasicOpenFile(const char *fileName, int fileFlags, bool polar_vfs)  
BasicOpenFilePerm(const char *fileName, int fileFlags, mode_t fileMode, bool polar_vfs)  
PathNameOpenFile(const char *fileName, int fileFlags, bool polar_vfs)  
PathNameOpenFilePerm(const char *fileName, int fileFlags, mode_t fileMode, bool polar_vfs)  
OpenTransientFile(const char *fileName, int fileFlags, bool polar_vfs)  
OpenTransientFilePerm(const char *fileName, int fileFlags, mode_t fileMode, bool polar_vfs)  
AllocateDir(const char *dirname, bool polar_vfs)  
AllocateDirPerm(const char *dirname, bool polar_vfs)  
MakePGDirectory(const char *directoryName, bool polar_vfs)  

还有很多不一一举例  

所以简单的让pg_bulkload编译通过的话, 可以直接写死true or false. 例如

找到pg_bulkload源码中调用了BasicOpenFilePerm的代码文件

cd /tmp/pg_bulkload  

grep -r BasicOpenFilePerm *  
lib/writer_direct.c: self->lsf_fd = BasicOpenFilePerm(self->lsf_path,  
lib/writer_direct.c: fd = BasicOpenFilePerm(fname, O_CREAT | O_WRONLY | PG_BINARY, S_IRUSR | S_IWUSR);  
lib/writer_binary.c: fd = BasicOpenFilePerm(fname, O_WRONLY | O_CREAT | O_EXCL | PG_BINARY,  

假设全部改成 true

vi lib/writer_binary.c  

#if PG_VERSION_NUM >= 110000  
       fd = BasicOpenFilePerm(fname, O_WRONLY | O_CREAT | O_EXCL | PG_BINARY,  
                                          S_IRUSR | S_IWUSR, true);  
vi lib/writer_direct.c  

#if PG_VERSION_NUM >= 110000  
       self->lsf_fd = BasicOpenFilePerm(self->lsf_path,  
               O_CREAT | O_EXCL | O_RDWR | PG_BINARY, S_IRUSR | S_IWUSR, true);  

#if PG_VERSION_NUM >= 110000  
       fd = BasicOpenFilePerm(fname, O_CREAT | O_WRONLY | PG_BINARY, S_IRUSR | S_IWUSR, true);  

重新编译就正常了.

$ cd /tmp/pg_bulkload     
$ USE_PGXS=1 make install    

现在可以在PolarDB 11中创建pg_bulkload插件了:

$ psql    
psql (11.9)    
Type "help"forhelp.    

postgres=# select version();    
           version                
--------------------------------    
PostgreSQL 11.9 (POLARDB 11.9)    
(1 row)    

postgres=# create extension pg_bulkload;  
CREATE EXTENSION  

但是在使用pg_bulkload时又遇到了报错, 请看下回分解.

文末彩蛋:国产数据库周边生态

当然一款数据库要流行起来, 除了自己要强大, 还离不开生态. 用好周边生态工具, 管理水平战胜90%老司机!!! 下面简单介绍一下国产数据库周边生态.

1、管控软件

鸣嵩(前阿里云数据库总经理 / 研究员)等大佬们创业创办的云猿生, 核心产品是KubeBlocks. 他们的理念是让管理数据库和搭积木一样简单, 如果你要管理很多套并且种类(OLTP\OLAP\NoSQL\KV\TS\MQ等)很多的数据库产品, 推荐首选.

  • https://github.com/apecloud/kubeblocks

PG中文社区核心委员唐成老师的公司乘数开源的Clup, 专用管理PostgreSQL和PolarDB的集群管理软件, 如果你要管理很多套数据库, 推荐选择. 并且Clup还提供了企业版、自研的连接池、分布式存储、一体机、备份平台等, 是企业用户推荐之选.

  • https://www.csudata.com/

若航老司机开源的pigsty, 集成了300多个PG插件的PG集群和PolarDB集群管理软件, 如果你要管理很多套PG或PolarDB数据库, 且对插件有特别多的需求, 推荐选择.

  • https://pigsty.cc/zh/

2、审计监控诊断优化

翟总(曾经是我背后的男人)到海信聚好看后研发的 DBdoctor, 采用ebpf技术, 在对数据库几乎没有影响的情况下实时监控数据库和服务器的各项指标, 发现和诊断问题根因非常方便.

  • https://www.dbdoctor.cn/

天舟老哥的核心产品Bytebase 是位于您和数据库之间的中间件。它是数据库 DevOps 的 GitLab/GitHub,专为开发人员、DBA 和平台工程师打造。

  • https://bytebase.cc/docs/introduction/what-is-bytebase/

PawSQL, SQL优化和诊断产品.  

D-Smart, Oracle老前辈白老大出品, 专注企业级市场, 将业界顶级DBA经验的产品化作品, 产品功能包括数据库监控、诊断、优化等.

  • https://www.modb.pro/db/567140

3、国产数据库IDE

IDE是开发者的必备工具,例如社区有pgAdmin, 国产IDE则可以看看老程序猿达刚老师的DeskUI:

  • https://www.deskui.com

4、数据同步&迁移&备份恢复

NineData, 老领导出去创业做的产品, 产品涵盖了数据同步、迁移、备份、比对、devops、chatDBA等.

  • https://www.ninedata.cloud/home

DSG, 非常老牌的数据库同步迁移企业级产品, 支持各种数据库的异构和同构迁移, 用他们的话说, 没有dsg搞不定的迁移, 比goldengate还牛.

  • https://www.dsgdata.com/

公开课

如果你对PolarDB学习感兴趣可以阅读这个公开课系列:

除了PolarDB还非常值得关注的几款PG栈国产数据库:

  • HaloDB(基于PG兼容PostgreSQL、Oracle、MySQL. http://www.halodbtech.com/ )、
  • IvorySQL(基于开源PG兼容PG、Oracle. https://www.ivorysql.org/zh-cn/ )、
  • ProtonBase(云原生分布式数仓. https://protonbase.com/ )、
  • 成都文武数据库(https://ww-it.cn)

参考文档点击阅读原文获得


感谢关注我的github (https://github.com/digoal/blog) 及视频号:

Image