开源软件联盟PostgreSQL分会

实战PG vector构建DBA个人知识库之一:基础环境搭建以及大模型部署

作者

陈旭,捷银消费金融有限公司DBA

大家好,今天和大家聊聊PG vector 构建个人知识库。
可能大多数朋友和我一样,工作了几年之后,累计了大量的邮件,PPT,word 文档,PDF 资料, 很难有时间自己整理。
到后来可以采用麦库笔记(已停服),印象笔记, 墨天轮博客来记录总结一些自己学习的知识和分享一些工作中的经验。
到了AI时代,我想做的是把手头的一些资料,笔记通过LLM+RAG的方式重新来归纳到新的知识库管理平台中。

本次实践设计到的一些技术堆栈:(技术面有点广,但是每项都不深入。。。)

前台技术:html, js, vue
后端技术:Python , Longchain 框架
数据库技术:PG 16 以及 pg_vector 插件
机器部署篇:私有机器部署或者租用云服务部署LLM已经embedding 模型
LLM模型:Llama3-Chinese
Embedding 模型:text2vec-large-chinese
AI 开发助手:CHATGPT 4 , 通义千问平台
数据集:麦库笔记,印象笔记, wiki 网页, word 技术文档…

作为实战PG的开场篇,我们需要搭建一下大模型的环境,需要物理机1一台(需要带有GPU的显卡),如何朋友们没有机器的话,可以租用云服务…
或者部署在自己的笔记本上(总之根据个人的财力量力而行就行)

我们这里采用的是算力云的服务: https://www.autodl.com/home
Image
关于模型的GPU选型问题可以参考: https://www.autodl.com/docs/gpu/

我们去算力市场租用一台机器: https://www.autodl.com/market/list 我们选择一台 RTX 3090的机器 , 每个小时才1.66 元
Image
我们选择按流量计费,这样在不用的时候可以选择关闭机器来节省一部分经费。
Image
关于基础镜像我们选择:pytorch 1.7.0 版本对应的 cuda版本 11.0
Image
我们可以通过web终端来访问这台机器:
Image
查看显卡信息:nvidia-smi
Image
对于大模型的选择,我们可以选择一个支持中文的大模型,我们访问一下魔搭社区 https://www.modelscope.cn/home

最终我们选择了 Llama3-Chinese 版本 https://www.modelscope.cn/models/seanzhang/Llama3-Chinese

Llama3-Chinese是以Meta-Llama-3-8B为底座,使用 DORA + LORA+ 的训练方法,在50w高质量中文多轮SFT数据 + 10w英文多轮SFT数据 + 2000单轮自我认知数据训练而来的大模型。

Image
模型下载界面: 主要的模型文件分为5个文件,大小不到20GB。
Image
我们在刚才租用的算力云的服务器上,设置一下资源下载加速:

  1. root@autodl-container-16e1448282-c9f3dde0:~/autodl-tmp# source /etc/network_turbo

  2. 设置成功

  3. 注意:仅限于学术用途,不承诺稳定性保证

安装git-lfs

  1. root@autodl-container-16e1448282-c9f3dde0:~/autodl-tmp# curl -s https://packagecloud.io/install/repositories/github/git-lfs/script.deb.sh | sudo bash

  2. Detected operating system asUbuntu/bionic.

  3. Checkingfor curl...

  4. Detected curl...

  5. Checkingfor gpg...

  6. Detected gpg...

  7. Detected apt version as1.6.12ubuntu0.1

  8. Running apt-get update...done.

  9. Installing apt-transport-https...done.

  10. Installing/etc/apt/sources.list.d/github_git-lfs.list...done.

  11. Importing packagecloud gpg key...Packagecloud gpg key imported to /etc/apt/keyrings/github_git-lfs-archive-keyring.gpg

  12. done.

  13. Running apt-get update...done.

  14. The repository is setup!You can now install packages.

  15. root@autodl-container-16e1448282-c9f3dde0:~/autodl-tmp# apt install git-lfs

  16. Readingpackage lists...Done

  17. Building dependency tree

  18. Reading state information...Done

  19. The following NEW packages will be installed:

  20. git-lfs

  21. 0 upgraded,1 newly installed,0 to remove and176not upgraded.

  22. Need to get7,168 kB of archives.

  23. Afterthis operation,15.6 MB of additional disk space will be used.

  24. Get:1 https://packagecloud.io/github/git-lfs/ubuntu bionic/main amd64 git-lfs amd64 3.2.0 [7,168 kB]

  25. Fetched7,168 kB in3s(2,139 kB/s)

  26. debconf: delaying package configuration, since apt-utils isnot installed

  27. Selecting previously unselected package git-lfs.

  28. (Reading database ...42112 files and directories currently installed.)

  29. Preparing to unpack .../git-lfs_3.2.0_amd64.deb ...

  30. Unpacking git-lfs (3.2.0)...

  31. Setting up git-lfs (3.2.0)...

  32. Git LFS initialized.

下载大模型文件:总大小在30GB左右

  1. root@autodl-container-16e1448282-c9f3dde0:~/autodl-tmp/LLM# git clone https://www.modelscope.cn/seanzhang/Llama3-Chinese.git

  2. Cloninginto'Llama3-Chinese'...

  3. remote:Enumerating objects:43,done.

  4. remote:Counting objects:100%(43/43),done.

  5. remote:Compressing objects:100%(41/41),done.

  6. remote:Total43(delta 11), reused 0(delta 0), pack-reused 0

  7. Unpacking objects:100%(43/43),done.

  8. Filtering content:100%(5/5),14.95GiB|50.68MiB/s,done.

  9. root@autodl-container-16e1448282-c9f3dde0:~/autodl-tmp/LLM# ls -lhtr

  10. total 4.0K

  11. drwxr-xr-x 4 root root 4.0KJul2216:23Llama3-Chinese

  12. root@autodl-container-16e1448282-c9f3dde0:~/autodl-tmp/LLM# du -h ./ --max-depth=1

  13. 30G./Llama3-Chinese

  14. 30G./

安装依赖的包

  1. pip install transformers

测试部署的大模型代码

  1. from transformers importAutoTokenizer,AutoModelForCausalLM

  2. model_id ="/root/autodl-tmp/LLM/Llama3-Chinese"

  3. tokenizer =AutoTokenizer.from_pretrained(model_id)

  4. model =AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="auto", device_map="auto")

  5. messages =[

  6. {"role":"system","content":"You are a helpful assistant."},

  7. {"role":"user","content":"""鲁迅(1881年9月25日—1936年10月19日),原名周樟寿,后改名周树人,

  8. 字豫山,后改字豫才,浙江绍兴人。

  9. 中国著名文学家、思想家、革命家、教育家、美术家、书法家、民主战士,新文化运动的重要参与者,中国现代文学的奠基人之一.

  10. 请问鲁迅活了多少岁?"""},

  11. ]

  12. input_ids = tokenizer.apply_chat_template(

  13. messages, add_generation_prompt=True, return_tensors="pt"

  14. ).to(model.device)

  15. outputs = model.generate(

  16. input_ids,

  17. max_new_tokens=2048,

  18. do_sample=True,

  19. temperature=0.7,

  20. top_p=0.95,

  21. )

  22. response = outputs[0][input_ids.shape[-1]:]

  23. print(tokenizer.decode(response, skip_special_tokens=True))

测试一下大模型的理解计算能力:

输入为:鲁迅(1881年9月25日—1936年10月19日),原名周樟寿,后改名周树人,
字豫山,后改字豫才,浙江绍兴人。
中国著名文学家、思想家、革命家、教育家、美术家、书法家、民主战士,新文化运动的重要参与者,中国现代文学的奠基人之一.

提问是:请问鲁迅活了多少岁?

大模型的回答是:鲁迅活了55岁。答案源于计算:(1881年9月25日—1936年10月19日)\
Image

最后我们尝试为本地部署的大模型发布一个API的接口, 这里我们用PYTHON自带 fastAPI 包实现:
安装包:

  1. pip3 install fastapi

发布一个查询接口:

  1. from transformers importAutoTokenizer,AutoModelForCausalLM

  2. import uvicorn

  3. from fastapi importFastAPI

  4. app =FastAPI()

  5. model_id ="/root/autodl-tmp/LLM/Llama3-Chinese"

  6. tokenizer =AutoTokenizer.from_pretrained(model_id)

  7. model =AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="auto", device_map="auto")

  8. @app.get("/llm_query/{query}")

  9. def llm_query(query):

  10. message =[

  11. {"role":"system","content":"You are a helpful assistant."},

  12. {"role":"user","content":"""{}""".format(query)},

  13. ]

  14. input_ids = tokenizer.apply_chat_template(

  15. message, add_generation_prompt=True, return_tensors="pt"

  16. ).to(model.device)

  17. outputs = model.generate(

  18. input_ids,

  19. max_new_tokens=2048,

  20. do_sample=True,

  21. temperature=0.7,

  22. top_p=0.95,

  23. )

  24. response = outputs[0][input_ids.shape[-1]:]

  25. return tokenizer.decode(response, skip_special_tokens=True);

  26. if __name__ =='__main__':

  27. print('-----start-------')

  28. #init_load()

  29. uvicorn.run(app, host="0.0.0.0", port=8868)

启动程序发布接口:

  1. root@autodl-container-16e1448282-c9f3dde0:~/autodl-tmp/scripts# vi LLMHandler.py

  2. root@autodl-container-16e1448282-c9f3dde0:~/autodl-tmp/scripts# python LLMHandler.py &

  3. [1]1404

  4. root@autodl-container-16e1448282-c9f3dde0:~/autodl-tmp/scripts#

  5. root@autodl-container-16e1448282-c9f3dde0:~/autodl-tmp/scripts# Special tokens have been added in the vocabulary, make sure the associated word embeddings are fine-tuned or trained.

  6. Loading checkpoint shards:100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████|5/5[00:03<00:00,1.26it/s]

  7. -----start-------

  8. INFO:Started server process [1404]

  9. INFO:Waitingfor application startup.

  10. INFO:Application startup complete.

  11. INFO:Uvicorn running on http://0.0.0.0:8868 (Press CTRL+C to quit)

发布服务后,我们还需要打开端口,参考: https://www.autodl.com/docs/port/

点击自定义服务:
Image

浏览器调用测试一下:

输入问题:你好,分享一下夏天的解暑方式?
LLM回答:“夏天的解暑方式有很多种,比如饮用清凉饮料,多喝水,做些轻松的户外活动,利用空调等冷却设备等等。您有什么特别的需求或偏好吗?”

Image

当然你也可以写一段小代码调用LLM的接口:

Image

最后总结:
1.我们完成大模型Llama3-Chinese 在算力云上的私有化部署
2.我们通过fastAPI发布了大模型的调用接口

ImageImage