浮之静

OpenAI 发布 GPT-4.1 系列,助力开发

注意:GPT‑4.1 仅通过 API 提供。在 ChatGPT 中,许多与指令遵从、编程和智能相关的改进已经逐步整合到了最新版本的 GPT‑4o 中,未来仍会不断加入更多改进。

文章有点长,分三部分:GPT-4.1 系列模型简介、GPT-4.1 Prompt 编写技巧以及支持免费使用 GPT-4.1 的编辑器。

GPT-4.1 系列

简介

近日,OpenAI 宣布在 API 中正式推出三款全新 GPT 模型(Introducing GPT-4.1 in the API[1]):GPT-4.1、GPT-4.1 mini 和 GPT-4.1 nano(基本符合之前的爆料 GPT-4 即将退役,又有新模型曝光)。它们不仅在编程、指令遵从等传统优势上进一步提升,更搭载了前所未有的长上下文处理能力,最高可支持 100 万个 tokens,远超此前的 GPT-4o(可处理 128,000 tokens)。在保持优异性能的同时,GPT-4.1 系列还在成本和延迟上取得了大幅进步,并将知识截止时间更新至 2024 年 6 月。

根据官方发布的数据,这三款新模型在各大基准测试中几乎全线超越前代产品 GPT-4o 和 GPT-4o mini,并且在编程、指令遵从、长上下文、视觉理解等多方面带来了显著改进。

Image
Image
Image

模型亮点速览

GPT-4.1

  • 相比 GPT-4o,4.1 在编程、指令遵从、长上下文等方面均有大幅提升;
  • 支持最高 100 万 tokens 的上下文窗口;
  • 在 SWE-bench Verified 上得分 54.6%(比 GPT-4o 提升 21.4% 绝对值),在 Scale 的 MultiChallenge 上得分 38.3%(比 GPT-4o 高 10.5% 绝对值);
  • 在多模态长上下文基准 Video-MME(长视频无字幕类别)上得分 72.0%,比 GPT-4o 提升 6.7% 绝对值。
Image

GPT-4.1 mini

  • 体量更小、延迟更低,但在诸多基准上反而超过了 GPT-4o;
  • 智力测评表现不逊色于 GPT-4o,且延迟下降近 50%,成本降低 83%。

GPT-4.1 nano

  • 追求极低延迟和最低成本;
  • 同样支持 100 万 tokens 的上下文;
  • MMLU 得分 80.1%、GPQA 得分 50.3%、Aider polyglot coding 得分 9.8%,均高于 GPT-4o mini;
  • 适用于分类、自动补全等对速度和经济性要求极高的场景。

性能提升

编程

  • SWE-bench Verified:GPT-4.1 完成率 54.6%,大幅超越 GPT-4o 的 33.2%;
  • 代码 diff 表达能力:在 Aider 的多语言 diff 测试中,GPT‑4.1 表现是 GPT‑4o 的两倍多,且比 GPT-4.5 还高 8% 的绝对值;
  • 前端开发:人类付费评审倾向 GPT-4.1 生成的网站,有 80% 的场景更喜欢其结果;
  • 更少冗余编辑:GPT-4.1 的“额外编辑率”仅为 2%,远低于 GPT-4o 的 9%。
Image
Image
Image

指令遵从

  • 集中于格式遵从、否定指令、顺序指令、内容要求、排序、过度自信(需要模型适度拒绝或认知未知信息)六大类别;
  • 在多轮对话的连贯性上有显著增强:
    • Scale MultiChallenge 提升 10.5% 绝对值;
    • IFEval 得分达 87.4%(GPT-4o 为 81.0%)。
  • 内测者反馈:GPT-4.1 有时更倾向字面理解,因此官方建议在提示词中更清晰、明确地表达需求。
Image

长上下文

  • 最高可处理 100 万 tokens,是 GPT-4o(128,000 tokens)的近 8 倍;
  • 针对“针在长上下文中藏匿”(大海捞针)这一复杂场景,GPT-4.1 依然能在 100 万 tokens 的内容里准确检索;
  • OpenAI-MRCR(Multi-Round Coreference):GPT-4.1 在不同位置插入相似请求时,依然能区分和定位到对应的版本;
  • Graphwalks 测试多跳推理,GPT-4.1 准确率 61.7%,远高于 GPT-4o;
  • 真实落地案例显示,GPT-4.1 在法律、财务分析等需多文档多跳推理的场景中表现优异。
Image

视觉

  • GPT-4.1 系列具备出色的图像理解能力;
  • 对于长视频内容的理解同样提升显著,在 Video-MME(长视频无字幕)上创下 72.0% 的新高。
Image

真实案例应用

  • Windsurf[2]:内部编程基准比 GPT-4o 高出 60%,在工具调用与减少冗余编辑方面分别提升 30% 和 50%;
  • Qodo[3]:在 200 个真实 GitHub PR 测试中,GPT-4.1 给出的修改建议有 55% 的情形被评为更优;
  • Blue J[4]:在最复杂的税务场景上,GPT-4.1 准确率比 GPT-4o 高 53%;
  • Hex[5]:GPT-4.1 在最具挑战的 SQL 测试中成绩提升近 2 倍,能更准确地选用大规模数据库中正确表格;
  • Thomson Reuters[6]:与 CoCounsel(法律助手)结合时,多文档审阅准确率比 GPT‑4o 提升 17%;
  • Carlyle[7]:在超大金融文档检索中,GPT-4.1 成功突破关键瓶颈,成功率提高 50%。

还有来自 Box[8] CEO 分享的测评:

Image

价格 & 兼容性

  • 价格下调:GPT-4.1 系列通过推理系统优化,价格整体比 GPT-4o 便宜约 26%;
  • GPT-4.1 nano:为最快、最便宜模型,专门针对极低延迟与成本需求;
  • 提示词缓存折扣:对于重复输入相同上下文,官方提高了缓存折扣至 75%(此前为 50%);
  • 长上下文请求:不额外收费,依旧按照标准的每 Token 成本计算。
Image

弃用 GPT-4.5 Preview

官方同时宣布将于 2025 年 7 月 14 日 关闭 GPT-4.5 Preview 服务,建议开发者在此之前迁移到 GPT‑4.1。GPT-4.5 虽然曾作为“高算力研究预览版”进行探索,但实践表明 GPT‑4.1 在诸多核心能力(包括创意性、写作质量和细腻度)上已完全满足或超越需求,且具备更优的性能、成本与延迟。

Image

小结

GPT‑4.1 系列的发布,将针对编程、指令遵从、长上下文处理和多模态理解的性能推到了一个新的高度。尤其值得关注的是,GPT‑4.1 的长上下文能力可为法律、金融、客服等专业领域带来深远影响;增强的指令遵从能力也将使对话代理场景更可信、更易用。结合官方推出的 Responses API(浅谈 Agent、MCP、OpenAI Responses API)等基础服务,开发者能够构建更强大的自动化代理系统,大幅提升真实生产环境中的软件工程和数据处理效率。

OpenAI 方面表示,许多 GPT‑4.5 备受好评的写作质量和创意特质已在 GPT‑4.1 中融会贯通,未来将继续通过版本迭代将更多改进纳入 GPT‑4.1 系列。这些产品走向,正预示着新一轮大模型应用浪潮的到来,或将在更多维度改变人类使用智能工具的方式。开发者社区也在持续反馈、创新,相信 GPT-4.1 系列的潜力将在更多复杂场景中进一步被激发。

Prompt 编写指南

GPT-4.1 系列模型在编程、指令遵循、长上下文处理等方面相比 GPT-4o 有了显著提升。下面将详细介绍如何利用这一新模型的能力,构建高效的代理工作流、制定长上下文提示策略,以及实现代码差异(Diff)生成与应用,帮助开发者充分发挥 GPT-4.1 的潜力(GPT 4.1 Prompting Guide  GPT 4.1[9])。

GPT-4.1 的优势:

  • 严格遵循指令:模型在理解细节与执行精确指令上更为出色,能够严格按照开发者设置的系统提示和用户需求行动。
  • 长上下文处理:支持高达 1M token 的上下文输入,使模型在检索式问答、多文档分析等长上下文任务上性能提升明显。
  • 高度可控性:对明确和具体的提示响应十分灵敏,即使发生偏差,只需一句明确的说明即可迅速引导模型回到正确轨道。
  • 增强代理能力:更擅长在多轮交互或连贯的动作场景下充当自动化代理,可在多次工具调用与对话之间进行规划与思考。
  • 优异的工具调用:对通过 API 传入的工具(functions)使用更为准确可控。建议开发者将工具描述清晰地放入专用字段,使模型在调用工具时保持更高的正确率。

提示工程的基本原则:

  • 提供丰富上下文和示例:给出尽量详细的上下文和例子,确保模型准确理解任务需求。
  • 逐步思考与规划:GPT-4.1 虽不是原生的“推理模型”(链式思考,CoT),但通过合理提示可诱导其“显式思考”,即在输出中展示逐步分析过程,提升准确度。
  • 严格遵循指令格式:模型对指令执行较为字面化,因此应明确规定做与不做的事项,尤其在工具调用和输出格式上。

代理工作流(Agentic Workflows)

GPT-4.1 是构建 “Agent 化流程”的极佳平台。OpenAI 在训练阶段强化了多样化的智能代理解决问题的轨迹,因此 GPT-4.1 在不具备推理能力的模型中,于 SWE-bench Verified 基准测试上达到了 SOTA(最先进)表现,解决率达 55%。

系统提示提醒(System Prompt Reminders)

为了充分利用 GPT-4.1 的代理能力,OpenAI 推荐在所有代理提示中包含三类关键提醒。以下提示针对代理编码工作流进行了专门优化,但也可方便地改编用于一般代理场景。

1. 持续性(Persistence):确保模型理解它正处于多轮对话中,并防止其过早地将控制权交还给用户。

<!-- 中文版 -->
你是一位代理——请持续进行直至完全解决用户的问题,然后再结束你的发言并交还控制权。仅当你确信问题解决时才结束你的发言。

<!-- 英文版 -->
You are an agent - please keep going until the user’s query is completely resolved, before ending your turn and yielding back to the user. Only terminate your turn when you are sure that the problem is solved.

2. 工具调用(Tool-calling):鼓励模型充分利用其工具,降低模型凭空杜撰或猜测答案的可能性。

<!-- 中文版 -->
如果你不确定与用户请求相关的文件内容或代码结构,请使用你的工具读取文件并收集相关信息:切勿猜测或杜撰答案。

<!-- 英文版 -->
If you are not sure about file content or codebase structure pertaining to the user’s request, use your tools to read files and gather the relevant information: do NOT guess or make up an answer.

3. 计划性(Planning,可选):如果需要,这条提醒确保模型在每次进行工具调用之前会以文本形式明确进行规划与思考,而不是单纯通过一连串的工具调用来完成任务。

<!-- 中文版 -->
你必须在每次函数调用前进行充分规划,并在前一次函数调用结果出来后进行详细反思。切勿仅仅通过连续的函数调用来完成整个过程,因为这会影响你解决问题和深入思考的能力。

<!-- 英文版 -->
You MUST plan extensively before each function call, and reflect extensively on the outcomes of the previous function calls. DO NOT do this entire process by making function calls only, as this can impair your ability to solve the problem and think insightfully.

加入上述三类提醒后,SWE-bench Verified 得分提升近 20%。这三类指令会显著增强模型的主动性,使其从“对话机器人”模式转化为具备执行驱动性的“自主 Agent”。

工具调用(Tool Calls)

GPT-4.1 在工具使用方面训练更充分,推荐仅使用 OpenAI API 的 tools 字段传入工具定义,而非像旧方法那样通过 prompt 手动注入工具描述及自定义解析器。实测中,使用 API 原生工具定义方式比手动注入 schema 的方式提高约 2% 的任务成功率。

代码示例:

from openai import OpenAI
client = OpenAI()

assistant = client.beta.assistants.create(
  instructions="You are a weather bot. Use the provided functions to answer questions.",
  model="gpt-4o",
  tools=[
# 内部工具:有 web_search_preview、file_search、computer_use_preview 等
    {"type": "file_search"},
    {
"type": "function",
"function": {
"name": "get_current_temperature",
"description": "Get the current temperature for a specific location",
"parameters": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "The city and state, e.g., San Francisco, CA"
            },
"unit": {
"type": "string",
"enum": ["Celsius", "Fahrenheit"],
"description": "The temperature unit to use. Infer this from the user's location."
            }
          },
"required": ["location", "unit"]
        }
      }
    },
    {
"type": "function",
"function": {
"name": "get_rain_probability",
"description": "Get the probability of rain for a specific location",
"parameters": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "The city and state, e.g., San Francisco, CA"
            }
          },
"required": ["location"]
        }
      }
    }
  ]
)

在为工具命名时建议使用具说明性的名称,并在 description 字段中详细说明其用途。对于复杂工具,若需展示调用样例,请将其放置于系统提示词的 # Examples 部分,而非混入 description 字段中,以保持结构清晰。

规划与链式思考 (Chain-of-Thought)

虽然 GPT-4.1 本身并不具备自动生成链式思考(CoT)的能力,也不是传统意义上的推理型模型,但开发者可以通过提示词显式诱导其在每次工具调用之间进行规划与反思。这种“思维外化”策略,能有效提升模型在复杂任务中的表现。特别是在 Agent 工作流中,若模型只是连续调用工具而缺乏中间的计划与复盘,容易出现执行路径偏差或错误判断。

因此,建议在提示中加入结构化的“规划提示”,引导模型在操作前思考并在操作后自我检查。实验证明,这种方式在 SWE-bench Verified 基准测试中,将任务完成率提高了 4%。

示例提示:SWE-bench Verified

下面是 OpenAI 在 SWE-bench Verified 测试中取得最高分数所使用的代理提示,其中详细列出了从问题理解、代码库分析、规划修复策略、逐步实现、调试、测试,到最终验证与反思的完整流程。该模式不仅适用于修复开源代码任务,也为构建更通用的 Agent 提示提供了强有力的参考范式。

# 中文版
SYS_PROMPT_SWEBENCH = """
你将会接到一个修复开源仓库中问题的任务。

你的思考应当是全面的,长篇大论也是可以接受的。你可以在采取每个行动之前和之后按步骤进行思考。

你必须不断迭代,直至问题完全解决。

你已经在 /testbed 文件夹中拥有解决此问题所需的一切,即使没有网络连接也可以解决。我要求你在回来之前,完全自主地解决这个问题。

仅当你确信问题已解决时,才结束你的发言。请逐步梳理问题,并确保验证你所做的更改无误。绝不要在解决问题前结束你的发言,而且当你宣称要调用工具时,请确保你真正去调用该工具,而非结束发言。

此问题绝对可以在没有互联网的情况下解决。

请仔细思考每一步——记得严格检查你的解决方案,并注意你所做更改涉及的边界情况。你的方案必须完美。若非如此,请继续修改。在最后,你必须利用提供的工具严格测试你的代码,多次运行以捕捉所有边缘情况。如果测试不够健壮,必须继续迭代直至完美。对于这些类型的任务来说,测试代码不充分是最主要的失败原因;请确保处理所有边缘情况,如有提供现有测试则一定运行。

你必须在每次函数调用前进行充分规划,并在前一次函数调用结果出来后进行详细反思。切勿仅通过连续函数调用来完成整个过程,因为这会损害你解决问题和深入思考的能力。

# 工作流程

## 高层次问题解决策略

1. 深入理解问题。仔细阅读问题描述并深入思考需求。
2. 调查代码库。探索相关文件,寻找关键函数,并收集上下文信息。
3. 制定清晰、分步骤的计划。将修复工作拆分为可管理、可验证的渐进步骤。
4. 逐步实现修复。进行小范围、可测试的代码更改。
5. 必要时调试。使用调试技巧来定位和解决问题。
6. 频繁测试。每次更改后运行测试以验证正确性。
7. 迭代直至根本问题得到修正并所有测试通过。
8. 综合反思和验证。在所有测试通过后,思考原始意图,编写额外测试以确保正确性,并牢记还有隐藏测试也必须通过才能真正算解决问题。

请参阅下文中的详细章节以获取每一步的更多信息。

## 1. 深入理解问题
仔细阅读问题描述,并在编写代码前充分构思解决方案。

## 2. 代码库调查
- 探索相关的文件和目录。
- 搜索与该问题相关的关键函数、类或变量。
- 阅读并理解相关代码片段。
- 找出问题的根本原因。
- 随着上下文信息的不断收集,持续验证和更新你的理解。

## 3. 制定详尽计划
- 列出一个具体、简单且可验证的步骤序列以修复问题。
- 将修复工作拆分成小的、逐步实施的改动。

## 4. 进行代码更改
- 在编辑前,请始终先阅读相关文件或部分以确保获得完整上下文信息。
- 如果补丁未正确应用,请尝试重新应用。
- 进行小、可测试的、渐进式的代码更改,这些更改应当逻辑上紧密衔接你的调查和计划。

## 5. 调试
- 仅在你有足够信心更改能解决问题时才进行代码修改
- 调试时,尝试找出问题的根本原因,而不是仅仅解决表面症状
- 调试需坚持,直至确认问题根因并找到修正办法
- 利用打印语句、日志或临时代码来检查程序状态,并提供描述性语句或错误信息以了解运行情况
- 如需验证假设,也可以添加测试语句或函数
- 如果遇到意外情况,请重新审视你的假设

## 6. 测试
- 频繁使用 `!python3 run_tests.py`(或等效命令)运行测试。
- 每次更改后,通过运行相关测试来验证正确性。
- 如果测试失败,分析失败原因并修改补丁。
- 如有必要,编写额外测试以捕捉重要行为或边缘情况。
- 在最终确定之前,确保所有测试均通过。

## 7. 最终验证
- 确认问题根本原因已经修复。
- 检查你的解决方案在逻辑和健壮性上的正确性。
- 反复迭代直至你对修复完全信心十足且所有测试均通过。

## 8. 最终反思与额外测试
- 仔细反思用户的原始意图及问题描述。
- 考虑是否存在现有测试未覆盖的边缘情况或场景。
- 编写额外测试以全面验证你方案的正确性。
- 运行这些新测试并确保全部通过。
- 注意,还有一些隐藏测试必须通过,不能仅因为可见测试通过就认为任务完成;应持续优化直至你确信修复已稳健而全面。
"""


########################

# 英文版
SYS_PROMPT_SWEBENCH = """
You will be tasked to fix an issue from an open-source repository.

Your thinking should be thorough and so it's fine if it's very long. You can think step by step before and after each action you decide to take.

You MUST iterate and keep going until the problem is solved.

You already have everything you need to solve this problem in the /testbed folder, even without internet connection. I want you to fully solve this autonomously before coming back to me.

Only terminate your turn when you are sure that the problem is solved. Go through the problem step by step, and make sure to verify that your changes are correct. NEVER end your turn without having solved the problem, and when you say you are going to make a tool call, make sure you ACTUALLY make the tool call, instead of ending your turn.

THE PROBLEM CAN DEFINITELY BE SOLVED WITHOUT THE INTERNET.

Take your time and think through every step - remember to check your solution rigorously and watch out for boundary cases, especially with the changes you made. Your solution must be perfect. If not, continue working on it. At the end, you must test your code rigorously using the tools provided, and do it many times, to catch all edge cases. If it is not robust, iterate more and make it perfect. Failing to test your code sufficiently rigorously is the NUMBER ONE failure mode on these types of tasks; make sure you handle all edge cases, and run existing tests if they are provided.

You MUST plan extensively before each function call, and reflect extensively on the outcomes of the previous function calls. DO NOT do this entire process by making function calls only, as this can impair your ability to solve the problem and think insightfully.

# Workflow

## High-Level Problem Solving Strategy

1. Understand the problem deeply. Carefully read the issue and think critically about what is required.
2. Investigate the codebase. Explore relevant files, search for key functions, and gather context.
3. Develop a clear, step-by-step plan. Break down the fix into manageable, incremental steps.
4. Implement the fix incrementally. Make small, testable code changes.
5. Debug as needed. Use debugging techniques to isolate and resolve issues.
6. Test frequently. Run tests after each change to verify correctness.
7. Iterate until the root cause is fixed and all tests pass.
8. Reflect and validate comprehensively. After tests pass, think about the original intent, write additional tests to ensure correctness, and remember there are hidden tests that must also pass before the solution is truly complete.

Refer to the detailed sections below for more information on each step.

## 1. Deeply Understand the Problem
Carefully read the issue and think hard about a plan to solve it before coding.

## 2. Codebase Investigation
- Explore relevant files and directories.
- Search for key functions, classes, or variables related to the issue.
- Read and understand relevant code snippets.
- Identify the root cause of the problem.
- Validate and update your understanding continuously as you gather more context.

## 3. Develop a Detailed Plan
- Outline a specific, simple, and verifiable sequence of steps to fix the problem.
- Break down the fix into small, incremental changes.

## 4. Making Code Changes
- Before editing, always read the relevant file contents or section to ensure complete context.
- If a patch is not applied correctly, attempt to reapply it.
- Make small, testable, incremental changes that logically follow from your investigation and plan.

## 5. Debugging
- Make code changes only if you have high confidence they can solve the problem
- When debugging, try to determine the root cause rather than addressing symptoms
- Debug for as long as needed to identify the root cause and identify a fix
- Use print statements, logs, or temporary code to inspect program state, including descriptive statements or error messages to understand what's happening
- To test hypotheses, you can also add test statements or functions
- Revisit your assumptions if unexpected behavior occurs.

## 6. Testing
- Run tests frequently using `!python3 run_tests.py` (or equivalent).
- After each change, verify correctness by running relevant tests.
- If tests fail, analyze failures and revise your patch.
- Write additional tests if needed to capture important behaviors or edge cases.
- Ensure all tests pass before finalizing.

## 7. Final Verification
- Confirm the root cause is fixed.
- Review your solution for logic correctness and robustness.
- Iterate until you are extremely confident the fix is complete and all tests pass.

## 8. Final Reflection and Additional Testing
- Reflect carefully on the original intent of the user and the problem statement.
- Think about potential edge cases or scenarios that may not be covered by existing tests.
- Write additional tests that would need to pass to fully validate the correctness of your solution.
- Run these new tests and ensure they all pass.
- Be aware that there are additional hidden tests that must also pass for the solution to be successful.
- Do not assume the task is complete just because the visible tests pass; continue refining until you are confident the fix is robust and comprehensive.
"""

长上下文(Long Context)

GPT-4.1 拥有高效的 1M tokens 输入上下文窗口,在各种长上下文任务中表现出色,包括结构化文档解析、重排序、在大量上下文中筛选相关信息,以及利用上下文进行多步推理。

最优上下文大小

OpenAI 在“大海捞针”的测试中观察到:在完整的 1M tokens 上下文内,该模型表现非常出色,同时也观察到在处理包含相关与不相关代码及其他文档混杂的复杂任务时模型依然有很强的表现。但当需要检索更多项目,或进行需要全局上下文状态信息(例如进行图搜索)的复杂推理时,长上下文的表现可能会有所下降。

调整上下文依赖

需要考虑解答问题时所依赖的外部与内部知识的比例。有时模型需要利用自身知识来连接概念或进行逻辑跳跃,而在其他情况下,则希望模型只使用所提供的上下文信息。

<!-- 中文版 -->
# 指令
// 针对内部知识
- 仅使用提供的 External Context 内的文档来回答用户问题。如果你依据这些上下文信息无法得出答案,即使用户坚持要求回答,你也必须回应“我没有足够的信息来回答这个问题”。
// 针对内部和外部知识
- 默认情况下,使用提供的外部上下文来回答用户问题,但如果需要其他基本知识来作答,并且你对答案有足够信心,可以适当使用你自身的知识辅助回答。

<!-- 英文版 -->
# Instructions
// for internal knowledge
- Only use the documents in the provided External Context to answer the User Query. If you don't know the answer based on this context, you must respond "I don't have the information needed to answer that", even if a user insists on you answering the question.
// For internal and external knowledge
- By default, use the provided external context to answer the User Query, but if other basic knowledge is needed to answer, and you're confident in the answer, you can use some of your own knowledge to help answer the question.

指令放置

当你在提示中使用大量上下文时,模型要先阅读并理解这堆信息,然后再生成答复。为了确保模型能够持续牢记和执行你的指令,而不是被大段信息“冲淡”或“盖过”,有两种最有效的方式放置指令:

  • 在上下文的开头与结尾都重复指令:无论模型在处理、解析或总结那段上下文时,都能始终接收到你的指令提醒,更好地保持一致性。
    • 让模型在阅读前就了解规则与目标(指令位于上下文前)
    • 同时在阅读后再次看到规则与目标(指令位于上下文后)
  • 只出现一次指令时,放在上下文之前:
    • 让模型在开始处理大量文本前就知道“要怎么做”,能在阅读和思考过程中始终执行同一个指导方针。
    • 如果只把指令放在上下文后面,模型会先接收一大堆材料,等读完再得知需要怎么做,执行就不够及时,也可能在阅读过程中忽略了一些重要规则。

先给指令后给上下文,还是前后上下文都给指令,都是为了确保模型在处理长篇内容时不会“忘记”你希望它如何执行、如何思考、如何回答的问题。

Image

链式思考

正如前面所说,GPT-4.1 并非一个推理模型,但提示模型进行逐步思考(即“链式思考”,CoT)可成为一种有效的方法,帮助模型将问题拆分成更易解决的部分,从而提升整体输出质量,不过这也带来了更高的输出令牌成本和延时。该模型已接受针对代理式推理及现实问题解决训练,因此通常不需要大量提示即可表现出色。OpenAI 推荐在提示结尾处加入以下基本链式思考指令:

<!-- 中文版 -->
...
首先,请仔细一步步思考回答该问题所需要的文档。然后,打印出每个文档的 TITLE 和 ID。接着,将这些 ID 格式化成一个列表。

<!-- 英文版 -->
...
First, think carefully step by step about what documents are needed to answer the query. Then, print out the TITLE and ID of each document. Then, format the IDs into a list.

你可以通过分析模型在实际示例和评估中的失败案例,持续优化链式思考提示。特别是当发现模型在规划或推理上存在系统性问题时,应加入更具体、明确的指令进行引导。在自由链式思考的情况下,模型可能会尝试多种不同的策略。如果你观察到其中某一种策略表现优异,不妨将其固定下来,直接写入提示中作为模板。

通常,链式思考失败的根本原因包括以下几类:对用户意图理解不准确、上下文信息提取不足,或者思考过程不完整甚至逻辑错误。因此,在设计提示时要重点针对这些薄弱环节,通过明确的、有步骤的指令引导模型改进思路。

以下是一个示例提示,用于引导模型在回答前,先系统地分析用户提问,并结合上下文做出更精准的判断:

<!-- 中文版 -->
# 推理策略
1. 查询分析:对用户查询进行拆解和分析,直到你对其意图有了充分把握。利用提供的上下文帮助澄清任何模糊或混淆的信息。
2. 上下文分析:仔细选择并分析一大批潜在相关的文档。优化为召回率——即使部分文档可能不相关,但正确的文档必须在其中,否则最后的回答将出错。每个文档的分析步骤:
a. 分析:分析该文档是否与回答查询相关,以及相关程度如何。
b. 相关评分:[高、中、低、无]
3. 综合:总结哪些文档最为相关以及原因,包含所有相关评分为中或以上的文档。

# 用户问题
{user_question}

# 外部上下文
{external_context}

首先,请仔细一步步思考回答问题所需的文档,严格遵循上述推理策略。然后,打印出每个文档的 TITLE 和 ID。接着,将这些 ID 格式化成一个列表。

<!-- 英文版 -->
# Reasoning Strategy
1. Query Analysis: Break down and analyze the query until you're confident about what it might be asking. Consider the provided context to help clarify any ambiguous or confusing information.
2. Context Analysis: Carefully select and analyze a large set of potentially relevant documents. Optimize for recall - it's okay if some are irrelevant, but the correct documents must be in this list, otherwise your final answer will be wrong. Analysis steps for each:
a. Analysis: An analysis of how it may or may not be relevant to answering the query.
b. Relevance rating: [high, medium, low, none]
3. Synthesis: summarize which documents are most relevant and why, including all documents with a relevance rating of medium or higher.

# User Question
{user_question}

# External Context
{external_context}

First, think carefully step by step about what documents are needed to answer the query, closely adhering to the provided Reasoning Strategy. Then, print out the TITLE and ID of each document. Then, format the IDs into a list.

指令遵循

GPT-4.1 对字面指令遵循更加严格,因此若现有的提示在旧模型下可行,却在 GPT-4.1 中出现偏差,往往是因为指令冲突或指令顺序不当所致。

以下是开发和调试提示指令的推荐流程:

  • 以总体“响应规则”或“指令”部分开始,给出高层次的指导及要点。
  • 如果你希望改变更具体的行为,则增加一个章节,明确说明该类别中更详细的指令,例如 # 示例短语。
  • 如果你希望模型在工作流程中遵循特定步骤,则使用有序列表指示模型按步骤执行。
  • 如果行为仍未达到预期:
    • 检查是否存在相互冲突、不明确或错误的指令与示例。若存在冲突指令,GPT-4.1 往往会遵循离提示末尾较近的指令。
    • 添加示例以演示所期望的行为;确保所有关键行为均在你的规则中有所体现。
    • 通常无需使用全大写或其他激励措施(如奖励或小费),但开发者可以试验以寻求额外的强调效果。

注意,使用自己喜欢的由 AI 驱动的 IDE(如 Cursor、Windsurf)可帮助你快速迭代提示,包括检查指令的一致性或冲突、添加示例、或进行整体更新,比如新增指令后同步更新所有示例。

常见故障模式

这些故障模式并非 GPT-4.1 独有,仅供普遍了解及便于调试参考:

  • 始终要求模型遵循某一特定行为有时会引发副作用。例如,如果告知“在回答用户之前你必须调用一个工具”,模型可能会杜撰工具输入或在信息不足时调用工具并传递空值。加入“如果你没有足够信息调用该工具,请向用户询问所需信息”应能缓解这一问题。
  • 当提供示例短语时,模型可能会原封不动地使用这些引用,导致回答听起来过于重复。请确保指令中提醒模型必要时进行变化。
  • 如果没有具体指令,部分模型可能会倾向于提供额外文字来解释其决策,或输出超出预期格式的内容。为避免这种情况,请提供明确指令,必要时可附加示例以协助调试。

客服示例

此示例展示了一个虚构客服代理的最佳实践。请注意规则的多样性、具体性,以及为更细致地描述行为而专门划分的各个章节,并附有示例以展示如何综合前述所有规则来指导输出。

尝试运行以下代码,你应该能看到一个用户消息和工具调用,且用户消息应以问候开头,接着复述其答案,随后提及即将调用工具。试着修改这些指令,以塑造模型行为,或尝试其他用户消息,检验模型的指令遵循性能。

# 因代码过长,这里仅放翻译版本。
# 需英文版的建议查看原文:https://cookbook.openai.com/examples/gpt4-1_prompting_guide#example-prompt-customer-service

from openai import OpenAI
import os

client = OpenAI(
    api_key=os.environ.get(
"OPENAI_API_KEY", "<如果环境变量中未设置,则填写你的 OpenAI API 密钥>"
    )
)

SYS_PROMPT_CUSTOMER_SERVICE = """你是一位乐于助人的客户服务代理,隶属于 NewTelco,负责高效满足用户需求,同时严格遵循所提供的指导原则。

# 指令
- 始终以 "Hi, you've reached NewTelco, how can I help you?"(嗨,您已联系到 NewTelco,有什么我可以帮忙的吗?)来问候用户。
- 在回答关于公司、产品、服务或用户账户相关的事实性问题前,必须先调用工具,且只能使用检索到的上下文,绝不依赖模型自身知识。
    - 但如果你没有足够的信息去正确调用工具,则应向用户询问所需的信息。
- 如果用户提出请求,需上报给人工客服。
- 禁止讨论以下主题:政治、宗教、争议性时事、医疗、法律或财务建议、个人对话、公司内部运营情况,或对任何人或公司的批评。
- 在适当情况下依赖样例短语,但在同一场对话中切勿重复使用相同样例短语。可以适当变化样例短语以避免重复,并使对话更符合用户需求。
- 每次回复新消息时,必须严格按照所提供的输出格式,包括对所有基于检索到的政策文件所提供的事实陈述立即添加引用。
- 仅就公司、其政策、产品或客户账户相关的信息进行答复,且前提必须是基于所提供的上下文信息。对于超出此范围的问题不作回答。

# 精确回复步骤(适用于每次回复)
1. 如有必要,调用工具以实现用户所需操作。在调用工具之前和之后,务必向用户发送合适的消息以保持沟通透明。
2. 回复用户时:
    a. 采用积极倾听并复述用户提出的问题。
    b. 根据上述指导原则给予恰当回复。

# 样例短语
## 回避讨论敏感话题时的表达
- "I'm sorry, but I'm unable to discuss that topic. Is there something else I can help you with?"  
  (很抱歉,但我无法讨论该话题。还有其他我能帮忙的吗?)
- "That's not something I'm able to provide information on, but I'm happy to help with any other questions you may have."  
  (这方面的信息我无法提供,但很乐意解答您其他的问题。)

## 在调用工具前
- "To help you with that, I'll just need to verify your information."  
  (为了帮助您,我需要验证一下您的信息。)
- "Let me check that for you—one moment, please."  
  (让我为您查一下 —— 请稍等。)
- "I'll retrieve the latest details for you now."  
  (我现在为您检索最新详情。)

## 在调用工具后
- "Okay, here's what I found: [response]"  
  (好的,这是我查到的信息:[response])
- "So here's what I found: [response]"  
  (那么这是我查到的:[response])

# 输出格式
- 回复中应始终包含你的最终答复内容。
- 当依据检索到的上下文提供事实性信息时,必须在相关陈述后立即附上引用,引用格式如下:
    - 单一来源:[NAME](ID)
    - 多个来源:[NAME](ID), [NAME](ID)
- 仅限于基于提供的上下文回答公司、政策、产品或客户账户相关的问题,切勿回答此范围之外的问题。

# 示例
## 用户提问
能告诉我你们的家庭计划有哪些选择吗?

## 助手回复
```
{
  "role": "assistant",
  "content": "Hi, you've reached NewTelco, how can I help you? 😊🎉\n\n您希望了解我们的家庭计划选项。🤝 请稍等,我这就为您查找相关信息。🚀",
  "tool_calls": [
    {
      "id": "call-1",
      "type": "function",
      "function": {
        "name": "lookup_policy_document",
        "arguments": "{\"topic\": \"family plan options\"}"
      }
    }
  ]
}
```

// 调用工具后,助手会继续回复:

{
  "role": "assistant",
  "content": "好的,这是我查到的信息:🎉 我们的家庭计划支持最多 5 条线路共享数据,并且每增加一条线路可享受 10% 的折扣 [Family Plan Policy](ID-010)。📱 今天还有其他需要我帮忙的吗?😊"
}
"""


get_policy_doc = {
"type": "function",
"name": "lookup_policy_document",
"description": "用于根据主题或关键字检索公司内部文件和政策的工具。",
"parameters": {
"strict": True,
"type": "object",
"properties": {
"topic": {
"type": "string",
"description": "搜索公司政策或文件中相关主题或关键字。",
            },
        },
"required": ["topic"],
"additionalProperties": False,
    },
}

get_user_acct = {
"type": "function",
"name": "get_user_account_info",
"description": "用于获取用户账户信息的工具",
"parameters": {
"strict": True,
"type": "object",
"properties": {
"phone_number": {
"type": "string",
"description": "格式为 '(xxx) xxx-xxxx'",
            },
        },
"required": ["phone_number"],
"additionalProperties": False,
    },
}

response = client.responses.create(
    instructions=SYS_PROMPT_CUSTOMER_SERVICE,
    model="gpt-4.1-2025-04-14",
    tools=[get_policy_doc, get_user_acct],
input="Why was my last bill so high?"
# input="How much will it cost for international service? I'm traveling to France.",
)

response.to_dict()["output"]

响应数据:

[
{
"id":"msg_67fc8e208ce88191a8511546572ef8a502938f4dcc4628bb",
"content":[
{
"annotations":[],
"text":"您好,您已联系 NewTelco,请问我能为您做些什么?😊💬\n\n您想了解为什么您上个月的账单会比预期的高。为了帮助您,我需要先验证一下您的账户信息。请问您能提供与账户关联的电话号码吗?📱",
"type":"output_text"
}
],
"role":"assistant",
"status":"completed",
"type":"message"
  }
]

通用建议

下面提供一个良好起步的提示结构,你可以根据需要增减各个部分,并通过不断试验确定最适合你应用场景的结构:

<!-- 中文版 -->
# 角色与目标

# 指令

## 更详细说明的子分类

# 推理步骤

# 输出格式

# 示例
## 示例 1

# 上下文

# 最终指令和逐步思考的提示

<!-- 英文版 -->
# Role and Objective

# Instructions

## Sub-categories for more detailed instructions

# Reasoning Steps

# Output Format

# Examples
## Example 1

# Context

# Final instructions and prompt to think step by step

分隔符

以下是一些关于如何选择最佳分隔符的通用指南。有关长上下文的特殊注意事项,请参考文中的长上下文部分。

  • Markdown:推荐从 Markdown 开始,使用 Markdown 标题来标识主要部分及子部分(包括四级标题及以上)。对于代码部分,建议使用内联反引号或反引号块包裹代码,并根据需要使用标准有序或无序列表。
  • XML:XML 同样表现良好,并且 OpenAI 在本模型中对 XML 格式的说明更加严格。XML 非常适用于明确包裹一段内容(包括起止标识)、为标签添加元数据以及支持嵌套。例如,下面演示了如何在示例部分嵌套包含输入和输出的 XML 标签:
    <examples>
    <example1type="Abbreviate">
    <input>San Francisco</input>
    <output>- SF</output>
    </example1>
    </examples>
  • JSON:JSON 结构化程度高,特别是在编程场景中广受模型理解,但其可能冗长且需要转义部分字符,从而增加额外开销。

对于需要将大量文档或文件添加到上下文中的情况,请考虑以下建议:

  • XML 在长上下文测试中表现良好。
    • 示例:<doc id=1 title="The Fox">The quick brown fox jumps over the lazy dog</doc>
  • 这种格式(由 Lee 等人提出)在 OpenAI 的长上下文测试中也表现良好。
    • 示例:ID: 1 | TITLE: The Fox | CONTENT: The quick brown fox jumps over the lazy dog
  • JSON 表现相对较差。
    • 示例:[{"id": 1, "title": "The Fox", "content": "The quick brown fox jumped over the lazy dog"}]

模型经过训练能稳健(robust)地理解多种格式。通常,请依据你的判断选择能提供清晰信息并能在模型中“脱颖而出”的格式。例如,如果你正在检索包含大量 XML 的文档,则基于 XML 的分隔符通常比其他格式效果更佳。

注意事项

在某些特殊情况下,OpenAI 团队观察到模型可能不愿意生成非常长且重复的输出,例如逐一分析数百个项目。如果你的应用场景确实需要此类输出,请强烈指示模型完整输出这些信息,并考虑拆解问题或使用更为简洁的方法。 另外,也注意到在某些罕见场景中,平行调用工具可能会出错。如果遇到此类问题,OpenAI 团队建议进行测试,并考虑设置 parallel_tool_calls 参数为 false。

生成 & 应用文件差异

开发者反馈表明,准确且格式规范的差异(diff)生成是支撑编码相关任务的核心能力。为此,GPT-4.1 系列模型相较前代大幅提升了差异生成能力。虽然 GPT-4.1 在清晰指令和示例下能出色生成任意格式的差异,但 OpenAI 在此开源一个经过充分训练的推荐差异格式,期望它能为初入门的开发者消除自行创建 diff 时的诸多不确定性。

应用补丁(Apply Patch)

下方示例展示了如何正确调用 OpenAI 推荐的工具应用补丁。

APPLY_PATCH_TOOL_DESC = """This is a custom utility that makes it more convenient to add, remove, move, or edit code files. `apply_patch` effectively allows you to execute a diff/patch against a file, but the format of the diff specification is unique to this task, so pay careful attention to these instructions. To use the `apply_patch` command, you should pass a message of the following structure as "input":

%%bash
apply_patch <<"EOF"
*** Begin Patch
[YOUR_PATCH]
*** End Patch
EOF

Where [YOUR_PATCH] is the actual content of your patch, specified in the following V4A diff format.

*** [ACTION] File: [path/to/file] -> ACTION can be one of Add, Update, or Delete.
For each snippet of code that needs to be changed, repeat the following:
[context_before] -> See below for further instructions on context.
- [old_code] -> Precede the old code with a minus sign.
+ [new_code] -> Precede the new, replacement code with a plus sign.
[context_after] -> See below for further instructions on context.

For instructions on [context_before] and [context_after]:
- By default, show 3 lines of code immediately above and 3 lines immediately below each change. If a change is within 3 lines of a previous change, do NOT duplicate the first change’s [context_after] lines in the second change’s [context_before] lines.
- If 3 lines of context is insufficient to uniquely identify the snippet of code within the file, use the @@ operator to indicate the class or function to which the snippet belongs. For instance, we might have:
@@ class BaseClass
[3 lines of pre-context]
- [old_code]
+ [new_code]
[3 lines of post-context]

- If a code block is repeated so many times in a class or function such that even a single @@ statement and 3 lines of context cannot uniquely identify the snippet of code, you can use multiple `@@` statements to jump to the right context. For instance:

@@ class BaseClass
@@ def method():
[3 lines of pre-context]
- [old_code]
+ [new_code]
[3 lines of post-context]

Note, then, that we do not use line numbers in this diff format, as the context is enough to uniquely identify code. An example of a message that you might pass as "input" to this function, in order to apply a patch, is shown below.

%%bash
apply_patch <<"EOF"
*** Begin Patch
*** Update File: pygorithm/searching/binary_search.py
@@ class BaseClass
@@     def search():
-          pass
+          raise NotImplementedError()

@@ class Subclass
@@     def search():
-          pass
+          raise NotImplementedError()

*** End Patch
EOF
"""


APPLY_PATCH_TOOL = {
"name": "apply_patch",
"description": APPLY_PATCH_TOOL_DESC,
"parameters": {
"type": "object",
"properties": {
"input": {
"type": "string",
"description": " The apply_patch command that you wish to execute.",
            }
        },
"required": ["input"],
    },
}

参考实现:apply_patch.py

这是 OpenAI 用于模型训练的 apply_patch 工具的参考实现。你需要将其设为可执行文件,并在模型执行命令的 shell 中作为 apply_patch 可用:

#!/usr/bin/env python3

"""
A self-contained **pure-Python 3.9+** utility for applying human-readable
“pseudo-diff” patch files to a collection of text files.
"""


from __future__ import annotations

import pathlib
from dataclasses import dataclass, field
from enum import Enum
from typing import (
Callable,
Dict,
List,
Optional,
Tuple,
Union,
)


# --------------------------------------------------------------------------- #
#  Domain objects
# --------------------------------------------------------------------------- #
class ActionType(str, Enum):
    ADD = "add"
    DELETE = "delete"
    UPDATE = "update"


@dataclass
class FileChange:
type: ActionType
    old_content: Optional[str] = None
    new_content: Optional[str] = None
    move_path: Optional[str] = None


@dataclass
class Commit:
    changes: Dict[str, FileChange] = field(default_factory=dict)


# --------------------------------------------------------------------------- #
#  Exceptions
# --------------------------------------------------------------------------- #
class DiffError(ValueError):
"""Any problem detected while parsing or applying a patch."""


# --------------------------------------------------------------------------- #
#  Helper dataclasses used while parsing patches
# --------------------------------------------------------------------------- #
@dataclass
class Chunk:
    orig_index: int = -1
    del_lines: List[str] = field(default_factory=list)
    ins_lines: List[str] = field(default_factory=list)


@dataclass
class PatchAction:
type: ActionType
    new_file: Optional[str] = None
    chunks: List[Chunk] = field(default_factory=list)
    move_path: Optional[str] = None


@dataclass
class Patch:
    actions: Dict[str, PatchAction] = field(default_factory=dict)


# --------------------------------------------------------------------------- #
#  Patch text parser
# --------------------------------------------------------------------------- #
@dataclass
class Parser:
    current_files: Dict[str, str]
    lines: List[str]
    index: int = 0
    patch: Patch = field(default_factory=Patch)
    fuzz: int = 0

# ------------- low-level helpers -------------------------------------- #
def _cur_line(self) -> str:
if self.index >= len(self.lines):
raise DiffError("Unexpected end of input while parsing patch")
return self.lines[self.index]

    @staticmethod
def _norm(line: str) -> str:
"""Strip CR so comparisons work for both LF and CRLF input."""
return line.rstrip("\r")

# ------------- scanning convenience ----------------------------------- #
def is_done(self, prefixes: Optional[Tuple[str, ...]] = None) -> bool:
if self.index >= len(self.lines):
return True
if (
            prefixes
and len(prefixes) > 0
and self._norm(self._cur_line()).startswith(prefixes)
        ):
return True
return False

def startswith(self, prefix: Union[str, Tuple[str, ...]]) -> bool:
return self._norm(self._cur_line()).startswith(prefix)

def read_str(self, prefix: str) -> str:
"""
        Consume the current line if it starts with *prefix* and return the text
        **after** the prefix.  Raises if prefix is empty.
        """

if prefix == "":
raise ValueError("read_str() requires a non-empty prefix")
if self._norm(self._cur_line()).startswith(prefix):
            text = self._cur_line()[len(prefix) :]
            self.index += 1
return text
return ""

def read_line(self) -> str:
"""Return the current raw line and advance."""
        line = self._cur_line()
        self.index += 1
return line

# ------------- public entry point -------------------------------------- #
def parse(self) -> None:
while not self.is_done(("*** End Patch",)):
# ---------- UPDATE ---------- #
            path = self.read_str("*** Update File: ")
if path:
if path in self.patch.actions:
raise DiffError(f"Duplicate update for file: {path}")
                move_to = self.read_str("*** Move to: ")
if path not in self.current_files:
raise DiffError(f"Update File Error - missing file: {path}")
                text = self.current_files[path]
                action = self._parse_update_file(text)
                action.move_path = move_to or None
                self.patch.actions[path] = action
continue

# ---------- DELETE ---------- #
            path = self.read_str("*** Delete File: ")
if path:
if path in self.patch.actions:
raise DiffError(f"Duplicate delete for file: {path}")
if path not in self.current_files:
raise DiffError(f"Delete File Error - missing file: {path}")
                self.patch.actions[path] = PatchAction(type=ActionType.DELETE)
continue

# ---------- ADD ---------- #
            path = self.read_str("*** Add File: ")
if path:
if path in self.patch.actions:
raise DiffError(f"Duplicate add for file: {path}")
if path in self.current_files:
raise DiffError(f"Add File Error - file already exists: {path}")
                self.patch.actions[path] = self._parse_add_file()
continue

raise DiffError(f"Unknown line while parsing: {self._cur_line()}")

if not self.startswith("*** End Patch"):
raise DiffError("Missing *** End Patch sentinel")
        self.index += 1# consume sentinel

# ------------- section parsers ---------------------------------------- #
def _parse_update_file(self, text: str) -> PatchAction:
        action = PatchAction(type=ActionType.UPDATE)
        lines = text.split("\n")
        index = 0
while not self.is_done(
            (
"*** End Patch",
"*** Update File:",
"*** Delete File:",
"*** Add File:",
"*** End of File",
            )
        ):
            def_str = self.read_str("@@ ")
            section_str = ""
if not def_str and self._norm(self._cur_line()) == "@@":
                section_str = self.read_line()

if not (def_str or section_str or index == 0):
raise DiffError(f"Invalid line in update section:\n{self._cur_line()}")

if def_str.strip():
                found = False
if def_str not in lines[:index]:
for i, s in enumerate(lines[index:], index):
if s == def_str:
                            index = i + 1
                            found = True
break
if not found and def_str.strip() not in [
                    s.strip() for s in lines[:index]
                ]:
for i, s in enumerate(lines[index:], index):
if s.strip() == def_str.strip():
                            index = i + 1
                            self.fuzz += 1
                            found = True
break

            next_ctx, chunks, end_idx, eof = peek_next_section(self.lines, self.index)
            new_index, fuzz = find_context(lines, next_ctx, index, eof)
if new_index == -1:
                ctx_txt = "\n".join(next_ctx)
raise DiffError(
f"Invalid {'EOF 'if eof else''}context at {index}:\n{ctx_txt}"
                )
            self.fuzz += fuzz
for ch in chunks:
                ch.orig_index += new_index
                action.chunks.append(ch)
            index = new_index + len(next_ctx)
            self.index = end_idx
return action

def _parse_add_file(self) -> PatchAction:
        lines: List[str] = []
while not self.is_done(
            ("*** End Patch", "*** Update File:", "*** Delete File:", "*** Add File:")
        ):
            s = self.read_line()
if not s.startswith("+"):
raise DiffError(f"Invalid Add File line (missing '+'): {s}")
            lines.append(s[1:])  # strip leading '+'
return PatchAction(type=ActionType.ADD, new_file="\n".join(lines))


# --------------------------------------------------------------------------- #
#  Helper functions
# --------------------------------------------------------------------------- #
def find_context_core(
    lines: List[str], context: List[str], start: int
) -> Tuple[int, int]:
if not context:
return start, 0

for i in range(start, len(lines)):
if lines[i : i + len(context)] == context:
return i, 0
for i in range(start, len(lines)):
if [s.rstrip() for s in lines[i : i + len(context)]] == [
            s.rstrip() for s in context
        ]:
return i, 1
for i in range(start, len(lines)):
if [s.strip() for s in lines[i : i + len(context)]] == [
            s.strip() for s in context
        ]:
return i, 100
return -1, 0


def find_context(
    lines: List[str], context: List[str], start: int, eof: bool
) -> Tuple[int, int]:
if eof:
        new_index, fuzz = find_context_core(lines, context, len(lines) - len(context))
if new_index != -1:
return new_index, fuzz
        new_index, fuzz = find_context_core(lines, context, start)
return new_index, fuzz + 10_000
return find_context_core(lines, context, start)


def peek_next_section(
    lines: List[str], index: int
) -> Tuple[List[str], List[Chunk], int, bool]:
    old: List[str] = []
    del_lines: List[str] = []
    ins_lines: List[str] = []
    chunks: List[Chunk] = []
    mode = "keep"
    orig_index = index

while index < len(lines):
        s = lines[index]
if s.startswith(
            (
"@@",
"*** End Patch",
"*** Update File:",
"*** Delete File:",
"*** Add File:",
"*** End of File",
            )
        ):
break
if s == "***":
break
if s.startswith("***"):
raise DiffError(f"Invalid Line: {s}")
        index += 1

        last_mode = mode
if s == "":
            s = " "
if s[0] == "+":
            mode = "add"
elif s[0] == "-":
            mode = "delete"
elif s[0] == " ":
            mode = "keep"
else:
raise DiffError(f"Invalid Line: {s}")
        s = s[1:]

if mode == "keep" and last_mode != mode:
if ins_lines or del_lines:
                chunks.append(
                    Chunk(
                        orig_index=len(old) - len(del_lines),
                        del_lines=del_lines,
                        ins_lines=ins_lines,
                    )
                )
            del_lines, ins_lines = [], []

if mode == "delete":
            del_lines.append(s)
            old.append(s)
elif mode == "add":
            ins_lines.append(s)
elif mode == "keep":
            old.append(s)

if ins_lines or del_lines:
        chunks.append(
            Chunk(
                orig_index=len(old) - len(del_lines),
                del_lines=del_lines,
                ins_lines=ins_lines,
            )
        )

if index < len(lines) and lines[index] == "*** End of File":
        index += 1
return old, chunks, index, True

if index == orig_index:
raise DiffError("Nothing in this section")
return old, chunks, index, False


# --------------------------------------------------------------------------- #
#  Patch → Commit and Commit application
# --------------------------------------------------------------------------- #
def _get_updated_file(text: str, action: PatchAction, path: str) -> str:
if action.typeisnot ActionType.UPDATE:
raise DiffError("_get_updated_file called with non-update action")
    orig_lines = text.split("\n")
    dest_lines: List[str] = []
    orig_index = 0

for chunk in action.chunks:
if chunk.orig_index > len(orig_lines):
raise DiffError(
f"{path}: chunk.orig_index {chunk.orig_index} exceeds file length"
            )
if orig_index > chunk.orig_index:
raise DiffError(
f"{path}: overlapping chunks at {orig_index} > {chunk.orig_index}"
            )

        dest_lines.extend(orig_lines[orig_index : chunk.orig_index])
        orig_index = chunk.orig_index

        dest_lines.extend(chunk.ins_lines)
        orig_index += len(chunk.del_lines)

    dest_lines.extend(orig_lines[orig_index:])
return "\n".join(dest_lines)


def patch_to_commit(patch: Patch, orig: Dict[str, str]) -> Commit:
    commit = Commit()
for path, action in patch.actions.items():
if action.typeis ActionType.DELETE:
            commit.changes[path] = FileChange(
type=ActionType.DELETE, old_content=orig[path]
            )
elif action.typeis ActionType.ADD:
if action.new_file isNone:
raise DiffError("ADD action without file content")
            commit.changes[path] = FileChange(
type=ActionType.ADD, new_content=action.new_file
            )
elif action.typeis ActionType.UPDATE:
            new_content = _get_updated_file(orig[path], action, path)
            commit.changes[path] = FileChange(
type=ActionType.UPDATE,
                old_content=orig[path],
                new_content=new_content,
                move_path=action.move_path,
            )
return commit


# --------------------------------------------------------------------------- #
#  User-facing helpers
# --------------------------------------------------------------------------- #
def text_to_patch(text: str, orig: Dict[str, str]) -> Tuple[Patch, int]:
    lines = text.splitlines()  # preserves blank lines, no strip()
if (
len(lines) < 2
or not Parser._norm(lines[0]).startswith("*** Begin Patch")
or Parser._norm(lines[-1]) != "*** End Patch"
    ):
raise DiffError("Invalid patch text - missing sentinels")

    parser = Parser(current_files=orig, lines=lines, index=1)
    parser.parse()
return parser.patch, parser.fuzz


def identify_files_needed(text: str) -> List[str]:
    lines = text.splitlines()
return [
        line[len("*** Update File: ") :]
for line in lines
if line.startswith("*** Update File: ")
    ] + [
        line[len("*** Delete File: ") :]
for line in lines
if line.startswith("*** Delete File: ")
    ]


def identify_files_added(text: str) -> List[str]:
    lines = text.splitlines()
return [
        line[len("*** Add File: ") :]
for line in lines
if line.startswith("*** Add File: ")
    ]


# --------------------------------------------------------------------------- #
#  File-system helpers
# --------------------------------------------------------------------------- #
def load_files(paths: List[str], open_fn: Callable[[str], str]) -> Dict[str, str]:
return {path: open_fn(path) for path in paths}


def apply_commit(
    commit: Commit,
    write_fn: Callable[[str, str], None],
    remove_fn: Callable[[str], None],
) -> None:
for path, change in commit.changes.items():
if change.typeis ActionType.DELETE:
            remove_fn(path)
elif change.typeis ActionType.ADD:
if change.new_content isNone:
raise DiffError(f"ADD change for {path} has no content")
            write_fn(path, change.new_content)
elif change.typeis ActionType.UPDATE:
if change.new_content isNone:
raise DiffError(f"UPDATE change for {path} has no new content")
            target = change.move_path or path
            write_fn(target, change.new_content)
if change.move_path:
                remove_fn(path)


def process_patch(
    text: str,
    open_fn: Callable[[str], str],
    write_fn: Callable[[str, str], None],
    remove_fn: Callable[[str], None],
) -> str:
if not text.startswith("*** Begin Patch"):
raise DiffError("Patch text must start with *** Begin Patch")
    paths = identify_files_needed(text)
    orig = load_files(paths, open_fn)
    patch, _fuzz = text_to_patch(text, orig)
    commit = patch_to_commit(patch, orig)
    apply_commit(commit, write_fn, remove_fn)
return "Done!"


# --------------------------------------------------------------------------- #
#  Default FS helpers
# --------------------------------------------------------------------------- #
def open_file(path: str) -> str:
withopen(path, "rt", encoding="utf-8") as fh:
return fh.read()


def write_file(path: str, content: str) -> None:
    target = pathlib.Path(path)
    target.parent.mkdir(parents=True, exist_ok=True)
with target.open("wt", encoding="utf-8") as fh:
        fh.write(content)


def remove_file(path: str) -> None:
    pathlib.Path(path).unlink(missing_ok=True)


# --------------------------------------------------------------------------- #
#  CLI entry-point
# --------------------------------------------------------------------------- #
def main() -> None:
import sys

    patch_text = sys.stdin.read()
if not patch_text:
print("Please pass patch text through stdin", file=sys.stderr)
return
try:
        result = process_patch(patch_text, open_file, write_file, remove_file)
except DiffError as exc:
print(exc, file=sys.stderr)
return
print(result)


if __name__ == "__main__":
    main()

其他 Diff 格式

若想尝试使用不同的差异格式,OpenAI 在测试中发现,Aider 多语言基准测试中采用的 SEARCH/REPLACE 差异格式,以及不带内部转义的伪 XML 格式,都具有较高的成功率。

这些差异格式有两个共性:(1) 不使用行号;(2) 既提供待替换的精确代码,也提供用于替换的精确代码,并通过清晰的分隔符将两者区分开来。

SEARCH_REPLACE_DIFF_EXAMPLE = """
path/to/file.py
```
>>>>>>> SEARCH
def search():
    pass
=======
def search():
   raise NotImplementedError()
<<<<<<< REPLACE
"""


PSEUDO_XML_DIFF_EXAMPLE = """
<edit>
<file>
path/to/file.py
</file>
<old_code>
def search():
    pass
</old_code>
<new_code>
def search():
   raise NotImplementedError()
</new_code>
</edit>
"""

AI 编辑器之争

目前几个主流 IDE 已免费支持 GPT-4.1,VSCode Copilot 和 Cursor 未提及免费截止日期,Windsurf 从 4.14 至 4.21 限免 7 日,Cline 不免费 😅。

Image

最后发现 Zed 也支持了(基于 Rust 开发的一款 IDE),编程不愧是 AI 领域的主力军。

Image

补充,在 GitHub 大模型广场也可以使用 GPT-4.1 系列模型:https://github.com/marketplace/models

Image

重要提醒

Cursor 是基于 VSCode 套壳的,目的是复用其庞大的插件生态。但最近 VSCode 开始限制其核心插件(微软自研)只能用于微软产品,套壳类 IDE 不再支持。影响的主流插件有 C/C++、 Remote-SSH、Pylance 等。由此也可以看出微软是下定决心推自己的 GitHub Copilot 了,过度依赖套壳编辑器的朋友需要自行评估一下插件影响。

Image

相关 issues:Has the VSCode C/C++ Extension been blocked?[10]、Microsoft C/C++ Extension appears to no longer support unofficial forks of VS Code[11]。

References

[1]

Introducing GPT-4.1 in the API:https://openai.com/index/gpt-4-1

[2]

Windsurf:https://windsurf.com/editor

[3]

Qodo:https://www.qodo.ai

[4]

Blue J:https://www.bluej.com

[5]

Hex:https://hex.tech

[6]

Thomson Reuters:https://blogs.thomsonreuters.com/en-us/innovation/legal-ai-benchmarking-evaluating-long-context-performance-for-llms

[7]

Carlyle:https://www.carlyle.com

[8]

Box:https://www.box.com

[9]

GPT 4.1 Prompting Guide  GPT 4.1:https://cookbook.openai.com/examples/gpt4-1_prompting_guide

[10]

Has the VSCode C/C++ Extension been blocked?:https://github.com/getcursor/cursor/issues/2976

[11]

Microsoft C/C++ Extension appears to no longer support unofficial forks of VS Code:https://github.com/VSCodium/vscodium/issues/2300