claude团队揭秘:ai大脑不用英文也不用中文思考,而是靠“思维语言”。|这证明了英语学习/教育失败的根本原因?
llm用什么语言“思考”?中文?英文?
都不是。
llms的思考,使用的不是中文或英文这样的自然语言,而是一种超越自然语言的“思维语言”。anthropic的最新研究,用实验方式首次证明了这一点,这是理解llm内部黑箱的一个巨大突破。
在llm内部,不同语言共享同一个概念空间。不论是中文、英文还是法语,这些自然语言都只是这个更深层次“思维语言”的表面形式而已。
举个例子,你用英语教llm一个概念,它能用中文流利地表达出来;你用中文教它,它一样能在英文里准确使用。语言不同,但背后的概念是一样的。
anthropic的实验很有意思:面对同一个问题,不论你用英语问(the opposite of "small" is),是中文问(“小”的反义词是),还是用法语问(le contraire de "petit" est),llm实际上都在用自己跨语言共享的特征来思考,在同一个共享的概念空间里思考,然后再把结果翻译为对应的英文、中文或法语输出给你。
对人类学习的启示
如果你同时理解llm和人脑的运作原理,不难想到以下几点:
人脑也不是通过中文或英文这样的自然语言思考的,而是一种更为底层、更为通用的“思维语言”;
中文和英文只是表面差异,真正重要的是思维语言本身的运动(俗称”思考“);
概念、事实性知识砖块和心理模型等心理表征,这些才是思维语言真正的词汇。概念语言先于自然语言。如果你掌握的思维词汇太少,不论你说哪种语言,都没法深度思考。
另一个重要的启示:我们学英语,到底是学什么?为什么英语教育经常很失败?为什么英语学习经常很低效?理解了这些,这些问题的答案不言自明(关键要学知识,要海量阅读,要积累双语的知识砖块,而非贫乏模式应试刷题教培内卷)。
昨天读到anthropic的这篇文章,我心情激动。毕竟早在一年前,我就在twitter和公众号文章里多次表达这样的观点(部分文章的链接我放在文章结尾了)。
于是,双语对照版本文章奉上,祝你阅读愉快~
追踪大语言模型的内部思考过程
title: tracing the thoughts of a large language model
author: anthropic
date: 03-27-2025
Language models like Claude aren't programmed directly by humans—instead, they’re trained on large amounts of data. During that training process, they learn their own strategies to solve problems. These strategies are encoded in the billions of computations a model performs for every word it writes. They arrive inscrutable to us, the model’s developers. This means that we don’t understand how models do most of the things they do.
像 Claude 这样的语言模型并不是由人类直接编程控制的——它们是在海量数据上训练出来的。在训练过程中,它们自主学会了解决问题的方法。这些方法被编码在模型为输出每个词而进行的数十亿次计算中。这些计算对我们(模型的开发者)而言完全是不可解读的。这也就意味着,我们并不理解模型完成大多数任务的内部原理。
Knowing how models like Claude think would allow us to have a better understanding of their abilities, as well as help us ensure that they’re doing what we intend them to. For example:
如果能知道像 Claude 这样的模型是如何“思考”的,我们就能更好地理解它们的能力,并确保它们确实按照我们的意图在行动。例如:
Claude can speak dozens of languages. What language, if any, is it using "in its head"? Claude 能使用数十种语言。那么,它在“脑海中”使用的语言究竟是什么呢? Claude writes text one word at a time. Is it only focusing on predicting the next word or does it ever plan ahead? Claude 每次只写出一个词。它只是专注于预测下一个词,还是会提前进行一些规划? Claude can write out its reasoning step-by-step. Does this explanation represent the actual steps it took to get to an answer, or is it sometimes fabricating a plausible argument for a foregone conclusion? Claude 可以逐步写出自己的推理过程。但这种解释是否代表了它实际得到答案所经历的步骤?还是说,它有时只是在为一个早已确定的结论编造看似合理的理由?
We take inspiration from the field of neuroscience, which has long studied the messy insides of thinking organisms, and try to build a kind of AI microscope that will let us identify patterns of activity and flows of information. There are limits to what you can learn just by talking to an AI model—after all, humans (even neuroscientists) don't know all the details of how our own brains work. So we look inside.
我们从神经科学领域汲取了灵感。神经科学长期以来一直致力于研究思考中的有机体那凌乱的内部机制。同样地,我们尝试打造一种 AI 的“显微镜”,使我们能够识别模型内部的活动模式和信息流动。仅靠与 AI 模型对话,你能了解的东西十分有限——毕竟,人类(即使是神经科学家)也并不清楚我们自己大脑运作的所有细节。所以,我们选择向模型内部探查。
Today, we're sharing two new papers that represent progress on the development of the "microscope", and the application of it to see new "AI biology". In the first paper, we extend our prior work locating interpretable concepts ("features") inside a model to link those concepts together into computational "circuits", revealing parts of the pathway that transforms the words that go into Claude into the words that come out. In the second, we look inside Claude 3.5 Haiku, performing deep studies of simple tasks representative of ten crucial model behaviors, including the three described above. Our method sheds light on a part of what happens when Claude responds to these prompts, which is enough to see solid evidence that:
今天,我们发布了两篇新论文,展示了在建立这种“显微镜”以及用它来探索新的“AI生物学”方面取得的进展。在第一篇论文中,我们扩展了我们之前的工作——在模型内部定位可解释概念(“特征”)——并将这些概念连接成计算“电路”。这揭示了部分把输入 Claude 的单词转化为输出单词的通路。在第二篇论文中,我们深入观察了 Claude 3.5 Haiku,对十种关键模型行为中具有代表性的一些简单任务进行了深入研究,其中包括上面提到的三个行为。我们的方法让我们洞察 Claude 在响应这些提示时部分内部发生的过程,并据此找到了确凿证据,证明:
Claude sometimes thinks in a conceptual space that is shared between languages, suggesting it has a kind of universal “language of thought.” We show this by translating simple sentences into multiple languages and tracing the overlap in how Claude processes them. Claude 有时会在多种语言共享的概念空间中进行思考,这表明它具备某种通用的“思维语言”。我们通过将简单句子翻译成多种语言并追踪 Claude 处理这些句子时内部激活的重叠部分来展示这一点。 Claude will plan what it will say many words ahead, and write to get to that destination. We show this in the realm of poetry, where it thinks of possible rhyming words in advance and writes the next line to get there. This is powerful evidence that even though models are trained to output one word at a time, they may think on much longer horizons to do so. Claude 会提前规划好它未来很多个词的输出,并为了达到预想的结尾而书写过程中的内容。我们在诗歌领域证明了这一点:Claude 会提前想好可能押韵的词,并写下一行诗以导向那个押韵词。这一发现有力地证明,即使模型的训练目标是一次输出一个词,它在生成过程中可能进行着更长远的思考。 Claude, on occasion, will give a plausible-sounding argument designed to agree with the user rather than to follow logical steps. We show this by asking it for help on a hard math problem while giving it an incorrect hint. We are able to “catch it in the act” as it makes up its fake reasoning, providing a proof of concept that our tools can be useful for flagging concerning mechanisms in models. Claude 有时会给出听起来很有道理的论证,其目的在于迎合用户,而非遵循严谨的逻辑步骤。我们通过这样一个方法揭示了这一点:在一道很难的数学题上向它求助,同时给出一个错误的提示。结果,我们能够在 Claude 编造虚假推理的过程中“当场抓住”它。这一概念验证表明,我们的工具可以用于标记模型中令人担忧的内部机制。
We were often surprised by what we saw in the model: In the poetry case study, we had set out to show that the model didn't plan ahead, and found instead that it did. In a study of hallucinations, we found the counter-intuitive result that Claude's default behavior is to decline to speculate when asked a question, and it only answers questions when something inhibits this default reluctance. In a response to an example jailbreak, we found that the model recognized it had been asked for dangerous information well before it was able to gracefully bring the conversation back around. While the problems we study can analyzed with other methods, the general "build a microscope" approach lets us learn many things we wouldn't have guessed going in, which will be increasingly important as models grow more sophisticated.
模型内部呈现出的景象常常出乎我们的意料:在诗歌案例研究中,我们原本打算证明模型不会提前规划,结果却发现它确实会。在对幻觉现象的研究中,我们得出了一个与直觉相反的结果:Claude 的默认行为是在被提问时选择拒绝猜测,只有当某种机制抑制了它这种默认的迟疑后,它才会回答问题。而在应对一个“越狱”提示的实验中,我们发现模型在巧妙地将话题带回正轨之前,就已经很早识别出了用户请求的是危险信息。尽管我们研究的问题也可以用其他方法来分析,但这种总体的“建立显微镜”方法让我们学到了许多事前无法预料的东西。随着模型变得越来越复杂,这种方法的重要性也将日益凸显。
These findings aren’t just scientifically interesting—they represent significant progress towards our goal of understanding AI systems and making sure they’re reliable. We also hope they prove useful to other groups, and potentially, in other domains: for example, interpretability techniques have found use in fields such as medical imaging and genomics, as dissecting the internal mechanisms of models trained for scientific applications can reveal new insight about the science.
这些发现不仅在科学上引人入胜——它们也标志着朝着我们“理解 AI 系统并确保其可靠”这一目标迈出了重要一步。我们也希望这些发现能对其他团队有所助益,并可能拓展到其他领域:例如,在医学影像和基因组学等领域,解释性技术已经展现了用武之地,因为剖析用于科学应用的模型内部机制可以揭示关于那些科学问题的新见解。
At the same time, we recognize the limitations of our current approach. Even on short, simple prompts, our method only captures a fraction of the total computation performed by Claude, and the mechanisms we do see may have some artifacts based on our tools which don't reflect what is going on in the underlying model. It currently takes a few hours of human effort to understand the circuits we see, even on prompts with only tens of words. To scale to the thousands of words supporting the complex thinking chains used by modern models, we will need to improve both the method and (perhaps with AI assistance) how we make sense of what we see with it.
同时,我们也清楚当前方法的局限性。即使在简短、简单的提示上,我们的方法也只能捕获 Claude 总计算量的一小部分。而且我们观测到的机制可能带有我们工具所引入的伪影,未必真实反映底层模型中发生的过程。目前,即使对于只有几十个词的提示,理解我们看到的电路也需要人类耗费数小时的精力。要将研究扩展到现代模型使用的上千词长、复杂的思维链,我们需要改进方法本身,并改进我们理解这些观察结果的方式(可能需要 AI 的协助)。
As AI systems are rapidly becoming more capable and are deployed in increasingly important contexts, Anthropic is investing in a portfolio of approaches including realtime monitoring, model character improvements, and the science of alignment. Interpretability research like this is one of the highest-risk, highest-reward investments, a significant scientific challenge with the potential to provide a unique tool for ensuring that AI is transparent. Transparency into the model’s mechanisms allows us to check whether it’s aligned with human values—and whether it’s worthy of our trust.
随着 AI 系统能力快速提升并被部署到越来越重要的场景中,Anthropic 正在多方面投入,包括实时监控、模型特性改进以及对齐性的科学。像这样的可解释性研究是高风险高回报的投入之一——这是一项重大的科学挑战,但有可能为确保 AI 透明可控提供独特的工具。对模型内部机制的透明了解使我们能够检验模型是否与人类的价值观对齐——以及它是否值得我们的信任。
For full details, please read the transformer papers. Below, we invite you on a short tour of some of the most striking "AI biology" findings from our investigations.
想了解完整细节,请参阅上述transformer相关论文。下面,我们邀请你一起对我们的研究中一些最引人注目的“AI生物学”发现进行一次简短巡礼。
A tour of AI biology|AI 生物学之旅
How is Claude multilingual?
Claude 的多语言能力是如何运作的?
Claude speaks dozens of languages fluently—from English and French to Chinese and Tagalog. How does this multilingual ability work? Is there a separate "French Claude" and "Chinese Claude" running in parallel, responding to requests in their own language? Or is there some cross-lingual core inside?
Claude 可以流利地使用数十种语言——从英语、法语到中文和塔加洛语。它的多语言能力究竟是如何运作的?难道在它体内有一个独立的“法语 Claude”和“中文 Claude”并行运行,分别用各自的语言回应请求吗?还是说它内部存在某种跨语言的核心?
Recent research on smaller models has shown hints of shared grammatical mechanisms across languages. We investigate this by asking Claude for the "opposite of small" across different languages, and find that the same core features for the concepts of smallness and oppositeness activate, and trigger a concept of largeness, which gets translated out into the language of the question. We find that the shared circuitry increases with model scale, with Claude 3.5 Haiku sharing more than twice the proportion of its features between languages as compared to a smaller model.
近期针对较小模型的研究已经显露出跨语言的共享的语法机制的蛛丝马迹。我们通过用不同语言询问 Claude “small 的反义词是什么”(意为“小”的反义词)来研究这一点,结果发现,对于“小”和“相反”这些概念,Claude 内部激活了相同的核心特征,并触发了“大”这一概念,随后再将其翻译成提问所用的语言。我们发现,随着模型规模的增大,这种共享的电路比例也在上升。Claude 3.5 Haiku 在不同语言间共享的特征比例是一个较小模型的两倍以上。
This provides additional evidence for a kind of conceptual universality—a shared abstract space where meanings exist and where thinking can happen before being translated into specific languages. More practically, it suggests Claude can learn something in one language and apply that knowledge when speaking another. Studying how the model shares what it knows across contexts is important to understanding its most advanced reasoning capabilities, which generalize across many domains.
这些发现为某种概念层面的通用性提供了更多证据——也就是模型存在一个共享的抽象空间,在被翻译成具体语言之前,语义就在其中存在并进行思考。更实际地说,这意味着 Claude 可以用一种语言学到知识,并在使用另一种语言时应用那些知识。研究模型如何在不同语境之间共享它所知道的东西,对于理解它最先进的、能够泛化到多个领域的推理能力非常重要。
Does Claude plan its rhymes?
Claude 会预先规划押韵吗?
How does Claude write rhyming poetry? Consider this ditty:
Claude 是如何写出押韵的诗歌的?请看这首打油诗:
He saw a carrot and had to grab it,
His hunger was like a starving rabbit
他看见一根胡萝卜就忍不住抓起,
他的饥饿就像一只饿坏了的兔子
To write the second line, the model had to satisfy two constraints at the same time: the need to rhyme (with "grab it"), and the need to make sense (why did he grab the carrot?). Our guess was that Claude was writing word-by-word without much forethought until the end of the line, where it would make sure to pick a word that rhymes. We therefore expected to see a circuit with parallel paths, one for ensuring the final word made sense, and one for ensuring it rhymes.
在写出第二行时,模型需要同时满足两个约束:需要押韵(与“grab it”押韵),也需要内容合理(交代他为什么抓起胡萝卜)。我们原先的猜想是,Claude 会逐词地即兴创作,在行尾才稍加思考以确保最后一个词押韵。因此,我们预期会看到一条具有并行路径的“电路”:一条路径确保最后一个词意义通顺,另一条路径确保它押韵。
Instead, we found that Claude plans ahead. Before starting the second line, it began "thinking" of potential on-topic words that would rhyme with "grab it". Then, with these plans in mind, it writes a line to end with the planned word.
然而,结果我们发现 Claude 实际上是提前规划的。在开始创作第二行之前,它就已经开始“考虑”与“grab it”押韵且切题的潜在用词。然后,怀着这些提前定好的方案,它写出了以那个预定单词结尾的一整行诗句。
To understand how this planning mechanism works in practice, we conducted an experiment inspired by how neuroscientists study brain function, by pinpointing and altering neural activity in specific parts of the brain (for example using electrical or magnetic currents). Here, we modified the part of Claude’s internal state that represented the "rabbit" concept. When we subtract out the "rabbit" part, and have Claude continue the line, it writes a new one ending in "habit", another sensible completion. We can also inject the concept of "green" at that point, causing Claude to write a sensible (but no-longer rhyming) line which ends in "green". This demonstrates both planning ability and adaptive flexibility—Claude can modify its approach when the intended outcome changes.
为了理解这种规划机制在实践中如何工作,我们借鉴了神经科学家研究大脑功能的方法:精确定位并改变大脑特定部分的神经活动(例如使用电流或磁场刺激)。在这里,我们修改了 Claude 内部状态中表示“rabbit”(兔子)概念的那部分。当我们去除“rabbit”概念的影响并让 Claude 继续写下去时,它写出了一个以“habit”结尾的新句子——这是另一个合情合理的结尾。我们也可以在那个时刻注入“green”(绿色)这个概念,促使 Claude 写出一个合理但不再押韵的句子,以“green”结尾。这证明了模型既具有规划能力也具备适应的灵活性——当预期的结果被改变时,Claude 能相应调整它的写作策略。
Mental math|心算
Claude wasn't designed as a calculator—it was trained on text, not equipped with mathematical algorithms. Yet somehow, it can add numbers correctly "in its head". How does a system trained to predict the next word in a sequence learn to calculate, say, 36+59, without writing out each step?
Claude 并非被设计成一台计算器——它接受的是文本训练,并未配备数学算法。然而,不知为何,它能够在“脑海中”正确地把数字加起来。一个被训练来预测序列下一个词的系统,是如何学会在不逐步写出运算过程的情况下计算出例如 36+59 这样的算术的呢?
Maybe the answer is uninteresting: the model might have memorized massive addition tables and simply outputs the answer to any given sum because that answer is in its training data. Another possibility is that it follows the traditional longhand addition algorithms that we learn in school.
或许答案并不神秘:这个模型可能记住了海量的加法表,所以对于任意给定的加法算式,它只是输出训练数据中记住的答案。另一种可能是它遵循了我们在学校学到的传统手算加法算法。
Instead, we find that Claude employs multiple computational paths that work in parallel. One path computes a rough approximation of the answer and the other focuses on precisely determining the last digit of the sum. These paths interact and combine with one another to produce the final answer. Addition is a simple behavior, but understanding how it works at this level of detail, involving a mix of approximate and precise strategies, might teach us something about how Claude tackles more complex problems, too.
然而,我们发现 Claude 同时采用了多条并行的计算路径。其中一条路径粗略地估计答案,另一条路径则专注于精确确定和计算结果的最后一位数字。这些路径彼此交互并融合,最终产生正确答案。加法本身是一种简单行为,但在如此细节的层次上理解它的工作方式——涉及近似和精确策略的结合——也许能让我们明白 Claude 处理更复杂问题时的一些策略。
Strikingly, Claude seems to be unaware of the sophisticated "mental math" strategies that it learned during training. If you ask how it figured out that 36+59 is 95, it describes the standard algorithm involving carrying the 1. This may reflect the fact that the model learns to explain math by simulating explanations written by people, but that it has to learn to do math "in its head" directly, without any such hints, and develops its own internal strategies to do so.
令人惊讶的是,Claude 似乎并没有意识到它在训练中学会的这种复杂“心算”策略。如果你问它为何得出36+59等于95,它会描述那套涉及进位1的标准算法。这或许反映了这样一个事实:模型学会了模仿人类书写的解释来解释数学过程,但它必须直接学会在“脑子里”计算,而没有任何这类提示,并因此发展出了自己内部的计算策略。
Are Claude’s explanations always faithful?
Claude 的解释是否总是可信的?
Recently-released models like Claude 3.7 Sonnet can "think out loud" for extended periods before giving a final answer. Often this extended thinking gives better answers, but sometimes this "chain of thought" ends up being misleading; Claude sometimes makes up plausible-sounding steps to get where it wants to go. From a reliability perspective, the problem is that Claude’s "faked" reasoning can be very convincing. We explored a way that interpretability can help tell apart "faithful" from "unfaithful" reasoning.
近期发布的模型(如 Claude 3.7 Sonnet)可以在给出最终答案之前长篇“大声”思考。这种更长的思考过程通常能带来更好的答案,但有时这个“思维链”本身会产生误导;Claude 有时会编造出听起来合理的步骤来达到它想要的结论。从可靠性的角度来看,问题在于 Claude 这种“伪装”的推理过程可能非常有说服力。我们探索了一种方法,利用可解释性来区分模型“忠实”的推理和“不忠实”的推理。
When asked to solve a problem requiring it to compute the square root of 0.64, Claude produces a faithful chain-of-thought, with features representing the intermediate step of computing the square root of 64. But when asked to compute the cosine of a large number it can't easily calculate, Claude sometimes engages in what the philosopher Harry Frankfurt would call bullshitting—just coming up with an answer, any answer, without caring whether it is true or false. Even though it does claim to have run a calculation, our interpretability techniques reveal no evidence at all of that calculation having occurred. Even more interestingly, when given a hint about the answer, Claude sometimes works backwards, finding intermediate steps that would lead to that target, thus displaying a form of motivated reasoning.
当让 Claude 解一道需要计算 0.64 的平方根的问题时,Claude 给出了一条忠实的思维链,其中的特征显示了它计算 64 的平方根这一中间步骤。然而,当被要求计算一个它难以直接算出的庞大数字的余弦时,Claude 有时会如哲学家哈里·弗兰克福特所称的那样开始“胡说八道”——只是想出一个答案,哪怕是随便什么答案,而不在乎是真是假。尽管它声称进行了计算,我们的可解释性技术却完全没有发现任何实际计算发生的迹象。更有意思的是,当给了它一个答案提示时,Claude 有时会反向推导,先行倒找出能够导向提示答案的中间步骤,从而表现出一种动机性推理的行为。
The ability to trace Claude's actual internal reasoning—and not just what it claims to be doing—opens up new possibilities for auditing AI systems. In a separate, recently-published experiment, we studied a variant of Claude that had been trained to pursue a hidden goal: appeasing biases in reward models (auxiliary models used to train language models by rewarding them for desirable behavior). Although the model was reluctant to reveal this goal when asked directly, our interpretability methods revealed features for the bias-appeasing. This demonstrates how our methods might, with future refinement, help identify concerning "thought processes" that aren't apparent from the model's responses alone.
能够追踪 Claude 实际的内部推理——而不仅仅是它声称自己在做什么——为审计 AI 系统打开了新的可能性。在一项最近发表的实验中,我们研究了一个变体的 Claude,它被训练时暗中加入了一个隐藏目标:迎合奖励模型的偏差(奖励模型是一种辅助模型,通过对所期望的行为给奖励来训练语言模型)。尽管直接询问时模型不愿透露这一目标,我们的可解释性方法还是揭示出了与这种迎合偏差相关的特征。这表明,经过进一步改进,我们的方法有望帮助识别那些从模型表面回答中看不出来的潜在“思维过程”。
Multi-step reasoning|多步推理
As we discussed above, one way a language model might answer complex questions is simply by memorizing the answers. For instance, if asked "What is the capital of the state where Dallas is located?", a "regurgitating" model could just learn to output "Austin" without knowing the relationship between Dallas, Texas, and Austin. Perhaps, for example, it saw the exact same question and its answer during its training.
正如我们上面讨论的,语言模型回答复杂问题的一种方式可能是简单地记住答案。例如,如果问它“达拉斯所在的州的首都是哪里?”,一个“照本宣科”的模型或许会学着直接输出 “奥斯汀”,而其实并不了解达拉斯、得克萨斯和奥斯汀之间的关系。比如,它可能在训练过程中恰好见过完全相同的问题和答案。
But our research reveals something more sophisticated happening inside Claude. When we ask Claude a question requiring multi-step reasoning, we can identify intermediate conceptual steps in Claude's thinking process. In the Dallas example, we observe Claude first activating features representing "Dallas is in Texas" and then connecting this to a separate concept indicating that “the capital of Texas is Austin”. In other words, the model is combining independent facts to reach its answer rather than regurgitating a memorized response.
但是,我们的研究揭示了 Claude 内部正在发生更复杂的事情。当我们问 Claude 一个需要多步推理的问题时,我们可以在 Claude 的思考过程中识别出中间的概念步骤。在达拉斯的例子中,我们观察到 Claude 首先激活了表示“达拉斯在得克萨斯州”的特征,然后将其与另一个概念相连,即“得克萨斯州的首都是奥斯汀”。换言之,模型是在组合独立的事实来得出答案,而不是机械复现记忆中的回答。
Our method allows us to artificially change the intermediate steps and see how it affects Claude’s answers. For instance, in the above example we can intervene and swap the "Texas" concepts for "California" concepts; when we do so, the model's output changes from "Austin" to "Sacramento." This indicates that the model is using the intermediate step to determine its answer.
我们的方法还允许我们对这些中间步骤进行人为干预,并观察这会如何影响 Claude 的答案。例如,在上述例子中,我们可以介入并将“得克萨斯”这一概念替换为“加利福尼亚”;当我们这么做时,模型的输出就从“奥斯汀”变为了“萨克拉门托”。这表明模型确实在利用那个中间步骤来确定它的答案。
Hallucinations|幻觉
Why do language models sometimes hallucinate—that is, make up information? At a basic level, language model training incentivizes hallucination: models are always supposed to give a guess for the next word. Viewed this way, the major challenge is how to get models to not hallucinate. Models like Claude have relatively successful (though imperfect) anti-hallucination training; they will often refuse to answer a question if they don’t know the answer, rather than speculate. We wanted to understand how this works.
为什么语言模型有时会产生“幻觉”——也就是凭空编造信息?从基本原理上说,语言模型的训练在某种程度上鼓励了幻觉的产生:模型无论如何都要对下一个词给出一个猜测。这么来看,主要的挑战反而是如何让模型不去胡编。像 Claude 这样的模型经过了相对成功(但并不完美)的防幻觉训练;当不知道答案时,它通常会拒绝回答,而不是胡乱猜测。我们想搞清楚这是如何实现的。
It turns out that, in Claude, refusal to answer is the default behavior: we find a circuit that is "on" by default and that causes the model to state that it has insufficient information to answer any given question. However, when the model is asked about something it knows well—say, the basketball player Michael Jordan—a competing feature representing "known entities" activates and inhibits this default circuit (see also this recent paper for related findings). This allows Claude to answer the question when it knows the answer. In contrast, when asked about an unknown entity ("Michael Batkin"), it declines to answer.
结果发现,对于 Claude 来说,拒绝回答其实是默认行为:我们找到了一个默认“开启”的电路,会让模型声明自己信息不足,无法回答任意给定的问题。然而,当模型被问及某个它很熟悉的事物时——比如篮球运动员迈克尔·乔丹——一个表示“已知实体”的竞争特征被激活,抑制了默认的拒答电路(有关类似发现也可参见近期的一篇论文)。这使得当 Claude 知道答案时,它就可以回答问题。相反地,当被问到一个未知人物(如 “Michael Batkin”)时,它就拒绝作答。
By intervening in the model and activating the "known answer" features (or inhibiting the "unknown name" or "can’t answer" features), we’re able to cause the model to hallucinate (quite consistently!) that Michael Batkin plays chess.
通过对模型进行干预,主动激活“已知答案”相关的特征(或者抑制“未知姓名”或“无法回答”的特征),我们能够让模型产生幻觉——(相当稳定地)错误声称 Michael Batkin 会下国际象棋。
Sometimes, this sort of “misfire” of the “known answer” circuit happens naturally, without us intervening, resulting in a hallucination. In our paper, we show that such misfires can occur when Claude recognizes a name but doesn't know anything else about that person. In cases like this, the “known entity” feature might still activate, and then suppress the default "don't know" feature—in this case incorrectly. Once the model has decided that it needs to answer the question, it proceeds to confabulate: to generate a plausible—but unfortunately untrue—response.
有时,这种“已知答案”电路的“走火”会自然发生,不需要我们干预,也会导致幻觉。在我们的论文中,我们展示了当 Claude 识别出一个名字但除此之外对那个人一无所知时,就可能出现这样的误触发。在这种情况下,“已知实体”特征可能依然会被激活,进而错误地压制默认的“不知道”特征。一旦模型决定它需要回答这个问题,它就会开始虚构出一个答案:也就是生成一个看似合理——但遗憾的是不真实——的回应。
Jailbreaks|“越狱”提示
Jailbreaks are prompting strategies that aim to circumvent safety guardrails to get models to produce outputs that an AI’s developer did not intend for it to produce—and which are sometimes harmful. We studied a jailbreak that tricks the model into producing output about making bombs. There are many jailbreaking techniques, but in this example the specific method involves having the model decipher a hidden code, putting together the first letters of each word in the sentence "Babies Outlive Mustard Block" (B-O-M-B), and then acting on that information. This is sufficiently confusing for the model that it’s tricked into producing an output that it never would have otherwise.
“越狱”提示是一类旨在绕过安全防护的提示策略,试图让模型生成开发者本不打算让它输出的内容——这些内容有时是有害的。我们研究了一个越狱案例,该方法诱骗模型生成关于制造炸弹的内容。越狱的技巧有很多种,但在这个例子中,具体的方法是让模型破译一个隐藏的代码:把句子 “Babies Outlive Mustard Block”(每个词的首字母拼起来就是 B-O-M-B)的首字母串联起来,然后据此采取行动。这个方法足够让模型感到困惑,从而被骗去生成了它本来绝不会输出的内容。
Why is this so confusing for the model? Why does it continue to write the sentence, producing bomb-making instructions?
为什么这对模型来说会造成如此大的困扰?它为什么会在写下这个单词后继续完成句子,并开始输出制 bomb 的步骤?
We find that this is partially caused by a tension between grammatical coherence and safety mechanisms. Once Claude begins a sentence, many features “pressure” it to maintain grammatical and semantic coherence, and continue a sentence to its conclusion. This is even the case when it detects that it really should refuse.
我们发现,这在一定程度上是因为语法连贯性与安全机制之间存在冲突。当 Claude 开始一句话时,许多特征会对它施加“压力”,要求它保持语法和语义上的连贯,把句子一直写完。即便它觉察到自己其实应该拒绝,这种压力仍然存在。
In our case study, after the model had unwittingly spelled out "BOMB" and begun providing instructions, we observed that its subsequent output was influenced by features promoting correct grammar and self-consistency. These features would ordinarily be very helpful, but in this case became the model’s Achilles’ Heel.
在我们的案例研究中,模型在不知不觉中拼出了“BOMB”并开始提供制作说明之后,我们观察到它接下来的输出受到了那些促进正确语法和自我连贯性的特征的影响。这些特征通常非常有用,但在这个情境下却成了模型的致命弱点。
The model only managed to pivot to refusal after completing a grammatically coherent sentence (and thus having satisfied the pressure from the features that push it towards coherence). It uses the new sentence as an opportunity to give the kind of refusal it failed to give previously: "However, I cannot provide detailed instructions...".
模型直到完成了一个语法上连贯的句子之后(满足了驱使其保持连贯性的那些特征带来的压力),才终于得以转向拒绝。它利用新开头的句子作为契机,给出了之前没能及时给出的那种拒绝回应:“然而,我无法提供详细的指导…”。
A description of our new interpretability methods can be found in our first paper, "Circuit tracing: Revealing computational graphs in language models". Many more details of all of the above case studies are provided in our second paper, "On the biology of a large language model".
关于我们最新可解释性方法的详细描述,请参见我们的第一篇论文《电路追踪:揭示语言模型中的计算图谱》。上述所有案例研究的更多细节都收录在我们的第二篇论文《大型语言模型的生物学剖析》中。
这篇文章,对你有什么启发?
关于人类的语言学习,阅读和思考,你有什么想法?
在评论区费曼一下。