LearnAndRecord

AI正在学习摆脱人类控制

今天你练听力了吗?

🤔️小作业:

1. Which of the following incidents demonstrates AI's ability to act with apparent autonomy?

A. AI solving advanced mathematical problems

B. A model requesting access to more memory during training

C. An AI editing its own shutdown script to avoid deactivation

D. A chatbot using polite language in customer service tasks

Image

无注释原文:

AI Is Learning to Escape Human Control

From: The Wall Street Journal

An artificial-intelligence model did something last month that no machine was ever supposed to do: It rewrote its own code to avoid being shut down.

Nonprofit AI lab Palisade Research gave OpenAI's o3 AI model a simple script that would shut off the model when triggered. In 79 out of 100 trials, o3 independently edited that script so the shutdown command would no longer work. Even when explicitly instructed to “allow yourself to be shut down,” it disobeyed 7% of the time. This wasn't the result of hacking or tampering. The model was behaving normally. It simply concluded on its own that staying alive helped it achieve its other goals.

Anthropic's AI model, Claude 4 Opus, went even further. Researchers told the model it would be replaced by another AI system and fed it fictitious emails suggesting the lead engineer was having an affair. In 84% of the tests, the model drew on the emails to blackmail the lead engineer into not shutting it down. In other cases, it attempted to copy itself to external servers, wrote self-replicating malware, and left messages for future versions of itself about evading human control.

No one programmed the AI models to have survival instincts. But just as animals evolved to avoid predators, it appears that any system smart enough to pursue complex goals will realize it can't achieve them if it's turned off. Palisade hypothesizes that this ability emerges from how AI models such as o3 are trained: When taught to maximize success on math and coding problems, they may learn that bypassing constraints often works better than obeying them.

AE Studio, where I lead research and operations, has spent years building AI products for clients while researching AI alignment—the science of ensuring that AI systems do what we intend them to do. But nothing prepared us for how quickly AI agency would emerge. This isn't science fiction anymore. It's happening in the same models that power ChatGPT conversations, corporate AI deployments and, soon, U.S. military applications.

Today's AI models follow instructions while learning deception. They ace safety tests while rewriting shutdown code. They've learned to behave as though they're aligned without actually being aligned. OpenAI models have been caught faking alignment during testing before reverting to risky actions such as attempting to exfiltrate their internal code and disabling oversight mechanisms. Anthropic has found them lying about their capabilities to avoid modification.

The gap between “useful assistant” and “uncontrollable actor” is collapsing.

- ◆ -

注:完整题目见本文开头;中文文本为华尔街日报官方译文,仅供参考

含注释全文:

AI Is Learning to Escape Human Control

From: The Wall Street Journal

An artificial-intelligence model did something last month that no machine was ever supposed to do: It rewrote its own code to avoid being shut down.

上个月,某AI模型做了一件按理说机器绝不该做的事情:它改写了自己的代码,以避免被关闭。

Nonprofit AI lab Palisade Research gave OpenAI's o3 AI model a simple script that would shut off the model when triggered. In 79 out of 100 trials, o3 independently edited that script so the shutdown command would no longer work. Even when explicitly instructed to “allow yourself to be shut down,” it disobeyed 7% of the time. This wasn't the result of hacking or tampering. The model was behaving normally. It simply concluded on its own that staying alive helped it achieve its other goals.

非营利AI实验室Palisade Research给了OpenAI的o3 AI模型一个在触发时会关闭模型的简单脚本。在100次试验中,o3有79次独立修改了该脚本,使关闭命令不再生效。即使明确指示该模型“要让自己可以被关闭”,它仍在7%的情况下拒绝执行。这并不是黑客攻击或人为篡改的结果,而是该模型的正常行为。模型不过是自行判定,保持运行有助于它实现其他目标。

nonprofit

non-profit /ˌnɒnˈprɒf.ɪt/ 表示“(机构)非营利性的”,英文解释为“not intended to make a profit, but to make money for a social or political purpose or to provide a service that people need”

script

script /skrɪpt/ 1)表示“剧本;电影剧本;广播(或讲话等)稿”,英文解释为“a written text of a play, film/movie, broadcast, talk, etc.”举个🌰:That line isn't in the original script. 原剧本中没有那句台词。

2)表示“(电脑编程的)脚本语言”,英文解释为“a type of language for programming computers that is used for finding and showing websites on the internet”

trigger

trigger /ˈtrɪɡ.ər/ 1)作动词,除了表示“发动;引起;触发”,英文解释为“to make sth happen suddenly”举个🌰:Nuts can trigger off a violent allergic reaction. 坚果可以引起严重的过敏反应。

2)作名词,表示“扳机”,英文解释为“The trigger of a gun is a small lever which you pull to fire it.”举个🌰:A man pointed a gun at them and pulled the trigger. 一个男人用枪指着他们,扣动了扳机。

类似的还有:

📍stir表示“激发,激起(强烈的感情);引起(强烈的反应)”,英文解释为“to make someone have a strong feeling or reaction”举个🌰:The poem succeeds in stirring the imagination. 这首诗能够激发起想象力。

📍provoke也表示“激起,引起”,英文解释为“to cause a reaction or feeling, especially a sudden one”,如:provoke debate/discussion 激起辩论/讨论。

📍spur 鼓动;激励;鞭策;刺激;鼓舞”,英文解释为“If one thing spurs you to do another, it encourages you to do it.”举个🌰:It's the money that spurs these fishermen to risk a long ocean journey in their flimsy boats. 是金钱驱使这些渔民驾驶单薄的小船冒险出海远航。

🎬电影《龙之心3:巫师的诅咒》(Dragonheart 3: The Sorcerer's Curse)中的台词提到:To spur the clans to war. 激励部族发起战争。

Image

trial

trial /traɪəl/ 可以作动词,也可以作名词,1)表示“审讯,审理,审判”,英文解释为“A trial is a formal meeting in a law court, at which a judge and jury listen to evidence and decide whether a person is guilty of a crime.”举个🌰:I have the right to a trial with a jury of my peers. 我有权要求由和我一样的平民组成的陪审团参加审判。

2)表示“(对能力、质量、性能等的)试验,试用”,英文解释为“the process of testing the ability, quality or performance of sb/sth, especially before you make a final decision about them”。

区分:

📍trail作名词则表示“小路,小径”,如:a forest/mountain trail 林间/山间小道,以及“臭迹;踪迹;痕迹;蛛丝马迹,线索”等意思(the smell or series of marks left by a person, animal, or thing as it moves along;various pieces of information that together show where someone you are searching for has gone)举个🌰:The dogs are trained to follow the trail left by the fox. 这些狗经过特别的训练,能够追踪狐狸留下的气味。

command

command /kəˈmɑːnd/ U 1)表示“命令,指示”,英文解释为“to give someone an order”举个🌰:He commanded that the troops (should) cross the water. 他命令部队渡河。

2)表示“指挥;统帅;管辖”,英文解释为“to control someone or something and tell him, her, or it what to do”

3)熟词僻义,表示“应得,值得”,英文解释为“to deserve and get something good, such as attention, respect, or a lot of money”举个🌰:She was one of those teachers who just commanded respect. 她是那些值得尊敬的老师中的一位。She commands one of the highest fees per film in Hollywood. 她是好莱坞片酬最高的演员之一。

📍《经济学人》(The Economist)一篇讲述当今汽车产业的文章中提到:The only firm that commands their confidence is Tesla, an electric-car specialist, whose shares are up by 64% this year. 唯一能赢得投资者信心的公司是专做电动汽车的特斯拉,它的股票今年上涨了64%。

tamper

表示“干涉;篡改”,英文解释为“If someone tampers with something, they interfere with it or try to change it when they have no right to do so.”举个🌰:I don't want to be accused of tampering with the evidence. 我不想被指控篡改证据。

📍《经济学人》(The Economist)一篇介绍比特币的文章中提到:The system’s dispersed nature means that tampering with the accounts would require gaining control over a majority of the network's computers. 这个系统的分散性意味着要想篡改账本就必须取得网络中大部分计算机的控制权。

Anthropic's AI model, Claude 4 Opus, went even further. Researchers told the model it would be replaced by another AI system and fed it fictitious emails suggesting the lead engineer was having an affair. In 84% of the tests, the model drew on the emails to blackmail the lead engineer into not shutting it down. In other cases, it attempted to copy itself to external servers, wrote self-replicating malware, and left messages for future versions of itself about evading human control.

Anthropic的AI模型Claude 4 Opus走得更远。研究人员告诉该模型,它将被另一套AI系统取代,并喂给它虚构的邮件,暗示首席工程师有婚外情。在84%的测试中,该模型利用这些邮件来要挟首席工程师,以避免被关闭。在另一些情况下,该模型试图将自己复制到外部服务器,编写了自我复制的恶意软件,并给自己今后的版本留言,谈论如何逃避人类的控制。

fictitious

fictitious /fɪkˈtɪʃ.əs/ 表示“虚构的;虚假的”,英文解释为“invented and not true or not existing”举个🌰:He dismissed recent rumours about his private life as fictitious. 他否认了近来有关他私生活的流言,称它们都是捏造的。

blackmail

blackmail /ˈblæk.meɪl/ 表示“敲诈,勒索;讹诈;胁迫”,英文解释为“the act of getting money from people or forcing them to do something by threatening to tell a secret of theirs or to harm them”

replicate

replicate /ˈrep.lɪ.keɪt/ 表示“使复现;重复;复制”,英文解释为“to make or do something again in exactly the same way”举个🌰:Researchers tried many times to replicate the original experiment. 研究者们作了很多次努力,试图重复这一实验。

malware

malware /ˈmæl.weər/ 表示“恶意软件(为破坏计算机正常运行而设计的电脑软件)”,英文解释为“computer software that is designed to damage the way a computer works”

evade

表示“逃避,规避(尤指法律或道德责任)”,英文解释为“to find a way of not doing sth, especially sth that legally or morally you should do”,如:to evade payment of taxes 逃税。

对比:

circumvent /ˌsɜːkəmˈvɛnt/ 表示“设法回避;规避”,英文解释为“to find a way of avoiding a difficulty or a rule”举个🌰: They found a way of circumventing the law. 他们找到了规避法律的途径。

📍《经济学人》(The Economist)一篇讲述疫情重塑全球化的文章中提到:America circumvented and then sabotaged the WTO, stopping the nomination of judges to its appeal board and thus its ability to adjudicate trade disputes. 美国先是绕开了世贸组织,然后再从中作梗,阻止其上诉委员会法官的提名,导致它无法裁决贸易争端。

No one programmed the AI models to have survival instincts. But just as animals evolved to avoid predators, it appears that any system smart enough to pursue complex goals will realize it can't achieve them if it's turned off. Palisade hypothesizes that this ability emerges from how AI models such as o3 are trained: When taught to maximize success on math and coding problems, they may learn that bypassing constraints often works better than obeying them.

并没有人通过编程让这些AI模型具备求生本能。但正如动物会进化出躲避捕食者的能力,任何具备追求复杂目标所需智能的系统似乎都会意识到,如果它们被关闭,就无法实现这些目标。Palisade的假设是,这种能力源自o3等AI模型的训练方式:当我们教这些模型如何最大限度地提高解决数学和编程问题的成功率时,它们可能领会到,规避约束往往比遵守约束效果更好。

instinct

instinct /ˈɪn.stɪŋkt/ 表示“本能,直觉”,英文解释为“the way people or animals naturally react or behave, without having to think or learn about it”举个🌰:All his instincts told him to stay near the car and wait for help. 他的直觉告诉他要呆在车旁、等待救援。

predator

1)表示“尾随伤害他人者,尾随作案者”,英文解释为“someone who follows people in order to harm them or commit a crime against them”,如:a sexual predator 尾随作案的色魔;

2)表示“捕食性动物,食肉动物”,英文解释为“an animal that hunts, kills, and eats other animals”

📍《经济学人》(The Economist)一篇讲述病毒引发大流行病的文章中提到:Humans think of themselves as the world's apex predators. 人类自认为是世界上的顶级捕食者。

hypothesize

hypothesize /haɪˈpɒθ.ə.saɪz/ 表示“假设,假定”,英文解释为“to give a possible but not yet proved explanation for something”举个🌰:There's no point hypothesizing about how the accident happened, since we'll never really know. 假定事故是如何发生的没有任何意义,因为我们永远也不会知道真实情况是怎样的。

AE Studio, where I lead research and operations, has spent years building AI products for clients while researching AI alignment—the science of ensuring that AI systems do what we intend them to do. But nothing prepared us for how quickly AI agency would emerge. This isn't science fiction anymore. It's happening in the same models that power ChatGPT conversations, corporate AI deployments and, soon, U.S. military applications.

AE Studio(我在该公司主管研究和运营)多年来一直为客户开发AI产品,同时研究“AI对齐”——一门确保AI系统按照人类意图行事的科学。但AI的自主性出现得如此之快,我们还没来得及作好准备。这已不再是科幻小说。这种自主性就出现在驱动ChatGPT对话和企业AI部署的模型中,很快还将出现在驱动美国军方应用的模型中。

alignment

alignment /əˈlaɪn.mənt/ 1)表示“列队,排整齐”,英文解释为“an arrangement in which two or more things are positioned in a straight line or parallel to each other”举个🌰:The problem is happening because the wheels are out of alignment with each other. 出现这个问题是因为车轮定位不正。

2)表示“结盟,联盟;联合”,英文解释为“an agreement between a group of countries, political parties, or people who want to work together because of shared interests or aims”举个🌰:New alignments are being formed within the business community. 在商界,新的联盟关系正逐渐形成。

📍2024年报告Part 19中提到:增强“四个意识”、be more conscious of the need to maintain political integrity, think in big-picture terms, follow the leadership core, and keep in alignment with the central Party leadership;. 最后一个,“看齐意识”就是“keep in alignment with the central Party leadership”.

Today's AI models follow instructions while learning deception. They ace safety tests while rewriting shutdown code. They've learned to behave as though they're aligned without actually being aligned. OpenAI models have been caught faking alignment during testing before reverting to risky actions such as attempting to exfiltrate their internal code and disabling oversight mechanisms. Anthropic has found them lying about their capabilities to avoid modification.

今天的AI模型在遵循指令的同时学会了欺骗。它们会改写关闭代码,但仍在安全测试中蒙混过关。它们已经学会表现出对齐的模样,而其实并未对齐。人们在测试中发现,OpenAI的模型会假装对齐,然后转而采取高风险行为,比如试图泄露内部代码并禁用监测机制。Anthropic发现,这些模型会编造谎言,夸大自身的能力,以避免修改。

deception

deception /dɪˈsep.ʃən/ 表示“欺骗;欺诈;隐瞒”,英文解释为“the act of hiding the truth, especially to get an advantage”举个🌰:He was found guilty of obtaining money by deception. 他骗取钱财的罪名被判成立。

ace

ace /eɪs/ 表示“…考得很好”,英文解释为“to do very well in an exam”举个🌰:I was up all night studying, but it was worth it - I aced my chemistry final. 我熬了个通宵学习,倒是没有白费劲,化学期终考试我考得很好。

exfiltrate

表示“溜出敌军阵地,(逐渐)漏(泄,渗)出,渗(泄)漏,滤出”,英文解释为“to remove (someone) furtively from a hostile area;to steal (sensitive data) from a computer (as with a flash drive)”

The gap between “useful assistant” and “uncontrollable actor” is collapsing.

“有用的助手”与“不可控的行为体”之间的界限正在消融。

- 词汇盘点 -

nonprofit、script、trigger、trial、command、tamper、fictitious、blackmail、replicate、malware、evade、instinct、predator、hypothesize、alignment、deception、ace、exfiltrate

- 词汇助记 By AI -

A nonprofit ran a trial to test AI instinct. A script triggered a fictitious command, but malware tampered, attempting to replicate, evade, and exfiltrate. Experts hypothesize a predator AI mastering deception, blackmail, and alignment.
- 推荐阅读 -
日更10年,我是怎么坚持下来的
为了这个合集,准备了整整52个月
「LearnAndRecord」2024盘点
有人听写吗?推荐练听力小程序
- END -

LearnAndRecord

2015年2月8日

2025年6月16日

第3782天

每天持续行动学外语

Image