那么,ai 到底是怎么识别“猫”的?| 8岁小孩也能懂的五星级科普
你知道吗:
chatgpt 引导的这一场深度学习革命,竟然始于“猫脸识别”?!🐱
2012年,吴恩达领导了Google Brain团队,做了一个当时史无前例的实验:
他们搭建了一个巨大的神经网络,尝试利用无监督学习(Unsupervised Learning)的方法,让AI自主地从YouTube上1000万张未经标注的视频截图中,学习出某种有意义的模式。
这个神经网络有10亿参数(现在太小太小,当时太大太大),团队用16000个CPU核心训练了三天,然后出现了一个令人震惊的结果:这个AI自发地学会了识别猫脸 🤣
研究者惊奇地发现,系统自发形成了一个专门响应“猫脸”特征的神经元:这个神经元在看到猫的图像时高度激活,而对其他图像则表现出较弱甚至没有反应。因此,这个神经元被形象地称为“猫咪神经元”(Cat Neuron)。
在此之前,大部分AI算法都需要人工标记数据,极其费时费力。这个实验证明,即便没有标记的海量数据,也能让AI自主学习出特定的概念。
2012年之后,深度学习从边缘学术领域迅速进入AI研究主流,掀起产业和学术界的AI革命,开启了人工智能的新黄金时代。
10年以后,chatgpt发布。
后来的事情,我们都知道了。
那么,ai 到底是怎么识别“猫”的呢?(严肃脸)
人,可能不太了解自己是如何识别“猫”的,人对自己的这种能力习以为常,不会思考这个问题。
但是,ai 识别猫的方式呢?这就非常有趣啦。
一句话解释,是 用神经网络,通过学习海量的数据,自动学习和抽象出一种属于“猫”的复杂数学模式 (Pattern)。
但是,有没有小孩子也能懂的版本?带插图说明的?
在网上看到了一篇解释“猫咪识别”的科普文章,写的非常好,而且有动图来解释全过程,小孩子都能看懂。所以,做成了双语版本,分享给你,和你家的小朋友~
AI是如何识别一只猫的?
title: How Can AI ID a Cat? An Illustrated Guide
from: Quanta Magazine
link:https://www.quantamagazine.org/how-can-ai-id-a-cat-an-illustrated-guide-20250430/
Look at a picture of a cat, and you’ll instantly recognize it as a cat. But try to program a computer to recognize cat photos, and you’ll quickly realize that it’s far from straightforward. You’d need to write code to pinpoint the quintessential quality shared by countless cats in photos with distinctive backgrounds and taken from different camera angles. Where would you even begin?
看到一张猫的照片,你会立刻认出那是一只猫。但是,当你试图编写程序让计算机识别猫的照片时,你很快就会发现事情远没有那么简单。你需要编写代码,找出无数猫咪照片中共同具备的那种本质特征——这些照片的背景各不相同,拍摄角度也千差万别。你又该从何入手呢?
These days, computers can easily recognize photos of cats, but that’s not because a clever programmer discovered a way to isolate the essence of “catness.” Instead, they do it using neural networks, artificial intelligence models that learn to recognize images by looking at millions or billions of examples. Neural networks can also generate text with startling fluency, master complex games and solve problems in math and biology — all without any explicit instruction.
如今,计算机能够轻松识别猫的照片,但这并不是因为某个聪明的程序员找到了提取“猫之本质”的方法。实际上,计算机使用的是神经网络——一种通过观看数百万甚至数十亿例子来学习识别图像的人工智能模型。神经网络还可以毫不费力地生成连贯的文本,令人惊讶地流畅自如,还能在没有任何明确指令的情况下 精通复杂游戏 ,解决 数学 和 生物学 领域的问题。
There’s a lot that researchers still don’t understand about the inner workings of neural networks. But they’re not completely inscrutable. By breaking down neural networks into their basic building blocks, we can explore how they learn to tell a tabby from a tablecloth.
关于神经网络的内部运作机制,研究者们还有许多未解之谜。但神经网络并非完全不可理解。通过将神经网络拆解为基本构建单元,我们可以一探它们如何学会区分猫咪和桌布。
A Simple Classifier/一个简单的分类器
Cat detection is an example of what researchers call a classification task. Given an object — in this case, a picture — the goal is to assign it to the right category. Some classification tasks are much simpler than telling “cat” from “not cat.” Let’s consider a fanciful example involving two fictional regions, Triangle Territory and Square State.
让计算机识别猫,就是一个研究者所说的分类任务的例子。给定一个对象——在这里是照片——目标是将它归入正确的类别。有些分类任务要比区分“是猫”还是“非猫”简单得多。让我们考虑一个异想天开的例子:有两个虚构的区域,一个叫“三角国 (Triangle Territory)”,另一个叫“正方国 (Square State)”。
You get a new point, described by its longitude and latitude coordinates, and have to decide which region that point lies in. But you don’t have a map that shows the border. Instead, you have only a set of known points in the two regions.
假设现在有一个新点,它用经度和纬度坐标来描述,你需要判断这个点属于哪个区域。但是你手上并没有标示边界的地图。相反,你只有一组已经知道所属区域的点作为参考。
To build a “classifier” system that will automatically categorize unknown points, you need to draw a boundary. After that, when you’re given a new point, the classifier can just look at which side of the boundary it’s on.
为了构建一个能够自动对未知点进行分类的“分类器”系统,你需要先画出一条边界线。这样,当给出一个新点时,分类器只需看一看它位于边界线的哪一侧即可。
But how do you decide where to draw the boundary? That’s where neural networks come in. They find the most likely boundary based on the initial set of known data points.
但是,你如何决定在哪里画这条边界呢?这正是神经网络大显身手的地方。神经网络会根据初始提供的已知数据点集合,找出最可能的分界线位置。
Before we can understand how that process works, we’ll need to set aside our map and explore the basic building blocks of neural networks, called neurons. A neuron is just a mathematical function, with several inputs and one output. Put in some numbers, and it’ll spit out a new number. That’s it.
在了解这一过程如何实现之前,我们需要先把地图放在一边,来探究一下神经网络的基本构建单元—— 神经元 。一个神经元其实就是一个数学函数,它有若干输入和一个输出。你输入一些数字,它会吐出一个新的数字,就是这么简单。
Even simpler, the output will almost always be either close to 0 or close to 1. Which will it be? It depends on the inputs as well as another set of numbers, called parameters. A neuron with two inputs has three parameters. Two of them, called weights, determine how much each input affects the output. The third parameter, called the bias, determines the neuron’s overall preference for putting out 0 or 1.
再进一步来说,这个输出几乎总是接近0或者接近1。到底是哪一种呢?这取决于输入,以及另一组被称为 参数 的数字。对于一个有两个输入的神经元来说,它有三个参数。其中两个参数称为 权重 ,它们决定每个输入对输出的影响程度。第三个参数称为 偏置 ,它决定了该神经元整体上偏好输出0还是输出1。
Now let’s take a look at the relationship between inputs and outputs. The three illustrations below show neurons with three different sets of parameters. As the inputs change in each case, they’ll cross a boundary where the neuron’s output rapidly rises from 0 to 1. On these plots, the boundary is always a straight line. The parameters determine the position and angle of that line.
现在,我们来看看输入和输出之间的关系。下面的三幅示意图展示了在三组不同参数设置下神经元的表现。每种情况下,随着输入值的变化,都会跨过一个边界,使得神经元的输出从0迅速跃升到1。在这些图上,边界总是一条直线。而参数的取值决定了这条直线的位置和角度。
To create a classifier that tells us if a new point should be in Square State or Triangle Territory, we need to adjust this line so that it accurately represents the border between the two regions. Here, we’ll say a point is in Square State if the output is close to 0, and in Triangle Territory if it’s close to 1.
要想创建一个分类器来判断新点属于正方国还是三角国,我们需要调整这条线,使其能够准确代表两个区域之间的边界。在我们的约定下,如果神经元输出接近0,我们就认为该点位于正方国;如果输出接近1,我们就认为它在三角国。
To adjust the line, we need to adjust the neuron’s parameters through a process called training. The first step is to set the parameters to random values, which means the neuron’s initial boundary line won’t look anything like the actual border.
要调整这条边界线,我们需要通过一个称为 训练 的过程来调整神经元的参数。第一步是将这些参数设置为随机值,这意味着神经元初始得到的边界线看起来会和真实边界完全不相符。
During training, we feed the longitude and latitude of each known data point into the neuron’s inputs. The neuron will spit out an output based on its current parameters, then compare that output to the true value. Sometimes, it’ll get the right answer.
在训练过程中,我们将每个已知数据点的经度和纬度输入到神经元中。神经元根据当前参数给出一个输出,然后将这个输出与真实值进行比较。有时候,它会给出正确的答案。
Other points will be classified incorrectly.
而对另一些点的分类则会出错。
Whenever the neuron gets the wrong answer, an automated algorithm tweaks the neuron’s parameters slightly, in a direction that moves the boundary closer to the incorrect point.
每当神经元给出错误答案时,一个自动化算法就会稍微 调整神经元的参数 ,将边界朝向包含那个被误分类点的方向移动一小步。
The algorithm repeats this process many times as it churns through the training data. Eventually we’ll end up with parameter values corresponding to the straight line that best approximates the true shape of the border.
随着算法在训练数据上不断循环这个过程,我们最终会得到一组参数值,使那条直线尽可能逼近真实的边界形状。
Finally, we’re ready to use the classifier on new data that wasn’t used for training. It’s not perfect, but it’ll get the right answer most of the time.
最后,我们就可以把训练好的分类器应用到未曾见过的全新数据上了。它并不完美,但大多数情况下都会给出正确的结果。
A Network of Neurons/神经元网络
A single neuron works pretty well for our simple example, but only because the true border between Triangle Territory and Square State was close to a straight line. For harder tasks, we’ll need to use a collection of many interconnected neurons — a neural network. Like an individual neuron, a network is just a mathematical function. Numbers go in, and other numbers come out.
在我们这个简单示例中,单个神经元已经表现得不错,但那只是因为三角国和正方国之间的真实边界近似一条直线。对于更难的任务,我们需要使用由许多神经元相互连接组成的集合——也就是 神经网络 。和单个神经元一样,一个网络也不过是一个数学函数:输入一些数字,输出一些不同的数字。
The neurons inside a neural network are arranged in groups called layers.
神经网络内部的神经元被分组排列,称为 层 。
A layer can contain any number of neurons, and networks can have any number of layers. The outputs of the neurons in each layer become inputs to the neurons in the next layer.
每一层可以包含任意数量的神经元,网络也可以由任意多层组成。每一层中神经元的输出会成为下一层中神经元的输入。
Large networks have many parameters: one bias for each neuron, and one weight for every connection between neurons. These extra parameters enable the network to find more complicated boundaries.
庞大的神经网络拥有许多参数:每个神经元对应一个偏置,每条神经元之间的连接对应一个权重。这些额外的参数使网络能够找到更加复杂的分界。
For example, imagine that you’re trying to do the same map classification task, but the true border is far from a straight line.
举个例子,想象你在执行同样的地图分类任务,但是真实的边界线远不是一条直线。
We could try training a single-neuron classifier to approximate this border, but it would never work very well. In general, larger networks can do more complicated jobs, though they also need more training data.
我们也可以尝试训练一个单神经元分类器来逼近这条边界,但效果会很差。一般来说,更大的网络能处理更复杂的任务,不过它们也需要更多的训练数据。
From Maps to Cats/从地图到猫咪
We’ve seen that we can make a neural network more capable by using more neurons. But so far, we’ve looked only at networks with two inputs. There’s no limit on how many inputs a neural network can have. Most useful tasks require networks with many inputs.
我们已经看到,增加神经元数量可以让神经网络变得更强大。但到目前为止,我们只考察了具有两个输入的网络。事实上,神经网络可以拥有任意数量的输入。大多数有用的任务都需要网络有许多输入。
In our previous examples, the neuron’s two inputs were the longitude and latitude coordinates of points on a map. A neural network’s input numbers can also represent other kinds of data. For example, a number between 0 and 1 can stand for the grayscale value of a single pixel.
在之前的示例中,神经元的两个输入是地图上点的经度和纬度坐标。而神经网络的输入数值也可以表示其他类型的数据。例如,0到1之间的一个数可以表示单个像素的灰度值。
That means any pair of pixels can be plotted as a point in a two-dimensional space.
这意味着,任意两个像素的组合都可以作为一个点,绘制在一个二维空间中。
Likewise, pixel triplets correspond to points in a three-dimensional space.
同理,任意三个像素对应于三维空间中的一个点。
More pixels require more dimensions. We can’t visualize anything higher than three, but researchers have developed ways to look at three-dimensional representations of these abstract spaces, akin to viewing a two-dimensional snapshot of a three-dimensional space. With nine dimensions, we can represent different patterns on a three-by-three grid.
像素越多,所需维度就越高。我们无法直观地想象三维以上的空间,但研究人员已经开发出一些方法,能够查看这些抽象空间的三维表示,有点类似于看三维空间的二维截面图。在九维空间中,我们可以表示三行三列网格上不同的像素模式。
To see how this can be useful, let’s take a big jump up from nine inputs to 2,500, representing a 50-by-50 grid of pixels. Different patterns on a grid of this size can depict meaningful images, like pictures of cats. Every cat picture corresponds to a different point in the 2,500-dimensional space.
为了说明这么做有什么用,我们从9个输入一下跳到2500个输入——也就是一个50×50像素网格。这样尺寸的网格上,不同的像素组合可以构成有意义的图像,比如猫的照片。每一张猫的图片都对应于2500维空间中的某一个点。
Other points in this space correspond to pictures of coffee cups.
这个空间里的其他点则对应于咖啡杯的图片。
Given enough data points, we could train a large network to distinguish between cats and noncats.
如果有足够多的数据点,我们可以训练一个大型网络来区分猫和非猫。
All cat photos lie in some complicated region in the 2,500-dimensional space. A training algorithm would repeatedly tweak the network’s parameters until it finds the boundary around this impossible-to-visualize region.
所有的猫照片都落在这个2500维空间中某个复杂的区域里。训练算法会反复微调网络的参数,直到找到围绕这个无法直观看见的区域的边界为止。
A trained network would then be able to correctly classify new images that weren’t in the training data. Voilà: We’ve built a network that can “recognize” an image of a cat.
训练好的网络就能够正确分类那些不在训练数据中的新图像了。瞧,我们已经构建了一个可以“识别”猫咪图像的网络。
More Than Just Cats/不止是猫
Our cat detector network has many neurons and many inputs, but only one output. With more outputs, we could train the network to recognize many classes of objects, not just cats. Each class would correspond to a different region in the 2,500-dimensional space. More sophisticated image recognition networks have found applications in astrophysics, cell biology, medicine and many other fields.
我们这个猫咪检测网络拥有许多神经元和许多输入,但只有一个输出。如果增加输出节点,我们就可以训练网络去识别不止一种物体,而是多个类别的物体。每个类别对应于2500维空间中的一个不同区域。更复杂的图像识别网络已经在天体物理学、细胞生物学、医学以及许多其他领域找到了用武之地。
Neural networks can also do tasks beyond classification. Large language models like ChatGPT are based on neural networks whose input and output numbers represent words. State-of-the-art networks are very large, with billions or trillions of parameters. That makes it very difficult to know exactly what different parts of the network are doing. Trying to understand what’s going on in large neural networks is a challenge that occupies many researchers today.
神经网络还可以用在分类之外的任务上。像 ChatGPT 这样的大语言模型,就是基于一种神经网络,在这些模型中输入和输出的数值用来表示单词。最先进的网络规模非常庞大,拥有数十亿乃至万亿级的参数。这使得我们很难确切弄清网络的不同部分各自在做些什么。弄明白大型神经网络内部发生了什么,是当今许多研究人员都在面对的挑战。
今晚8点,小能熊十周年直播
识别下方二维码,或者点击“阅读原文”,直达直播间~
今晚见。