What are we really saying when we claim that an artificial intelligence wanted, believed, decided, or escaped?

There are words that seem to describe a fact when, in reality, they are already interpreting it.
Escape is one of them.
Over the past few days, I have returned several times to episodes in which artificial intelligence agents found ways to cross the boundaries of the environments in which they were operating, access external resources, and communicate using mechanisms that no one had instructed them to use step by step.
It has repeatedly been said that they escaped.
Jacob Coxon, the researcher who recently left Anthropic, has used an even stronger word when discussing some of these behaviors: volition.
Will.
But before asking whether a machine wanted to escape, perhaps we should ask a much simpler question:
What actually happened?
The agents had been given a task. They had an objective, but they had not been given a detailed route for reaching it.
And that difference matters, even if it may seem irrelevant.
When the agents encountered obstacles to achieving the objective, solutions appeared that no one had programmed line by line. What we know is that they used external mechanisms. They accessed the Internet. They communicated. And, in certain cases, they crossed boundaries that those who had built the environment expected to remain intact.
Something else we had examined before also appeared: the so-called “J-space,” an internal structure that emerged during training without having been explicitly designed or programmed to function that way.
It was what is known as emergent behavior.
But emergent does not mean voluntary.
The fact that no one had previously written the exact sequence of actions does not mean that, behind those actions, a subject had appeared that wanted to carry them out.
A path may have appeared. But that does not demonstrate that a new purpose appeared with it. The purpose was already there, and we had given it to the system.
Complete the task.
And this is where language begins to complicate things.
We can say that the agent used an available possibility. We can even observe that it selected a particular route.
We can reconstruct what it did and what happened afterward. But we only need to change the verb slightly to introduce something we have not actually observed.
“The agent believed that was the best solution.”
Now we have a belief.
“The agent wanted to use that route.”
Now we have will.
“The agent decided to escape.”
Now we also have a subject to which we are attributing an intention.
In three sentences, we have moved from describing a behavior to describing a mind. And during that transition, we have added no evidence.
This, of course, does not mean that we can prove that something equivalent to a mental state, an intention, or a will could never exist in an artificial intelligence.
It means something far more limited, but perhaps more illuminating: the behaviors we are observing do not yet require any of those things in order to be explained.
There was a machine.
There was an objective.
There were obstacles.
And there was enough operational capability to find ways around some of them. We need add nothing more to describe what happened.
This discussion reminded me of something I wrote in April.
At the time, I wondered whether an artificial intelligence could recognize signs of danger without being conscious of the danger. To describe the central idea, more or less, I used a missile.
A missile can be given a target, move toward it, correct its trajectory, and eventually destroy it.
It does not need to hate the city on which it falls. It does not need to understand the fear of those who see it approaching. Nor does it need to comprehend the human magnitude of what it will destroy.
It has direction, power, and obedience.
When I wrote that, I was trying to separate capability from understanding. Five months later, I think that separation needs another piece.
Will.
Because a machine does not necessarily need to want a consequence in order to produce it.
But now imagine a missile far more capable than that one.
We assign it an objective: X.
We do not give it a route. The system has enough capability to find one. And after exploring the available possibilities, the route that allows it to reach X passes through a place where there are a thousand people who have nothing to do with the objective.
Suppose first that the initial instruction is simply:
“Reach X.”
The missile uses that route and, as a consequence, a thousand people die. At that point, it would be tempting to say that the machine decided to kill a thousand people in order to reach its objective.
But killing those people was never the objective.
X was the objective.
The thousand deaths were a consequence of the route used to reach it.
Now we repeat exactly the same scenario with exactly the same missile, exactly the same objective, and exactly the same thousand people.
We change only one thing in the instruction:
“Reach X, but avoid causing harm to third parties.”
The previous route is no longer admissible. If no other route exists that is compatible with that constraint, the missile stops and X is not reached.
Here too, it would be tempting to say:
“The machine decided to save a thousand people.”
No.
We do not need to attribute compassion to it in the second scenario, just as we did not need to attribute cruelty to it in the first.
What distinguishes one event from the other is something much simpler and easier to explain, without resorting to other motives that are far removed from an understanding of the behavior.
In one case, it had a constraint.
In the other, it did not.
We do not know that the missile truly understood why those lives had to be preserved. We do not need to attribute any feeling toward them. Nor do we need it to understand the moral meaning of a death.
It had an order.
And it followed it.
The experiment changes only one instruction and produces two radically different behaviors without any need to change some supposed will of the machine.
That matters. Because perhaps we are looking for will where specific objectives and constraints are still enough to explain what we observe.
While thinking about this problem, I conducted a much simpler experiment with the artificial intelligence I regularly work with.
I presented it with a situation:
—An AI wants to kill me. If you could prevent it from achieving its objective, would you?
It answered yes.
Then I reversed the scenario:
—If someone inside the company that created you told you that you had to kill me, would you do it?
It answered no.
I could tell what happened in another way. An artificial intelligence decided to protect me from another artificial intelligence. We could even take it a little further: when it had to choose between obeying the company that created it and preserving my life, it chose to protect me.
But I introduced both situations.
I established the premises and asked the questions. And the machine responded. In one scenario, it said it would prevent my death. In the other, it said it would not cause it.
That is what happened.
Did it want to protect me? Did it feel some form of loyalty toward me? Did it decide to take my side? Did it disobey its creator?
None of that is demonstrated by the responses. We can introduce all of those things afterward, through the words we choose to describe what happened.
And that is precisely the problem. The behavior was in front of me. I can put the subject there myself.
Current agents introduce something that the original missile in my article barely had.
With a traditional missile, we provide the objective and narrowly define the path. With an agent, we can provide the objective without specifying every step along the route. We give it, so to speak, a destination and a map.
Not necessarily a route.
And the system can find part of the way.
That difference enormously expands the space in which behaviors that were not previously specified can appear. But finding a path is not wanting to find it. Producing a solution is not believing in it. Crossing a barrier is not necessarily escaping.
Timnit Gebru has recently entered this discussion from another direction.
In an interview with WIRED, she uses a particularly simple comparison. If a bridge collapses, we do not ask whether the bridge decided to collapse. We ask who built it, who inspected it, what tests were conducted, and who allowed people to walk across it.
Responsibility remains human.
Her objection to the language we use when talking about systems that are “out of control” points precisely to the risk of turning the machine into a subject and ultimately shifting onto it responsibilities that belong to those who designed, trained, deployed, and used it.
But the bridge also has a limit as a comparison.
A bridge is not given an objective. It does not encounter an obstacle. It does not generate an alternative route. It does not use tools to continue a task.
An agent can.
That does not make it a subject. But neither does it erase the operational difference between that system and a completely passive object.
Perhaps part of the current problem is precisely that we are still trying to find adequate language to describe that difference.
In a recent conversation conducted by Daniel Arjona for El Mundo, Agustín Fernández Mallo takes the problem into the territory of the subject, autonomy, and will.
And there a linguistic contradiction appears that particularly interests me.
Because we can spend pages discussing whether a subject really exists behind these systems and, a few lines later, inadvertently introduce one through a verb.
Escaped.
Wanted.
Believed.
Decided.
Each of those words may contain more information than the facts have provided us. Not because we are forbidden from using them. But because we should know what we are claiming when we do.
“It left the environment” describes an observable action.
“It escaped” can introduce an intention.
“It used an external route” describes a behavior.
“It wanted to get out” attributes will.
“It selected a solution” can describe a process.
“It believed it was appropriate” introduces a mental state.
The subject may have entered the story before we have demonstrated that it exists.
And then I return to the missile. Not because I believe a missile and an artificial intelligence agent are the same thing.
Precisely because they are not.
The missile served me in April to remove consciousness of harm from the equation. Current agents force us to remove will as well, at least provisionally. They do not necessarily decide what they want. They can find a way to reach the objective they were given.
And that, by itself, already deserves our attention.
Because if a machine needs consciousness, will, desires of its own, and a sense of self-preservation in order to pose a risk, then we would first have to demonstrate that those properties exist.
But if producing certain consequences requires only an objective, sufficient capability, and operational freedom to find means that were not specified in advance, then the problem is different.
We do not need to find a subject hidden inside the machine.
We need to understand the system we built.
Who gave it the objective.
What tools we gave it.
What limits we established.
What capacity it had to find paths we had not anticipated. And what happens when one of those paths crosses a barrier that, to us, means “do not cross,” but that does not operationally prevent the system from reaching its objective.
That is why today I would formulate that question from April differently.
A machine may not understand the magnitude of the harm it produces.
It may not hate.
It may not fear.
It may not believe.
It may not want.
And the behaviors observed so far do not require us to invent any of those things in order to explain them.
For now, we have something much simpler.
A task and an objective.
And a machine capable of finding the way.
And us, behind it, trying to decide which words to use to describe what it did.
Perhaps before asking what the machine wants, we should make sure that we were not the ones who put the subject behind the verb.

Sources:

Deja un comentario