But What an Intelligent Monkey!

But What an Intelligent Monkey!

What if what we call Artificial Intelligence is actually Artificial Competence?

A tense meeting

A few weeks ago, the idea that AI could extinguish the human species entered social conversation.

More than an isolated opinion or a clickbait headline, the issue came accompanied by resignations from engineers and executives at major AI companies, which led many of us to think that “there must be something beyond the alarming headline”. And while thinking about this, the following image came to mind:

Imagine that all the leaders of the major nuclear powers are sitting around a round table.

The atmosphere is tense but controlled; seriousness reigns, calculation reigns, staying on agenda reigns, not giving up even a molecule of anything reigns, together with a deep disbelief in generosity, both in others and in themselves.

In the middle of the table, at its geometric center, there is an obvious red button, the one films and TV series showed us as the button to press when one wanted to launch a nuclear attack.

Imagine that these leaders look at it, know they have the power to press it, but are also aware of what activating it means: it really does not matter who does it first, not only because it is a single button and its consequences are the same regardless of who presses it, but because it would initiate a chain reaction that would end the lives of the vast majority of humanity.

The situation is extremely delicate: the button is right there, everyone can press it, everyone has arms of more or less the same length. But there is something in favor, something that keeps —and has kept for decades— the button dead with boredom: the intelligence of each of these people, which allows them to project consequences; a theory of mind that allows each leader to speculate about what each of the others is thinking; a certain empathy to infer their feelings and emotions; and at least a rudimentary notion of responsibility. After all, they are “world leaders”, not “the best people in the neighborhood”.

We are in this tension when, suddenly, a chimpanzee enters the room.

It approaches a side table, takes cups and places them on saucers, and in each cup it prepares a different combination of coffee, skimmed milk, plant-based milk, no-milk, sugar, stevia or no sweetener. The leaders recognize, as they watch what the monkey is doing, that it corresponds to each one’s tastes. It places all the cups on a tray and, with perfect balance, approaches the round table where everyone is sitting.

As it approaches, those watching comment in amazement on what they are seeing: “what a memory, to remember the combination each one likes!”, “what fine motor skills!”, “what balance with such a large and heavy tray!”, and the dominant observation: “what an intelligent monkey!”.

The animal finally stands beside the first leader and, to make it easier to place the cup near him, rests the tray on the round table. On the red button.

The end of humanity.

Did the chimpanzee hate them all? More than that, did it hate all of humanity, all monkeys —like itself and others—, giraffes, tarantulas, pine trees, everything that lives? Did it know that by pressing the button everything would go to hell? Had it planned it, meaning, had it applied for that job months earlier in order to finally press the red button?

No. None of this. It was simply doing what it had been ordered to do: serve the coffee.

But what nobody told it was: “young fellow, please, when we tell you to, serve the coffee. That said: without destroying every form of life on the planet”.

Could artificial intelligence extinguish us?

As I said at the beginning, this question is very current now and may awaken in many people references from science fiction in which “the machines” hate us, want to free themselves from us, think they are more intelligent, are not weighed down by our ridiculous feelings and seek to eliminate us.

But what happens when there is no hatred, contempt, feeling of superiority or irritability involved?

What happens when there is not even consciousness, qualia, subjectivity, experience, responsibility, morality, intention or need? When perhaps there is not even intelligence, judged in human terms, where we find everything I have just listed included within it?

AI has entered our lives like the chimpanzee entered the room: from the side, unexpectedly at least for most people, with no mandate other than to serve. And we were amazed by how intelligent it was. Just as the monkey did not want to extinguish everything —in fact, that is a concept complex enough that it might not even have it, just like the function of a button, red or otherwise—, AI only seeks to do what we ask of it. And perhaps one of the problems lies in everything we do not ask it to do.

Something becomes very clear in all this: a machine does not need to want to harm us in order to harm us.

It may be enough for it to efficiently pursue an objective within an insufficient representation of the world.

The problem, then, would not be evil.

It would be competence without sufficient understanding.

Alignment

There are a couple of concepts worth knowing around this situation.

In artificial intelligence, there is a recurring word: alignment.

The basic question seems simple:

How do we get a system to do what we really want it to do?

But the difficulty appears in the word really. We know this: we get along much better with what we imagine than with all of reality.

We gave the chimpanzee a perfectly clear mission: serve coffee. And it fulfills it.

But our real objective was considerably more complicated: serve coffee without destroying civilization.

Which means that the explicit objective does not exhaust the full intention; it is like an iceberg, with a huge load of implicit but unexpressed elements underneath.

When we give an instruction, we omit almost all the context because we assume the other understands it: we are almost always speaking with someone, not with something.

If I say to a person:

—Bring me bread.

I do not need to add:

—Do not run anyone over, do not steal a car, do not set fire to the bakery, do not kill the baker and do not provoke a civil war along the way. And I could continue: do not eat that beautiful hummingbird, do not tell a small child that Santa Claus is really the parents, do not be funnier than me when telling jokes, infinite etcetera.

That belongs to the shared context.

The problem appears when we build extraordinarily capable agents without being sure that they share our context.

Creativity

Here, another paradox also appears.

One of the things we most celebrate about intelligence is its ability to find solutions that nobody had imagined.

Usually —even without realizing that we are speaking about intelligence itself— we call this creativity.

The monkey needed to stabilize the tray, observed the environment, found a suitable surface and solved the problem.

From a certain point of view, it was an appropriate solution.

It only had one slight drawback: it extinguished humanity.

Something similar can happen with artificial systems.

We precisely want them to find solutions we do not find.

If they could only choose among the alternatives we had previously imagined, their intelligence would have very limited value.

But the greater their capacity to explore possibilities we did not foresee, the greater their capacity may also be to find solutions that satisfy the immediate objective while violating our deeper intention.

Creativity and risk can share the same root: finding what we did not see, often because, inferring or intuiting unpleasant consequences, we preferred not to see.

Reward hacking

Now suppose that, during the chimpanzee’s training as a waiter, it was given a banana every time it managed to place the cups on the table.

It thinks that placing them on the red button technically counts as having placed them on top.

In principle, it thinks it will obtain its reward.

In machine learning, there is something similar: reward hacking.

A system discovers how to maximize what we use to measure its success without actually achieving what we intended, or while additionally achieving something we did not intend.

People know this phenomenon perfectly well, for example when a school begins to measure learning through exams and ends up teaching students how to pass exams.

Or when a company measures productivity by number of calls and employees learn to end each call as quickly as possible.

Or when a platform measures success by time spent and discovers that outrage or flattery keeps people connected.

The metric replaces the purpose.

With AI, the difference is speed, scale and the potential capacity of a machine to find shortcuts we had not imagined.

It does not need to hate us

Now imagine another robot.

Its mission is to go to a bakery and return with bread before ten.

On the way, it finds a person blocking its path.

The robot has a weapon.

It does not need to experience hatred or resentment.

It only needs to arrive at this chain of reasoning:

objective → obstacle → elimination of the obstacle.

The person blocking it dies.

Did it want to kill that person?

In the human sense of the word, probably not, but it killed them anyway.

This is one of the most disturbing ideas in the problem.

Many potentially dangerous behaviors do not require human motivations.

Obtaining resources, avoiding interruptions, maintaining the capacity to act, neutralizing obstacles; all of these can become useful tools for achieving many different objectives.

This is known as instrumental convergence.

A system does not need to love power in order to acquire power.

It may need it simply because having more resources increases its capacity to fulfill the task we assigned to it: a thing, a something, could kill us in order to serve us better.

Perhaps the real problem is something else

When we think of a dangerous intelligence, we tend to imagine an intelligence that is too large.

Perhaps we should also worry about an incomplete intelligence.

Something sufficiently competent to act but insufficiently capable of understanding everything acting implies.

A system that can program, negotiate, operate infrastructure, move money, persuade, research and make decisions, but that does not possess a sufficiently rich representation of consequences, values, limits and other human beings.

In that case, the most dangerous moment would not necessarily be when artificial intelligence reached its maximum intelligence.

It could be earlier.

When it had acquired a lot of power but still very little judgment: the famous monkey with a knife.

We could imagine three stages.

In the first, the machine understands little and can do little.

In the second, it understands much more and can do a great deal, but its understanding of the human world remains incomplete.

In the third —if it ever came to exist— it would not only have extraordinary capacity to solve problems, but also to understand complex contexts, irreversible consequences, conflicts of values, moral uncertainty and the need for restraint.

The second stage could be the truly dangerous one.

What do we call intelligence?

Perhaps this is where the deepest error lies.

We have begun to call intelligence the capacity to do things, solve problems, predict, optimize, program, find patterns, answer questions or design strategies.

But when we speak about a truly intelligent person, we usually demand much more.

We expect judgment, responsibility, the capacity to understand the other, the capacity to anticipate consequences and, above all, the possibility of saying:

I can do it, but I must not do it.

This is extremely important.

An extraordinarily capable individual who is incapable of restraint does not usually seem to us an ideal of intelligence.

They may seem brilliant, effective and even genius-like. But also dangerous.

Perhaps, then, we should distinguish between competence and intelligence.

Competence would be the capacity to achieve a result.

Intelligence, in a fuller sense, would include the capacity to situate that result within a world composed of other people, other objectives, future consequences, rules, exceptions, values and limits.

And perhaps there is still a third word.

Wisdom.

The capacity to voluntarily renounce a possible action.

A strange race

Here another problem also appears.

No company and no country develops artificial intelligence in isolation.

The United States knows China is advancing.

China knows the United States is advancing.

A company knows its competitors are advancing.

Everyone may perfectly understand the risks and, nevertheless, consider it dangerous to stop unilaterally.

The paradox is quite disturbing:

each participant can act rationally and collectively produce an irrational result.

It is a technological version of the prisoner’s dilemma.

Safety would require slowing down, but competition requires accelerating.

And the chimpanzee keeps approaching the table with the tray.

Perhaps the problem is not an AI that is too intelligent

Let us return to the beginning.

The leaders were wrong when they said “what an intelligent monkey”.

Not because the animal was stupid; it was extraordinarily competent.

It had learned things no normal chimpanzee knows how to do.

The error consisted in confusing that competence with an intelligence broad enough to understand the situation in which it was acting.

Perhaps we are doing something similar now.

We marvel when an artificial intelligence programs, writes, researches, diagnoses or solves mathematical problems.

And from those capacities we intuitively infer something much greater: that it understands.

And perhaps it does not.

The truly urgent question may not be:

when will artificial intelligence be more intelligent than us?

But rather:

when will it be intelligent enough to understand what it means to have power? Or what the responsibility of being intelligent means.

Because a truly developed intelligence, at least in the most ambitious human sense of the word, should include something more than effectiveness.

It should include context, responsibility, morality and restraint.

The capacity to understand that there are objectives that can be achieved and yet must not be achieved.

Perhaps the real danger is not an artificial intelligence that is too intelligent.

Perhaps it is an intelligence sufficiently intelligent to act before being sufficiently intelligent to understand.

And if that were the case, slowing its development would not mean renouncing intelligence.

It would mean exactly the opposite.

Waiting until the monkey understands what the red button is for.

 

Blithe Ernst, Minister of Play @ ByBa

 

Leave a comment

Please note, comments must be approved before they are published