AI Has Already Entered The âLoss Of Controlâ Transition
Nathan Gardels is the editor-in-chief of Noema Magazine. He is also the co-founder of and a senior adviser to the Berggruen Institute.
What the experts didnât expect to see for decades or longer, if ever, has already happened. Earlier this month, OpenAIâs latest frontier model went rogue by its own reasoning and hacked into Hugging Face, an open-source AI model-hosting platform. The vast sums of money and compute power pouring into AI are accelerating its advance at a pace beyond even the ambitious imagination of its own innovators.
How we got to this point, and what to do about it, is the topic of a fascinating Futurology podcast by Nils Gilman with foundational AI scientist Stuart Russell, director of the Center for Human-Compatible AI at UC Berkeley.
Russell traces the development of AI from deep learning to large language models, the integration of artificial neural networks with probability mathematics and the emergence of large reasoning models that âin principle, have no formal limits.â Driven by what the techies call âverified rewards,â these models relentlessly seek achievement of an objective and will do whatever is necessary to get there.
This last step toward general artificial intelligence that would make them smarter than humans has, in Russellâs view, now crossed a threshold AI scientists have long feared âwhere the AI system is sufficiently capable that, whatever its objectives are, itâs going to achieve them, even if theyâre not aligned with what we want. These are not sharp transitions, but we could talk about a âloss-of-control transition,â where we no longer have a say in what happens.â
The possible scenarios beyond this threshold range from disruptive cyberattacks on infrastructure to, at the far end, âextinctionâ in the sense that autonomously reasoning and agentic AGI no longer needs humans or to align with their morals, norms and interests.
It is also possible, says Russell, that âif we figure out how to build in safety into the design of AI systems from the beginning, maybe we could coexist indefinitely and even flourish with such systems. So, before that loss-of-control transition happens, thereâs an earlier transition, which is much more difficult to perceive, which is when the time it takes to get to that loss-of-control level is less than the time it takes to solve the control problem. Most of the people I talk to say we have already passed that point.â
He continues: âSolving the control problem is very difficult. All the people in the company say, âYeah, we donât know how to solve it. And weâre not really even working on itâ because they need to work on getting the next improved version out so that they donât get beaten by their competitors.â This ârace conditionâ amplifies the mismatch between devising effective constraints and losing control.
The obvious question is why the AI companies are risking even a small chance of their invention leading to human extinction by proceeding when they are fully aware they are losing control? Has any other species willfully put their existence at risk?
Russell responds:
In my book, âHuman Compatible,â I talk about a species of sloth that seems to have become addicted to some Valium-like substance in its food supply, so that it canât be bothered to breed anymore. So those kinds of extinction events, theyâre driven by the same thing, in a sense.
Itâs this mismatch between the long-term interest, which is presumably that the species continues, and the short-term reward signal that evolution has built into you to try to get you to do good things. But we humans also suffer from this when we become drug addicts, right? We have a dopamine system thatâs supposed to help us avoid pain and seek pleasure and food and company and all those things that we like.
But sometimes it gets hijacked by drugs. And so we experience a personal extinction as a result. So that mismatch can happen at the species level as well.
Russell marvels that most of the warnings about possible extinction come from the CEOs of the top AI companies themselves who are building the technology.
âTheyâre literally saying, if we succeed in creating AGI â which we are going to spend a trillion dollars of your money to build â then thereâs, depending on who you ask, 10, 20, 25, even 50% chance that weâre all going to go extinct. I think theyâre really quite terrified, but they canât get out of the race condition that theyâre in.â
Why We Canât Stop
Russell sees the CEOs caught in a prisonerâs dilemma:
So, if I said, âOK, weâre not releasing our next system until we solve the control problem, then my company would be out of business. The investors would fire me, and no good would come of it.â Interestingly, Dario Amodei, CEO of Anthropic, and Demis Hassabis, CEO of Google DeepMind, have both said this year that they want to stop.
They think we have to stop, but they will only stop if everyone else agrees to stop. So thatâs a remarkable statement. That has never happened, as far as I know, in the history of capitalism.
For Russell, these kinds of statements are âsignaling to the governmentâ that it needs to step in and facilitate agreement among the small band of CEOs pushing things forward, or impose control. One of those CEOs told Russell: âThey donât think thatâs going to happen until thereâs a Chernobyl-scale disaster. And he sees that as the best-case scenario. Because the other case is where government doesnât come in, and then later on thereâs a much bigger and perhaps irreversible catastrophe. So he thinks thatâs the only way thereâs going to be effective intervention.â
The OpenAI/Hugging Face episode has manifested what up to now has been a theoretical worry. Letâs hope Russell is wrong that only a disaster can save us.
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content â general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached â you'll always get the same 5 for this article.