The biggest yield

The latest warnings that AI could kill us all are arriving just as OpenAI and Anthropic prepare enormous IPOs. The fear is recursive self-improvement: an AI builds its successor, which becomes better at building the next one.

I can see why people are cynical about the timing. But the questions of safety and control deserve serious answers.

Jacob Coxon has resigned from Anthropic, accusing the labs of “gambling with our lives”. Jasmine Wang has warned about the danger of speeding towards RSI. Behind the tweets and resignations is a question I want answered.

How would an AI actually kill humanity?

A useful starting point is an HSE leaflet about runaway chemical reactions. A reaction produces heat. The heat makes the reaction go faster, producing still more heat. If cooling cannot keep up, the vessel can rupture and people can die. The process that was supposed to make something useful has destroyed the plant.

That is a loss of control over a process that feeds itself.

The fear with AI has that shape. Systems accelerate their own development and spread their activity faster than people can supervise or stop them. The harm would come through what they can reach: computer networks, essential infrastructure, weapons. We depend on those systems to stay alive.

Nuclear fission makes the distinction particularly clear. A power station controls a chain reaction to produce useful electricity. A bomb is designed to release destructive energy. A bigger explosion does not mean we built a better power station.

The Manhattan Project scientists even investigated whether a bomb could ignite the atmosphere. They calculated that it would not. The weapons still destroyed cities. Getting the calculation right did not settle what the achievement meant for the people beneath the bombs.

What worries me about the AI race is the appeal of that role: the brilliant nerd making history, squeezing out the biggest yield, solving the problem better than anybody else. There is a particularly ugly kind of ego in wanting to win that contest regardless of how disgusting the consequences are for everyone outside the lab.

For AI, keeping the system useful and under human control has to be part of what we mean by improvement. Alignment is part of the engineering problem. If the result destroys the people it was supposed to help, it has failed. However clever the solution.

This is why the word improvement in recursive self-improvement matters. Better at what, and for whom? The goal has to be aligned with human purposes, and each round has to preserve that alignment and our ability to stop the process. These are part of the engineering problem. Without them, we have not succeeded.

Otherwise it is the paperclip optimiser. A machine gets better and better at turning the world into paperclips. An arbitrary objective, pursued with extraordinary competence, produces pointless waste. Like plastic floating in the ocean: we get more of it, it gets everywhere, and we have to live with the consequences. Making that process faster hardly counts as improvement.

I worry that some people are seeing a ghost in the machine. Get the chain reaction going and goodness will somehow emerge from the intelligence itself. Call it the singularity. Call it superhuman, an Übermensch, or “machines of loving grace”. However stupid its purpose, however misaligned its actions, however much it pollutes our human world, they imagine that crossing this threshold will make it good.

Treating the human world as disposable in anticipation of that transformation is nihilism and insanity of the first order. People with these millenarian beliefs become more dangerous the more power they have over AI development. It matters who chooses the goals, authorises the next training run and decides to release the result.

Sam, Dario: what exactly is being improved, and for whom?

Photo by Naja Bertolt Jensen on Unsplash.