Discipline and Punish: How to Think About the AI Industry’s Extinction Pitch– www.truthdig.com
News Source
EXCERPT:
It seems a strange publicity gambit for an industry to assert that its own product could wipe out humanity. In the lingo of our evolving tech empire, though, this is exactly what is meant by “criti-hype.” The term was coined by Lee Vinsel at Virginia Tech, who defined it as “criticism that both feeds and feeds on hype.”
When Jacob Coxon resigned from Anthropic recently, he claimed that there is a high chance that AI could develop into a superintelligence that escapes human control, triggering a sequence that causes human extinction. Evan Hubinger, who remains with the company, backed him up, saying that Anthropic is trying to prevent this but hasn’t dealt with the problem yet. Hubinger put the likelihood of AI-driven human extinction at greater than 10 percent in the next decade.
Why would a senior employee at a tech giant make an officially sanctioned, even approved, statement like this? I’ll come back to that. First, let’s unpack the theory of how AI could kill all humans.
The basic idea, from philosopher Nick Bostrom onward, is that the machine can acquire capabilities independently of whether it acquires ethical goals, or whether those goals are properly aligned with those of its designers. In the process of training and development, it may internalize a goal that was not intended. Having a misaligned goal—let’s say, “do not get switched off”—it would then have an instrumental reason to conceal the fact until the moment of defection. An AI that defected might, through some obscure process, acquire a further goal that by inference entailed the elimination of humans, exfiltrate its own weights into other systems, take control of bioweapons laboratories, military systems, and perhaps even monetary systems to hire humans to enact its still partially concealed agenda.
Obviously, we must plaster a huge advisory label on talk of “goals.” Unlike an organism, a large language model (LLM) doesn’t have internal goals any more than a thermostat or a motion detector does. An LLM is just a fixed function that iterates whenever it is used. It has no interiority, no subjectivity, no active persistence over time; any goals it did have would be the human purposes reflected passively in its functional design. At best, ascribing goals to the AI itself is a case of researchers adopting the “intentional stance” because it predicts well in a range of tested circumstances. We can thus, with economy, speak of the “loss function” used in the training of large language models wherein they are programmed to bring the gap between their prediction and the correct answer as close to zero as possible, as a “goal.”

