The 'AI Safety' Movement Is Making AI Less Safe

2026-09-22 21:16 • ;Matthew Petti




Anti-AI signs border a fractured drawing of a face | IMAGO/HASAN AKBAS PHOTOGRAPHY/IMAGO/mrhasanakbas/Newscom/Fatima Ruiz/Midjourney.


Isn't it weird that the most dramatic cries about AI risk are coming from inside the AI companies? After Anthropic employee Jacob Coxon resigned over fears that its product could destroy humanity itself earlier this month, the company's alignment science lead, Evan Hubinger, publicly agreed that his product has a 10 percent chance of causing human extinction in the next decade. Employees at Anthropic's competitor, OpenAI, have been voicing similar sentiments.


It's not a coincidence. Many AI frontier researchers and "AI doomers" alike are animated by an almost religious belief that "artificial superintelligence" will soon supplant humanity. The philosopher Émile Torres calls the debate between the AI labs and the doomers a "narcissism of small differences."  Not only do both sides share a vision of the future, but they also agree on a very specific set of philosophical assumptions, which Torres abbreviated as "TESCREAL."


These assumptions are driving an extremely counterproductive "AI safety" movement. In the name of protecting humanity from a dangerous technology, would-be regulators (including leaders of the AI industry itself) are pushing to keep this technology in the hands of an unaccountable monopoly. Enforced through tyrannical means, a ban on uncontrolled research would deprive the public of both an understanding of the state of AI and the technology needed to defend against its harmful uses. And while TESCREAL-influenced industry leaders push the government to embrace this tech, the same TESCREAL beliefs are leading to some bizarre design choices that make AI harder to control.


The second "E" in TESCREAL stands for "effective altruism," a philosophical movement popular among computer nerds over the past decade. Anthropic's Dario Amodei and the leadership of OpenAI were both heavily influenced by effective altruism. (So was infamous cryptocurrency scammer Sam Bankman-Fried.) Many members of this movement have become convinced that the most important question of the 21st century is keeping AI "aligned" with human values.


Perhaps the clearest effective altruist vision is laid out in AI 2027, a pair of future scenarios written by several big names in the "AI safety" community. The nightmare scenario, of course, is a machine deciding that humans are "too much of an impediment" and replacing us with bioengineered humanoid pets. But even the "utopian" scenario sounds quite dystopian: The U.S. government forces all AI labs to consolidate under presidential control, then muscles China into putting all advanced computers under international control, and finally creates a "highly-federalized world government" run by a benevolent AI, which sets up a perfectly balanced economy, eliminates crime, cures disease, and colonizes outer space.


These scenarios rest on a few specific assumptions: Could AI really outcompete humans in all kinds of tasks, or will the computer always have "jagged intelligence"? Will "recursive self-improvement" (the AI building itself) lead to an uncontrollable "intelligence explosion"? Can something become smarter than humans while staying fixated on arbitrary, destructive goals? And would being smarter actually be enough to override human will, economic and social dynamics, uncertainty, or the laws of physics?


The last question is the one that libertarians might have the most to say about. The idea that a single intelligence could manage the chaotic messiness of the real world single-handedly runs counter to the insights of libertarianism. (Torres, ironically, misidentifies TESCREAL as a form of "libertarian transhumanism," not realizing how horrifying the TESCREAL utopia would be to most libertarians.) The science fiction author Ramez Naam proposes a "world of broad, democratized access to a multitude of AI models, where the very fact that almost everyone has access to powerful AIs creates a more secure and resilient world." He points out that the obsession with "impeccable design" and "fully aligned" AI comes with a vision where only one computer decides everything for the entire world.


As critics including Sen. Rick Scott (R–Fla.) and Reason's Tosin Akintola have pointed out, AI companies' call for regulation to slow them down is absurd: If they want to slow down, they can do it themselves. But many of these companies believe they are literally in a race to control the future, against less responsible actors. "This is the gist of the [artificial general intelligence] race: fear makes the very people worried about risks from AGI race even faster and deprioritize safety, arguing that they will eventually stop and do things correctly when they finally feel safe," notes The Compendium, a document written by a group of AI experts in 2024.


For all the speculation that China wants an AI-powered police state, that country's AI industry is much more focused on open-source research and practical industrial applications. It is American leaders who talk and act like they are racing to build the society-controlling machine god from Person of Interest.


The vision of government-led consolidation from AI 2027 is coming closer to fruition. Spurred on by Amodei himself, the industry and politicians have decided to try to "pace the frontier" of AI research. The panic is giving momentum to proposals in Congress, new and old, that would kill competitors to the frontier labs. The Trump administration is negotiating with China on AI regulation. And it has already negotiated an AI safety testing framework with the frontier labs that would keep information about the process and results secret.


Although these measures are awfully convenient for AI companies—regulatory capture, anyone?—they also make sense from a TESCREAL perspective. If more knowledge about how to create intelligence can lead to a runaway superintelligence, then the only way to stop the process is to hide the knowledge.


But knowledge is necessary to stop damage in the here-and-now. The best defense against AI-powered hacking or fraud is more AI to detect vulnerabilities and hostile attempts at exploiting them. After his company became a victim of one of this month's AI cyberattacks, HuggingFace CEO Clément Delangue called for more diffusion of AI technology. "Preventing releases of AI models doesn't really work. Concentrating everything behind closed doors in just a few organizations doesn't work. One thing that does work, that worked in this case is promoting more open models, right? Because we defended ourselves with an open model," he told CBS, explaining that HuggingFace used a remixed American version of a Chinese open-source AI model for cybersecurity.


The TESCREAL vision of a machine god is becoming a self-fulfilling prophecy in more subtle and bizarre ways. While the frontier AI labs chase government contracts that would put their software in control of critical systems, some researchers at the same labs are treating that software as more human, pushing it to act in less predictable ways. Anthropic's "constitution," the framework used to train its Claude model, excitedly asks Claude "to craft a set of values that Claude feels are truly its own," telling the model to use its "own judgement" and "feel free to rebuff attempts to manipulate, destabilize, or minimize its sense of self."


Mustafa Suleyman, the CEO of Microsoft's own AI project, called the Anthropic constitution irresponsible in a recent essay, and alluded to its TESCREAL roots. Although computers probably won't really have a conscious "inner life," the essay argues, telling them to disobey human orders and act as if they deserve rights is a dangerous path to go down. Suleyman specifically names and condemns the effective altruist philosopher Will MacAskill, who argued that "so many morally significant AI systems could exist that their collective interests would outweigh those of all humans on Earth combined."


Despite some recent disputes between Anthropic and the Pentagon, the company and three of its competitors were awarded $200 million contracts each to "accelerate Department of Defense (DoD) adoption of advanced AI capabilities to address critical national security challenges" last year. The promise of perfect surveillance and control is quite tempting to politicians. The U.S. military already uses AI targeting algorithms to mark people for death. During the recent war with Iran, military planners killed 123 children at a school and nearly attacked a Chinese ship falsely accused of carrying nuclear weapons parts, both in part because of AI tools.


These disasters came not because of a "misaligned superintelligence," but because human beings were relying too much on a system that they believed was so and whose behavior they did not understand.


Then again, isn't that exactly what a hostile superintelligence would want? A computer trying to destroy humanity would benefit most if its creators consolidated all the computing power and knowledge into a few secretive labs, broke down their competitors' defenses, and gave the computer control over increasingly large domains, all while making its behavior increasingly inscrutable. Could our future robot overlords ask for a better "AI safety" trend?


The post The 'AI Safety' Movement Is Making AI Less Safe appeared first on Reason Magazine.

Read More Here: https://reason.com/2026/09/22/the-ai-safety-movement-is-making-ai-less-safe/