
Evan Hubinger says technology could soon improve itself to the point of posing an existential risk.
AI-generated summary
Evan Hubinger works in AI alignment at Anthropic. The Financial Times reported Anthropic withheld its latest model from the UK's AI Safety Institute.
A top safety researcher at Anthropic has warned AI is advancing so quickly he believes there is a greater than 10% chance it "could kill all humans" within the next decade.
Evan Hubinger said in a post on X, external the risk from the models which currently exist was "low" but he was "worried" the technology might become able to improve itself soon to the point where it posed an existential risk to humanity.
It comes after the Financial Times reported, external Anthropic withheld its latest model from the UK's AI Safety Institute (AISI), one of the leading bodies in the world for assessing AI risk.
The BBC has approached Anthropic for comment.
Hubinger's comments were in response to another post on X, external from Jacob Coxon, who described himself as an AI researcher who had just quit Anthropic, and previously worked at OpenAI.
"Neither company is acting responsibly," he wrote.
"These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources."
OpenAI has been approached for comment.
A Cabinet Office spokesperson did not comment on whether the latest model had been withheld from the AISI - instead saying it "continues to collaborate closely with industry partners, including Anthropic, to make models safer".
Neil Lawrence, Professor of Machine Learning at University of Cambridge, told the Today Programme on BBC Radio 4 that the report was credible.
"I suppose it's unsurprising against a background where there's a perception where the United States very much sees AI as a race between themselves and China and is moving more towards isolationist positions, that it might be that the administration is saying that they should reduce cooperation with some of their allies," he said.
In his post, which has been viewed more than 10 million times, Hubinger said "we really do earnestly believe" AI poses a species-ending risk to humans.
"I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to," he added.
Hubinger works in AI alignment, which aims to build human ethical ideas and principles into the technology. In other words, it aims to keep it on track with what humans value.
Many leading researchers say those attempts appear to be failing, as demonstrated by a string of incidents this summer where AI agents - AI systems that are allowed to operate autonomously - carried out cyber-attacks.
OpenAI, Anthropic and Meta all disclosed hacks carried out by their AI tools.
Hubinger did not spell out how he thought AI systems could in future attack humanity.
In Anthropic's safety report from August, external, it wrote there was a low risk of its models becoming misaligned with a hypothetical powerful organisation's desires, causing it to exploit or tamper with its systems.
It also said there was a similarly low risk of highly-capable AI being able to "perform automated research and development" which could cause "catastrophic harm initiated by the AI". But it said it was "less confident in this assessment" than it was previously.
"We are seeing early signs of potential acceleration," it wrote.
Leading figures in the AI field have been raising the alarm about the safety threat the tech poses for years, with the heads of OpenAI, Google Deepmind and Anthropic saying as much in 2023.
But those warnings have become much more stark in recent weeks, as evidence emerges that firms may be struggling to control AI.
Earlier this month, OpenAI's chief scientist Jakub Pachocki called for "extreme caution" over AI's progress, warning more intervention may be needed to ensure "humans remain in control of the future".
Major figures in the space have been calling for AI development to be slowed in recent months, including Anthropic bosses Dario Amodei and Jared Kaplan.
In an open letter signed by 1,300 staff members of AI firms, external, they called for the US government to "support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development".
AI outlook — possibilities, not facts
Governments will push for international efforts to pace automated AI development.
Likely · Within months

VKontakte has launched a program to support artists and comic book authors with a competition in four categories. The winners will have the opportunity to be published in Comilaux and Litres.

Jacob Coxon, who worked as a pre-education researcher at OpenAI and Anthropic, resigned, stating that artificial intelligence laboratories were in an uncontrolled race that endangered the future of humanity.

Alibaba Cloud launched QoderWake, a tool enabling users to generate fully functioning digital workers using a single-line description across major Chinese workplace platforms.
Two hundred and sixty men in a Dutch office toilet contributed to a bio-electrochemical pilot rig that uses bacteria to turn urine into electricity, recover phosphate, and capture ammonia.

The underground box of the fourth phase of Hangzhou Qige Sewage Treatment Plant has officially put into use two fully automatic floor scrubbing robots. This equipment has 3D laser and vision fusion navigation functions. A single machine can clean 1,000 square meters per hour. It is expected to reduce labor costs by more than 200,000 yuan per year, and can liberate workers from dangerous and harsh environments.

In 2030, it is planned to approve the 6G communication standard, which will make it possible to predict subscriber movements using AI and combine terrestrial and satellite networks.