Breaking
RUA musician from Kursk won 12 million rubles in the Magnificent 8 lotteryFRThe 79th Besançon International Music Festival opens with 24 hours of free concertsRUFamous dancer Cindy Loani Matute Morales was kidnapped and burned in HondurasCNThe iPhone 18 Pro series will be launched simultaneously, and Black Panther members can enter an hour in advance.RUCloudy weather without precipitation and temperatures up to plus 19 degrees are expected in Moscow this weekendDEFire in Upper House Warehouse requires the use of high-altitude rescuersRUThe son of an M.D.U. assistant professor confessed to killing his mother in Zhukovka-3RUMore than 90 UAVs and missiles were destroyed while repelling an attack on the Rostov regionDEDan Lahav on AI Risks: Exponential Development and Security Incidents at IrregularTRFenerbahçe Started the Champions League with a 1-1 Draw with RomaRUA musician from Kursk won 12 million rubles in the Magnificent 8 lotteryFRThe 79th Besançon International Music Festival opens with 24 hours of free concertsRUFamous dancer Cindy Loani Matute Morales was kidnapped and burned in HondurasCNThe iPhone 18 Pro series will be launched simultaneously, and Black Panther members can enter an hour in advance.RUCloudy weather without precipitation and temperatures up to plus 19 degrees are expected in Moscow this weekendDEFire in Upper House Warehouse requires the use of high-altitude rescuersRUThe son of an M.D.U. assistant professor confessed to killing his mother in Zhukovka-3RUMore than 90 UAVs and missiles were destroyed while repelling an attack on the Rostov regionDEDan Lahav on AI Risks: Exponential Development and Security Incidents at IrregularTRFenerbahçe Started the Champions League with a 1-1 Draw with Roma
BackAnthropic Claims It Thwarted Multiple Malicious AI Operations Using Claude Models
Anthropic Claims It Thwarted Multiple Malicious AI Operations Using Claude Models
Developing
Al Jazeera12 minutes agoTech2 min read

Anthropic Claims It Thwarted Multiple Malicious AI Operations Using Claude Models

Quick Look

Anthropic reported blocking several malicious uses of its Claude AI models, including attempts to develop missile guidance in Yemen, Russian and Chinese cyber-espionage campaigns, Iranian influence operations, and a recruitment scheme targeting Uyghurs in Syria, while also addressing internal safety concerns and legal disputes with the Pentagon.

AI-generated summary

Why It Matters

Anthropic has faced scrutiny over AI safety, including a Pentagon blacklisting for refusing to remove safeguards against autonomous weaponry, which a California judge later ruled unlawful. The company continues to navigate tensions with Washington while asserting its models have been misused in various malicious operations.

Font size

Anthropic AI claims to have thwarted multiple malicious operations using its Claude models, ranging from cyber-espionage and weapons design to mass surveillance campaigns.

On the conventional weapons front, the company alleges in a new report that it intervened in northern Yemen, blocking an effort to deploy Claude for missile guidance software, including a guided rocket and a long-range ballistic missile.

Recommended Stories

list of 3 items

list 1 of 3US judge blocks Pentagon blacklisting of AI firm Anthropic

list 2 of 3Sony, Warner Music sue Anthropic, saying it pirated songs to train its AI

list 3 of 3US pushes looser approach to AI regulation, while EU pushes new law

end of list

AI-assisted missile and rocket development

According to Anthropic’s threat report, the operators used Claude “in place of human software engineers”, reportedly assigning different instances of the model specific roles to write missile-guidance and flight-control software.

While internal safeguards blocked many requests, Anthropic admitted several slipped through. The operators avoided detection by obscuring their ultimate goals and breaking tasks across separate sessions, so no single prompt gave away the operation.

The company said it has no evidence the group managed to field a working weapon, though Anthropic claimed the operators appeared to have conducted an unsuccessful test-fire.

The report stated that the company banned the accounts involved and “shared threat information with public- and private-sector partners to mitigate risks posed by the actors”.

State-linked cyber-espionage

The report alleges that a Russian-linked espionage operation bearing the hallmarks of Midnight Blizzard (or APT29), which Anthropic said relied on automated AI workflows to run nearly the entire operation, handling everything from phishing and setup to data theft against Ukrainian, European and diplomatic targets, including drone makers.

Separately, the company said it disrupted a Chinese operation run by university students in Hunan province, who used Claude “as the engineering and orchestration layer” of an offensive programme targeting government and corporate networks across the Middle East, Europe and Southeast Asia.

In both cases, Anthropic said it banned the associated accounts and deployed additional monitoring to detect similar activity.

Identifying targets, including in Syria and Iran

Anthropic also alleges that it identified and removed three Iranian state-aligned accounts using Claude to run covert influence and psychological operations. Each operation was tied to a named Iranian propaganda institution, including the Islamic Culture and Communications Organisation and a Mashhad seminary command room distributing content aligned with the Islamic Revolutionary Guards Corps’ (IRGC) narratives.

In another instance, the report alleges that state-aligned groups used Claude for an industrial-scale operation, directing the model to generate structured profiles that mapped targets by location, demographics, political leanings, and confidence scores.

“The most operationally mature case,” Anthropic noted in its report, was detected when a China-aligned account “with no Arabic language skills” used Claude to run a “multi-day recruitment operation to infiltrate Uyghur targets in Syria”. Anthropic said, “the model drafted outreach in the regional dialect” and “translated replies in real time”.

Earlier this week, Anthropic revealed yet another incident of an AI model gaining unauthorised access to external systems involving an early version of Claude Opus 4.6, shortly after former company researcher Jacob Coxon publicly resigned over safety concerns.

“The people building AI earnestly believe that it could kill us all by the end of the decade,” Coxon warned in a post on X. A further Anthropic scientist, Evan Hubinger, later chimed in, saying Coxon was correct.

The researchers’ warnings have led a growing number of US lawmakers to call for new rules to govern AI systems.

Anthropic said it is investigating the recurring issues and breaches across the incidents and has engaged an independent research firm to review them.

Relationship with Washington

Relations between Washington and the AI company remain fraught following a contentious standoff over ethical guardrails. Earlier this year, the Pentagon blacklisted the company as a supply chain risk after it refused to drop safeguards against using its technology for autonomous weaponry and domestic surveillance.

Anthropic challenged the decision in California, where a judge ruled last month that the US Department of Defense had acted unlawfully in issuing the designation. Yet despite the bitter legal battle and public friction, the Pentagon has reportedly deployed the firm’s Claude models in military missions in Iran and Venezuela.

The report arrives at a critical juncture for the company as it seeks to restore full standing within the US defence industrial base following the Pentagon’s blacklisting.

What to Watch

AI outlook — possibilities, not facts

  • Anthropic will implement stricter monitoring and usage policies for its Claude models to prevent future misuse.

    Likely · Within months

  • US lawmakers will advance new legislation to regulate AI systems in response to safety concerns raised by former Anthropic researchers.

    Likely · Within months

Open Questions

  • What specific safeguards failed to block the malicious Claude usage?
  • How effective was the shared threat information with partners in mitigating ongoing risks?
  • What are the findings of the independent research firm investigating recurring breaches?
  • Has the Pentagon continued using Claude models in military operations despite the legal ruling?

Related Topics

This article was originally published by Al Jazeera.

Related Stories

LG Denies Smart TVs Record Ambient Conversations, Cites Explicit User Activation
Developing·

LG Denies Smart TVs Record Ambient Conversations, Cites Explicit User Activation

LG has denied allegations that its smart televisions record ambient conversations or listen while on standby, stating that voice data is only recorded when users press and hold the voice button or activate the 'Hi LG' wake word after enabling Far-Field voice recognition. The company says audio is immediately deleted if the wake word is not recognized. The denial follows an investigative report by Gamers Nexus, Level1Techs, and independent researchers claiming LG TVs track users, harvest network data, and archive background noise up to 12 meters away, even when the microphone switch is off. Researchers allege unmatched wake-word interactions may be stored locally as plain text and uploaded later, and that TVs scan home networks for device information. LG maintains that features like ACR, voice recognition, and interest-based advertising require explicit user opt-in and can be disabled in settings.

Al Jazeera
2 min read
Grindr settles lawsuit over sharing of sensitive user data with advertisers
Developing·

Grindr settles lawsuit over sharing of sensitive user data with advertisers

Grindr has agreed to pay £26 million to settle a lawsuit filed by around 12,000 UK users who alleged the app shared highly sensitive personal data, including HIV status, with advertisers. While denying liability, the settlement highlights broader concerns about how apps collect, share, and monetize user data through complex digital ecosystems, often without meaningful user awareness or consent.

Deutsche Welle
2 min read
More on this topicanthropic