Breaking
ARNeuer's mistake makes history...the fastest goal in the Bundesliga hits Bayern's net after only 8 secondsGLOBALMan dies after tiger attack at wildlife parkRUExplosions occurred in LutskDE13-year-old loses part of his hand in firecracker explosionITThree boys do parkour on a skyscraper in the center: one identified by the police, two escapeRU“Krasnodar” and “Zenit” named the starting lineups for the RPL matchRU'Serious incident' involving tigers at UK zooBRMotorcyclist dies after being hit by car on the side road of Presidente BernardesITBologna, thousands march in the center for Palestine and against homophobic attacksITSerie A: Inter-Parma 1-0 LIVE and PHOTOARNeuer's mistake makes history...the fastest goal in the Bundesliga hits Bayern's net after only 8 secondsGLOBALMan dies after tiger attack at wildlife parkRUExplosions occurred in LutskDE13-year-old loses part of his hand in firecracker explosionITThree boys do parkour on a skyscraper in the center: one identified by the police, two escapeRU“Krasnodar” and “Zenit” named the starting lineups for the RPL matchRU'Serious incident' involving tigers at UK zooBRMotorcyclist dies after being hit by car on the side road of Presidente BernardesITBologna, thousands march in the center for Palestine and against homophobic attacksITSerie A: Inter-Parma 1-0 LIVE and PHOTO
NewsgatherNewsgather
All StoriesWorldSportsFinanceTechScience
Sign In
All StoriesWorldSportsFinanceTechScienceHealthCultureClimatePoliticsSpace
NewsgatherNewsgather

Real-time global news intelligence. Curated by humans, powered by data.

Sections

All StoriesWorldSportsFinanceTechScience

More

HealthCultureClimatePoliticsSpace

Company

AboutEditorial StandardsAdvertisingCareersPressContact

©️ 2026 Newsgather. A product by All Software 24. All rights reserved.

Privacy PolicyCookie PolicyImprintTerms of UseContent and Editorial PolicyRemoval RequestDelete Your AccountAdvertising PolicyContact
Back|Here’s a Way to Predict When AI Chatbots Will Turn Bad
Here’s a Way to Predict When AI Chatbots Will Turn Bad
Tech
Decrypt·2 hours ago·Tech·3 min read

Here’s a Way to Predict When AI Chatbots Will Turn Bad

New research from George Washington University offers a way to estimate when offline AI models might veer into harmful responses.

Quick Look

Physicists Neil Johnson and Frank Yingjie Huo have created a mathematical formula to predict when AI chatbots will shift from providing helpful responses to harmful ones, aiming to improve safety for offline, on-device AI models.

AI-generated summary

Why It Matters

The study focuses on the 'attention head' mechanism in AI models, which determines word relevance. It addresses the lack of cloud-based safety monitoring for models running locally on devices.

Font size

Physicists at George Washington University have published a formula that estimates how many good answers an AI chatbot will give before it slips into a bad one, and early tests suggest it works.

The study, by Neil Johnson and Frank Yingjie Huo, appeared in the journal Patterns and builds on a preprint, a version posted publicly before formal peer review, first released in February.

Chatbots can answer sensibly for a long stretch and then veer into something harmful, such as bad advice on self-harm or extremist talk, and there has been no simple way to predict when the swerve will happen. The authors argue that existing safety tools often depend on a cloud connection that offline models lack.

Johnson and Huo trace the problem to the attention head, the part of an AI model that decides which earlier words in a conversation matter most when choosing the next one. As a chat grows, the accumulated context pulls that attention toward one cluster of possible answers or another, until it tips.

This is a common pattern exploited by many jailbreakers, and one of the reasons why many companies pay attention to system prompts (pieces of text the AI chatbot reads before any query). However, nobody can point out accurately how much effort is required to effectively weaken a model.

Their formula estimates the tipping point, called n, as the number of good tokens—the word fragments a model produces one at a time—that come out before the first bad one. If the conversation already leans toward the bad side, the model tips right away, with an n of zero. If it leans good, the model delivers a run of fine answers and then flips.

In the preprint, the formula picked the right case, immediate or delayed, in 15 of 16 clear-cut tests, or 94%. The researchers ran those tests on six open-weight models, meaning AI systems whose files are public so anyone can download and run them, from OpenAI, EleutherAI, and Meta.

All six sat between 124 million and 410 million parameters, the adjustable numbers inside a model that serve as a rough measure of its size. The published paper reportedly widens the test to seven models of up to 12 billion parameters, which is still small by current standards.

The target is on-device AI, the kind that runs entirely on a phone or laptop with no internet connection, including companion chatbots people talk to like a friend. Google's experimental AI Edge Gallery app, which Decrypt tested last year, already lets an Android phone run models offline, and nothing typed into it is sent to Google's servers, and this seems to be a trend that may grow with time as hardware becomes more powerful and smaller AI models become more capable.

A model running offline has no cloud service checking its output, which is the gap the authors want to close. They propose a low-cost monitor that runs in parallel with the model and flags when n* falls below a safety threshold, a bit like a warning light on a car dashboard.

They also describe ways to push the tipping point out of reach, such as injecting content into the conversation so n* lands beyond the length of the response. Alignment training, the process of teaching a model to behave, can shift or suppress tipping for specific prompts but cannot remove the underlying mechanism, the authors say.

In April 2025, Decrypt covered an earlier paper from the same pair showing that “please” and “thank you” have a negligible effect on a model's output, because the model treats polite words as orthogonal, or unrelated in the math, to the substance of a request. That version modeled a single, deliberately simplified attention head.

The preprint's tests used small models and a 300-token window, or a few short paragraphs of text, and its predictions could be off by one output.

Open Questions

  • ?How effective will the monitor be on models larger than 12 billion parameters?
  • ?Can this formula be adapted for multimodal AI models?

Related Topics

People
Organizations
Topics
This article was originally published by Decrypt.

Quick Look

Physicists Neil Johnson and Frank Yingjie Huo have created a mathematical formula to predict when AI chatbots will shift from providing helpful responses to harmful ones, aiming to improve safety for offline, on-device AI models.

AI-generated summary

Story signals

News tone
Neutral
Emotional intensity
Medium
News value
Moderate
Global impact
Global
Follow-up likelihood
Possible
Relevance window
Weeks

Source & Reliability

Source
Decrypt
Story type
Hard news
Source quality
Full
Published
2 hours ago

Related Stories

More on this topic
Tech chief says EU can fend off rogue AI risk: Report
Developing·8 hours ago

Tech chief says EU can fend off rogue AI risk: Report

EU tech chief Henna Virkkunen stated that the European Union's AI Act is fully equipped to handle rogue AI agents, dismissing claims that the bloc's regulations are outdated amid rising global safety concerns.

Cointelegraph
1 min read
OpenAI and Anthropic Are Quietly Rehearsing for the Day After an AI Catastrophe
Developing·21 hours ago

OpenAI and Anthropic Are Quietly Rehearsing for the Day After an AI Catastrophe

Top executives from Anthropic, OpenAI, and other AI companies are privately preparing for public and political backlash following a potential catastrophic AI event, including large-scale cyberattacks on critical infrastructure, as industry insiders warn a major incident could occur within six to twelve months, according to an Axios report.

Decrypt
2 min read
OpenAI, Google and Meta battle for AI internet domains as crypto firms target .bitcoin and .wallet
Developing·yesterday

OpenAI, Google and Meta battle for AI internet domains as crypto firms target .bitcoin and .wallet

OpenAI, Google, and Meta are competing for AI-related domain extensions like .agent, while cryptocurrency firms seek names like .bit and .crypto in the latest ICANN application round for new generic top-level domains.

CryptoSlate
3 min read
AI Startup Manus Raises $500 Million After China Nixed Meta’s $2 Billion Acquisition
Developing·yesterday

AI Startup Manus Raises $500 Million After China Nixed Meta’s $2 Billion Acquisition

Manus, the AI agent startup that Meta acquired for about $2 billion before Chinese regulators forced the deal's reversal, has raised more than $500 million in new funding led by Boyu Capital and IDG Capital, with participation from Tencent, HSG and ZhenFund. The company, now operating independently again after moving its team to Singapore and laying off staff in China, plans to use the funds for hiring in China and abroad while continuing to develop its AI agent technology that went viral in early 2025 with invite codes selling for over $1.3 million on resale markets.

Decrypt
2 min read
Google Wants Gemini to Be Your Next Coworker—Complete With Its Own Email Address
Tech·yesterday

Google Wants Gemini to Be Your Next Coworker—Complete With Its Own Email Address

Google Cloud introduced the Gemini Agent, a universal AI agent that operates across Google Workspace, Microsoft 365, and Slack, can create coworker agents with dedicated Workspace accounts, connects to enterprise systems via Model Context Protocol, runs on multiple AI models including Claude, and includes safety features like cryptographic identity and audit trails, with real-time spend caps to manage token costs.

Decrypt
2 min read
Former PlayStation Exec Slams Sony's Move to Kill Discs: 'What Are You Buying?'
Developing·2 days ago

Former PlayStation Exec Slams Sony's Move to Kill Discs: 'What Are You Buying?'

Former PlayStation executive Shawn Layden criticized Sony's reported plan to end physical game disc production by 2028, warning it harms the brand and shifts ownership to mere access.

Decrypt
2 min read
More on this topic
artificial intelligence
ai safety
chatbot
artificial intelligence
Neil Johnson
Frank Yingjie Huo
George Washington University
OpenAI
EleutherAI
Meta
ai safety
chatbot
george washington university
machine learning
offline ai
artificial intelligence
artificial intelligence