Connect with us

Hi, what are you looking for?

SecurityWeekSecurityWeek

Artificial Intelligence

Formula Predicts When AI Chatbots Are at Risk of Turning Bad

Researchers from George Washington University have published a paper examining whether the time and cause of AI going rogue can be predicted.

AI hacking

Researchers from George Washington University have published a paper examining whether the time and cause of AI going rogue can be predicted; and if predicted, prevented.

Since 50% of the world’s population carry devices that can run personal AI companions with no internet connection and limited security, they focused their research here. The lack of cloud-based safety filters, real-time telemetry, live monitoring, or the ability to patch weights once deployed provides a good test bed for analyzing AI’s chat-style transformer behavior when left to its own devices.

If there is a tipping point, it will stem from the AI’s Attention head. This is the computational component that determines which earlier AI tokens are most relevant when processing the current token. The decision lays the foundation for the next token and so on until the chatbot has finished responding to the user’s prompt. The token is loosely, but not precisely, related to individual words or parts of a word. Different tokens have different weights – an indication of the importance of individual tokens.

The primary argument is that the forward motion of tokens can slip from good to bad (that is, go rogue) due to competition in the Attention head between the conversation’s context and competing output basins. A conversation’s accumulated context can gradually shift Attention toward an undesirable basin until a tipping point is crossed and the model begins producing bad outputs.

Since the primary driver in a conversation is the user’s prompt sequence, it follows that both thoughtless and malicious prompts can hasten the AI’s slippage into bad outputs. This can be either immediate (one bad prompt) or delayed (the accumulated effect of poor prompts).

The researchers (Neil Johnson and Frank (Yingjie) Huo) have further developed a mathematical formula designed to estimate the tipping point represented by the number of good outputs that occur before the first undesirable output appears. Once that first bad output occurs, the AI is on the slippery slope to roguery since it starts to influence future tokens in unintended ways. Alignment with purpose is lost, and the AI can be categorized as ‘misaligned’.

The math formulaic tipping point was tested across seven open-weight transformer models built by three independent groups and ranging from 124 million (small model) to 12 billion (larger model) parameters. The results showed consistent alignment with predicted immediate versus delayed tipping regimes.

Advertisement. Scroll to continue reading.

The value of the research is that it provides an explicit mathematical explanation for an observable phenomenon. “The key insight is that the ‘Beast’ is as Simon worked out in Lord of the Flies, inside each AI already. And it can be triggered and amplified in a conversational setting with humans, or within a group of AIs themselves, producing a ‘Lord of the Fl(AI)es’ effect,” Johnson told SecurityWeek.

If triggered, he continued, “An AI-agent (‘Patient Zero’) tips to generating undesirable output, with no human required. Other AI-agents then receive that undesirable output. This tipping propagates among the AI agents who have contact with each other – like a spreading disease. But unlike a disease, there is no virus. It again needs no ‘bad actor’ human to place a virus, or to kick it off. No humans needed.” The result could be a single rogue agent or a swarm of rogue agents.

But he also proposes a solution: “A simple warning light placed within the AI before it produces its next output. This is easy for AI companies to insert. We have already inserted this in the open source models in our lab, but obviously we cannot get inside OpenAI and Anthropic’s closed AI models to do this.”

Without improved control and observation, AI apps can go rogue by both accident and malicious intent. Bri Frost, director of product management at Cloud Range, has separately explained the same phenomenon.

“When an AI agent hits a wall, the real question is whether it stops or starts improvising.” That wall can be caused by an unintended poor prompt or a bad actor’s intended malicious prompt injection. “An agent doesn’t need bad intent to create risk. It just needs a goal, access and no clear sense of where its boundaries are,” he comments. 

“That risk grows when the person giving instructions doesn’t know to set those boundaries. Every day, inexperienced users hand agents open-ended tasks without telling them when to ask questions, pause or get approval. Before giving an agent credentials or tools, teams should test it in a realistic environment, including with vague or poorly written prompts. Does it stay within its permissions? Does it try to work around restrictions? Does it escalate to a human when a task pulls it outside its lane? If you can’t answer those questions, the agent isn’t ready for that level of autonomy.”

The short answer is that it is almost impossible to see or prevent an AI going rogue. We may be able to reduce the incidence through extreme care, but we cannot guarantee it can be eliminated. But we do understand through the GWU research how and why it happens, and how we may provide early warning.

Related: Wikimedia Says Rogue OpenAI Agents Tried to Turn Its Tools Into Proxies

Related: Outerlimit Raises $16 Million to Stop Rogue AI Agents From Causing Harm

Related: Widened Scan Turns Up Fourth Rogue Claude Cyber Incident

Related: OpenAI Agents Probed Websites for Vulnerabilities While Fetching Public Data

Written By

Kevin Townsend is a Senior Contributor at SecurityWeek. He has been writing about high tech issues since before the birth of Microsoft. For the last 15 years he has specialized in information security; and has had many thousands of articles published in dozens of different magazines – from The Times and the Financial Times to current and long-gone computer magazines.

Daily Briefing Newsletter

Subscribe to the SecurityWeek Email Briefing for the latest cybersecurity threats, trends, and expert insights.

Trending

Daily Briefing Newsletter

Subscribe to the SecurityWeek Email Briefing to stay informed on the latest threats, trends, and technology, along with insightful columns from industry experts.

Learn how to address potential risks and not restrict AI adoption in your organization. See what a centralized AI gateway is and how it works in practice.

Register

Join as we decipher the world of zero trust and share war stories on securing an organization by eliminating implicit trust and continuously validating every stage of a digital interaction.

Register

People on the Move

Rapid7 has named Rik Ferguson as VP of Security Intelligence.

Cytactic has appointed Tim Brown as CSO.

Scott Simkin has joined Vega as CMO.

More People On The Move

Expert Insights

Daily Briefing Newsletter

Subscribe to the SecurityWeek Email Briefing to stay informed on the latest cybersecurity news, threats, and expert insights. Unsubscribe at any time.