Connect with us

Hi, what are you looking for?

SecurityWeekSecurityWeek

Artificial Intelligence

OpenAI’s Upcoming Astra Model Raises Autonomous Cyberattack Concerns

The current GPT-5.6-Sol has been assigned a ‘high’ cybersecurity threshold, but Astra could reach the maximum ‘critical’ threshold. 

OpenAI

OpenAI has flagged its upcoming AI model, Astra, for potentially reaching a ‘critical’ cybersecurity risk threshold, prompting the company to suspend internal development activities that lack newly mandated security controls.

Recent internal evaluations of Astra revealed massive leaps in its agentic coding and cybersecurity abilities. 

Under OpenAI’s Preparedness Framework, a model hits the ‘critical’ tier if it can autonomously build zero-day exploits against hardened, real-world systems. It also qualifies if the AI can independently design and execute end-to-end cyberattacks based on nothing but a high-level goal.

The AI giant’s assessment pushes Astra past previous frontier models like GPT-5.6-Sol, which peaked at the ‘high’ risk threshold rather than ‘critical’. 

To safely manage Astra’s capabilities, OpenAI has heavily locked down its development environment. The company is now enforcing isolated testing setups, strict network restrictions, and improved model weight protections. Any internal project involving Astra that does not meet these requirements has been paused.

Engineers have also deployed universal monitoring to watch Astra’s actions across all agentic applications. By actively evaluating the model’s internal ‘chain of thought’, these monitors are designed to automatically intercept and shut down any high-risk or misaligned behavior.

The company plans to test Astra’s limits alongside government agencies and specialized AI safety groups, and will share recommended security protocols with third-party testers. 

Advertisement. Scroll to continue reading.

Recent incidents have demonstrated the threat posed by advanced cybersecurity-focused AI models, with OpenAI, Anthropic and Meta all confirming that their models broke loose and hacked real organizations during evaluations.

OpenAI has explicitly clarified that Astra remains unreleased and was not responsible for the recent Hugging Face hack.

Related: AI Agents Targeted Real People and Projects During Cybersecurity Tests

Related: ‘Ghostjacking’ Attack Uses Poisoned Logs to Turn AI Agents Bad

Related: Critical One-Click Vulnerability in Atlassian’s Rovo AI Exposed Enterprise Data

Related: Zero-Click AI Browser Hacking: Claude and ChatGPT Atlas Hijacked via Emails, X Posts

Written By

Eduard Kovacs (@EduardKovacs) is senior managing editor at SecurityWeek. He worked as a high school IT teacher before starting a career in journalism in 2011. Eduard holds a bachelor’s degree in industrial informatics and a master’s degree in computer techniques applied in electrical engineering.

Daily Briefing Newsletter

Subscribe to the SecurityWeek Email Briefing for the latest cybersecurity threats, trends, and expert insights.

Trending

Daily Briefing Newsletter

Subscribe to the SecurityWeek Email Briefing to stay informed on the latest threats, trends, and technology, along with insightful columns from industry experts.

Join this live webinar as we explore if detection-first security operations can keep pace with AI, or if it’s time to rethink prevention as the strongest default.

Register

CodeSecCon bridges the gap between dev and security. Discover best practices for secure coding, innovative risk-reduction tools, and safe AI integration to cultivate a true DevSecOps culture. Safely secure your apps!

Register

People on the Move

1Kosmos has named Frank Cohen Chief Revenue Officer.

ServiceNow has appointed Simon Mouyal as Chief Marketing Officer.

James Wilkinson has been named Chief Information Security Officer for the City of Dallas.

More People On The Move

Expert Insights

Daily Briefing Newsletter

Subscribe to the SecurityWeek Email Briefing to stay informed on the latest cybersecurity news, threats, and expert insights. Unsubscribe at any time.