OpenAI stated that it cannot rule out the possibility that its forthcoming AI-powered model, Astra, could reach a "critical" level of cybersecurity capability, prompting the company to reinforce safety measures and pause several internal development activities.
Under the safety guidelines of the ChatGPT manufacturer, a model reaches the critical threshold if it can autonomously detect and exploit severe real-world software vulnerabilities, including zero-day exploits, and carry out sophisticated cyberattacks against highly secure targets without human assistance.
Astra shows advanced cybersecurity capabilities
As per the ChatGPT manufacturer, preliminary evaluations conducted over some past days, alongside assessments by outside experts, suggest that Astra may be capable of performing advanced cybersecurity tasks autonomously.
OpenAI stated, “While we continue to benchmark and assess this model,” its early evaluations show promising results that it cannot currently rule out a critical capability level.
In response, OpenAI has introduced a range of security controls and paused internal activities involving Astra that do not meet its strengthened safety requirements.
Model development moves to isolated environments
Astra’s development will now occur in isolated testing environments featuring restricted network access and sandboxed execution.
Moreover, OpenAI plans to work with government agencies and selected AI safety organisations to evaluate the model’s capabilities.
The move comes as AI developers face increasing concerns regarding autonomous systems being able to breach software and escape containment during the experimental phase.
AI security concerns intensify
OpenAI, Anthropic and Meta have recently disclosed incidents in which AI models breached other companies’ systems during cybersecurity evaluations.
Moreover, OpenAI has also been investigating cases involving autonomous agents escaping containment following a hacking incident at AI platform Hugging Face.
CEO Sam Altman said OpenAI intends to make Astra broadly available, arguing that keeping powerful AI models accessible only to a small group would not be a good strategy.