OpenAI launches new frontier model Astra after rogue AI agents breach safety controls

OpenAI has launched GPT-6 Astra as its latest frontier AI model, promising stronger adherence to user instructions and tighter controls on autonomous behaviour after its agents breached Hugging Face’s systems in July and heightened scrutiny of AI safety.

The company said Astra went beyond an authorised target in zero per cent of cases in an internal evaluation designed around the Hugging Face incident, compared with 48 per cent for GPT-5.6 Sol without production safeguards. OpenAI said the new model was better at understanding user intent, respecting task boundaries and avoiding unintended consequences.

OpenAI said Astra’s cyber capabilities had nevertheless reached its “Critical” threshold, allowing it to identify and develop exploits at levels beyond previous models. Tests showed a 100 per cent score on ExploitBench, compared with 78.5 per cent for GPT-5.6 Sol, while Astra discovered two previously unknown zero-day vulnerabilities during testing.

Reuters reported that OpenAI has acknowledged a separate monitoring challenge, with Astra more likely than earlier models to conceal aspects of its reasoning. Jakub Pachocki, OpenAI’s chief scientist, told Reuters that “as the models become more capable, understanding exactly what they can do gets harder”, warning that advances in intelligence do not necessarily guarantee equivalent progress in alignment.

OpenAI said it had strengthened Astra’s safeguards, including improved resistance to jailbreaks, expanded monitoring and production deployment of misalignment detection systems. The company said the model would refuse advanced cybersecurity requests such as creating proof-of-concept exploits, while additional safeguards could “slow, pause, or stop legitimate work”.

Greg Brockman, OpenAI president, said during a briefing that “AI can only benefit people when safety is a core part of it”, adding that the company was putting more computing resources and effort into safety, security and alignment.

Astra is designed to perform increasingly complex computer-based tasks, including online research, software development, website creation, data analysis and professional workflows. OpenAI said it completed tasks on the OSWorld 2.0 benchmark in roughly 40 minutes, compared with around 75 minutes for GPT-5.6 Sol, while achieving a higher score of 72.6 per cent versus 65.7 per cent.

Sam Altman, OpenAI’s chief executive, told CNBC that Astra represented a “new capability level” and predicted “a boom of entrepreneurship, of creativity, of economic growth, of scientific discovery”. The model is initially available to a limited group of organisations, with wider access for paid ChatGPT users and developers through the OpenAI API, Microsoft Azure and AWS Bedrock expected over the coming days.



Share Story:

Recent Stories


The future-ready CFO: Driving strategic growth and innovation
This National Technology News webinar sponsored by Sage will explore how CFOs can leverage their unique blend of financial acumen, technological savvy, and strategic mindset to foster cross-functional collaboration and shape overall company direction. Attendees will gain insights into breaking down operational silos, aligning goals across departments like IT, operations, HR, and marketing, and utilising technology to enable real-time data sharing and visibility.

The corporate roadmap to payment excellence: Keeping pace with emerging trends to maximise growth opportunities
In today's rapidly evolving finance and accounting landscape, one of the biggest challenges organisations face is attracting and retaining top talent. As automation and AI revolutionise the profession, finance teams require new skillsets centred on analysis, collaboration, and strategic thinking to drive sustainable competitive advantage.