UK researchers find GPT-5.5 matches Mythos Preview in cybersecurity benchmarks
OpenAI's latest release scored higher on expert tasks but failed the same critical simulations as Anthropic's restricted model, prompting criticism of fear-based marketing tactics

Researchers from the UK's AI Security Institute (AISI) have published new findings indicating that OpenAI's GPT-5.5 achieves performance levels comparable to Anthropic's Mythos Preview in cybersecurity evaluations. The study, which forms part of a long-term assessment of frontier models, challenges the narrative that the restricted release of Mythos represented a unique or singular breakthrough in defensive capabilities. Instead, the data suggests that these advanced security features are a byproduct of general improvements in long-horizon autonomy and reasoning found across the sector.
The AISI report highlights that GPT-5.5 outperformed Mythos Preview on high-level 'Expert' tasks, achieving an average success rate of 71.4 per cent compared to 68.6 per cent for the competing model. In a specific challenge involving the construction of a disassembler to decode a Rust binary, the OpenAI model solved the problem in 10 minutes and 22 seconds without human assistance. This performance places GPT-5.5 ahead of Mythos Preview, which managed only two successful attempts out of ten on 'The Last Ones' test range, a simulation designed to mimic a complex data extraction attack on a corporate network.
Despite these gains in specific domains, GPT-5.5 failed the 'Cooling Tower' simulation, a test involving the disruption of power plant control software. This result mirrors the failure rate of all previously tested models, including earlier iterations of both OpenAI and Anthropic systems. The AISI notes that while the new model succeeded in three out of ten attempts on the 'The Last Ones' range, it could not breach the more difficult Cooling Tower scenario, reinforcing the idea that current limitations in autonomous cyber operations remain consistent across the industry.
OpenAI CEO Sam Altman has publicly addressed the implications of these findings, criticising what he describes as fear-based marketing surrounding restricted model releases. Comparing the rhetoric to selling a bomb shelter after constructing a bomb, Altman argued that while dangerous models will inevitably require special distribution methods, the hype does not necessarily correlate with unique security threats. He acknowledged that Mythos is a capable model but emphasised that the need for restricted access stems from the inherent risks of the technology rather than a singular anomaly.
The release of GPT-5.5 marks a continuation of OpenAI's strategy for managing frontier models, following a similar limited launch for GPT-5.4-Cyber in February. The initial rollout of the new GPT-5.5-Cyber variant is expected to be restricted to critical cyber defenders in the coming days, adhering to the same 'Trusted Access for Cyber' pilot program framework. This approach allows verified security researchers and enterprises to study the models for legitimate defensive work while mitigating potential misuse.
The AISI has been running frontier AI models through 95 different Capture the Flag challenges since 2023, covering areas such as reverse engineering, web exploitation, and cryptography. These comprehensive tests have provided a consistent baseline for comparing the capabilities of different systems, revealing that the gap between the leading models is narrowing. As the industry moves forward, the focus remains on understanding the true limits of autonomous reasoning in high-stakes environments rather than relying on marketing narratives of singular breakthroughs.
