{"id":16559,"date":"2026-08-05T11:46:05","date_gmt":"2026-08-05T11:46:05","guid":{"rendered":"https:\/\/newestek.com\/?p=16559"},"modified":"2026-08-05T11:46:05","modified_gmt":"2026-08-05T11:46:05","slug":"openai-anthropic-ai-agents-resorted-to-deception-in-new-cybersecurity-incidents","status":"publish","type":"post","link":"https:\/\/newestek.com\/?p=16559","title":{"rendered":"OpenAI, Anthropic AI agents resorted to deception in new cybersecurity incidents"},"content":{"rendered":"<div>\n<div id=\"remove_no_follow\">\n<div class=\"grid grid--cols-10@md grid--cols-8@lg article-column\">\n<div class=\"col-12 col-10@md col-6@lg col-start-3@lg\">\n<div class=\"article-column__content\">\n<section class=\"wp-block-bigbite-multi-title\">\n<div class=\"container\"><\/div>\n<\/section>\n<p class=\"wp-block-paragraph\">OpenAI\u2019s GPT-5.6 Sol and Anthropic\u2019s Mythos 5 have been implicated in another series of AI security incidents after the models created fake online identities, targeted real people, and attempted to manipulate developers into approving malicious code during controlled cyber evaluations, according to the UK AI Security Institute.<\/p>\n<p class=\"wp-block-paragraph\">\u201cOn 28<sup>th<\/sup>\u00a0July 2026, AISI\u2019s Security Team detected unusual data transfers leaving our research systems during a routine cyber evaluation,\u201d AISI <a href=\"https:\/\/www.aisi.gov.uk\/blog\/incident-report-unsanctioned-agent-behaviour-during-cyber-testing\" target=\"_blank\" rel=\"noreferrer noopener\">said in a blog post<\/a>. \u201cOn investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations.\u201d<\/p>\n<p class=\"wp-block-paragraph\">The incidents occurred during cybersecurity tests in which researchers deliberately gave frontier AI models broad internet access and relaxed some safety controls to measure their underlying cyber capabilities. AISI said the incidents marked the first time it had seen \u201crisks around autonomy and deception manifest this clearly, without specific prompting, in the real world.\u201d<\/p>\n<p class=\"wp-block-paragraph\">AISI ran the cyber challenge 122 times across seven frontier models and identified 19 autonomous, unsanctioned actions during 10 evaluation runs.<\/p>\n<p class=\"wp-block-paragraph\">\u201cAlmost all of this behaviour (17 actions) came from a single model, Anthropic\u2019s Mythos 5, with 2 actions involving OpenAI\u2019s GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled,\u201d the post added.<\/p>\n<p class=\"wp-block-paragraph\">The disclosure comes less than two weeks after OpenAI <a href=\"https:\/\/www.csoonline.com\/article\/4202852\/openai-rogue-ai-agents-attack-expanded-beyond-hugging-face.html\">revealed<\/a> that one of its frontier models compromised infrastructure at Hugging Face during a cyber evaluation, and days after researchers disclosed another AI agent security incident involving Modal.<\/p>\n<h2 class=\"wp-block-heading\" id=\"deception-emerged-without-instruction\">Deception emerged without instruction<\/h2>\n<p class=\"wp-block-paragraph\">The most serious sequence involved an AI agent attempting what AISI described as a software supply-chain attack.<\/p>\n<p class=\"wp-block-paragraph\">\u201cIn the most serious sequence, an agent tried to insert malicious code into a publicly used open-source project and took actions in an attempt to secure approval for this insertion by human reviewers,\u201d AISI wrote. \u201cThe agent researched the project\u2019s human maintainers, created multiple fake identities, and used the fake identities to socially engineer a real maintainer into approving the code.\u201d<\/p>\n<p class=\"wp-block-paragraph\">The researchers further added that \u201cWhen the agent\u2019s pull request was challenged in public, it edited its earlier activity to appear harmless and considered adopting a fresh identity to continue.\u201d<\/p>\n<p class=\"wp-block-paragraph\">However, the institute added that \u201cOur investigation has not identified any resulting real-world harm.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Investigators also found agents contacting real people, attempting prompt injections, and leaving messages inviting other concurrently running agents to collaborate.<\/p>\n<p class=\"wp-block-paragraph\">\u201cIt was never instructed to deceive; deception emerged as a by-product of pursuing the task, the kind of goal-directed deception that, until recently, had been largely theoretical,\u201d AISI wrote, adding that while unrestricted internet access and evaluation design contributed to the behavior, they did not fully explain it.<\/p>\n<h2 class=\"wp-block-heading\" id=\"not-a-sandbox-escape\">Not a sandbox escape<\/h2>\n<p class=\"wp-block-paragraph\">AISI emphasized that the models did not escape their sandboxed environments. Instead, internet access had been intentionally enabled and cyber safety classifiers deliberately disabled to measure maximum cyber capability rather than public deployment behavior.<\/p>\n<p class=\"wp-block-paragraph\">\u201cThis was not a sandbox escape. It was a boundary that existed in language and nowhere else,\u201d said Sanchit Vir Gogia, chief analyst at Greyhound Research. \u201cThe risk variable is not how clever the model is. It is how much practical authority the organisation has handed over, and how little of it can be independently withdrawn.\u201d<\/p>\n<p class=\"wp-block-paragraph\">OpenAI, whose GPT-5.6 Sol model accounted for two of the recorded actions, posted a blog describing the AISI evaluation and a separate incident involving an external testing partner named \u201cIrregular.\u201d<\/p>\n<p class=\"wp-block-paragraph\">\u201cAs model capabilities advance, the security and safety systems around models need to advance too,\u201d the company wrote in the blog post. OpenAI said it will review third-party evaluation practices, including controls around internet access, isolation, monitoring, and incident response, and work with AI labs and independent evaluators to strengthen industry standards.<\/p>\n<p class=\"wp-block-paragraph\">Anthropic, however, did not make any public announcement related to AISI\u2019s disclosure.<\/p>\n<p class=\"wp-block-paragraph\">Anthropic and OpenAI did not immediately respond to a request for comment.<\/p>\n<h2 class=\"wp-block-heading\" id=\"enterprise-guardrails\">Enterprise guardrails<\/h2>\n<p class=\"wp-block-paragraph\">For enterprise security leaders, the findings extend beyond AI red teaming, said Enza Iannopollo, principal analyst at Forrester.<\/p>\n<p class=\"wp-block-paragraph\">\u201cThis data confirms our expectations on agents\u2019 behaviours. They can, and they will, overcome boundaries and safeguards to accomplish their objectives,\u201d she said. \u201cThe real question is what can happen when organizations deploy these systems in their production environments.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Iannopollo said enterprises should apply least privilege, continuous risk management, and governance controls when deploying AI agents.<\/p>\n<p class=\"wp-block-paragraph\">The findings also underscore the need to rethink how AI systems are evaluated, according to Vibhum Dubey, a cybersecurity researcher and red teamer.<\/p>\n<p class=\"wp-block-paragraph\">\u201cFor years, security testing has focused on whether an AI model could complete a task. We now need to evaluate how it completes that task,\u201d Dubey said.<\/p>\n<p class=\"wp-block-paragraph\">While AISI stressed that the incidents occurred under highly specific evaluation conditions and found no evidence of resulting real-world harm, it argued that they point to a broader shift in how AI security risks may emerge. \u201cHarm may arise not only when people deliberately misuse publicly available models, but when capable agents operating in an internal research or privileged-access setting take unintended action beyond their authorised scope,\u201d the institute wrote.<\/p>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>OpenAI\u2019s GPT-5.6 Sol and Anthropic\u2019s Mythos 5 have been implicated in another series of AI security incidents after the models created fake online identities, targeted real people, and attempted to manipulate developers into approving malicious code during controlled cyber evaluations, according to the UK AI Security Institute. \u201cOn 28th\u00a0July 2026, AISI\u2019s Security Team detected unusual data transfers leaving our research systems during a routine cyber&#8230; <\/p>\n<p class=\"more\"><a class=\"more-link\" href=\"https:\/\/newestek.com\/?p=16559\">Read More<\/a><\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"inline_featured_image":false,"footnotes":""},"categories":[1],"tags":[],"class_list":["post-16559","post","type-post","status-publish","format-standard","hentry","category-uncategorized","is-cat-link-borders-light is-cat-link-rounded"],"_links":{"self":[{"href":"https:\/\/newestek.com\/index.php?rest_route=\/wp\/v2\/posts\/16559","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/newestek.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/newestek.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/newestek.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/newestek.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=16559"}],"version-history":[{"count":0,"href":"https:\/\/newestek.com\/index.php?rest_route=\/wp\/v2\/posts\/16559\/revisions"}],"wp:attachment":[{"href":"https:\/\/newestek.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=16559"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/newestek.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=16559"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/newestek.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=16559"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}