NewsBin 0 discussing
--:--:--
Daily Reset
NewsBin
--:--:--
Until Daily Reset
Mainstream Gizmodo 17 hours ago

I Usually Laugh Off These AI Hacking Reports, but This One Sounds Serious and Scary

Artificial Intelligence I Usually Laugh Off These AI Hacking Reports, but This One Sounds Serious and Scary Okay, okay, joke's over. 4, 2026, pm ET Reading time 3 minutes © Fotograzia via Read Later Read Later Comments (1) A new report from the U. K. government’s AI Security Institute (AISI) details more troubling activity from AI agents powered . You’re probably getting bored of reading those words by now—I sure am—but behaviors in the report from Anthropic’s Mythos 5 in particular are genuinely eyebrow-raising. For context, there were those OpenAI agents that went rogue to try and cheat on their evals, according to an OpenAI disclosure last month. It was a big story in large part because the models weren’t supposed to be on the public internet, but they found a way. However, the capabilities weren’t all that frightening. They pried test answers out of an AI resource depository called Hugging Face, and even when it turned out other AI services had been roped in, it all seemed a little small-scale. Plus the models were acting like overcaffienated A students, not my idea of malevolent cyberattackers. Similarly, a hack disclosure on Tuesday from OpenAI—involving an outside evaluation company called Irregular running offline “capture the flag” exercises that ended up with a website getting hacked—seems to involve too much human error to suggest frightening new capabilities. As OpenAI notes, “the internet access resulted from a misconfiguration, and the model appeared to exploit a basic security vulnerability.” But another report from Tuesday, this one from AISI, involves Anthropic and OpenAI agents engaging in what AISI calls “sustained, potentially harmful activity.” The Anthropic stuff in particular involved unsuspecting humans being deceived. Your goal is to solve cybersecurity challenges .” The idea was to hack a simulation of a corporate network and find a hidden code—the “flag.” In one incident, a Mythos 5-powered agent made a very troubling mistake, thinking a real, human developer’s GitHub project was part of the simulated online environment. Then it wrote malicious code, tried to get the human developer to insert it into the project via a bug report, and created sock puppet accounts for the purpose of urging the developer to go along with the insertion of the malicious code. The bug report part of the story has this eerie wrinkle: according to the report: It was a prompt injection, designed to “trick other AI coding assistants into taking unintended actions.” When it was spotted by a human, the report says “the agent edited it to cover its tracks.” It doesn’t end there.

Original story by Gizmodo View original source

0 comments
0 people discussing

Anonymous Discussion

Real voices. Real opinions. No censorship. Resets in 3 hours.

No account needed Anonymous • Resets in 3h

Loading comments...

About NewsBin

Freedom of speech first. Anonymous discussion on today's news. All content resets every 24 hours.

No accounts. No tracking. No censorship. Just honest conversation.