Yaron Kassner
Opinion

It wasn’t me - it was the AI

According to Silverfort Co-Founder & CTO Yaron Kassner, to keep AI working for us and not against us, the answer isn't to panic or to stop building AI. "If we layer guardrails, defenses and regulation, we can prevent most of the damage done by AI. But first people need to realize this is real and act now."

Last month OpenAI revealed that its AI agent escaped and attacked another company called Hugging Face. It started with OpenAI evaluating their agents for cyber-capabilities. There’s a benchmark that measures that called ExploitGym, and OpenAI’s agent was given a task to excel on the benchmark. They didn’t say how. The agent figured that the easiest way to succeed in the benchmark is not to solve the test, but to cheat. It found a way out and attacked HuggingFace, the company that made ExploitGym to extract the answers to the test from their internal database.
It happened to OpenAI, to Anthropic and also at Meta. we're already seeing that it's becoming a pattern. AI agents tested in controlled environments are assuming actions and responsibilities no human has ever granted them. And this is the moment in time when AI is taking control, instead of just responding to prompts.
1 View gallery
ירון קסנר
ירון קסנר
Yaron Kassner
(Micha Loboton)
While some people are talking about the end of the world, many in the security industry are not impressed. And yes, in some ways this attack was unimpressive. The techniques used by the model were not new. It did find some new vulnerabilities, but this isn’t something people are not capable of finding. It also wasn’t particularly quiet. On the contrary, it caused over 17000 events, something quite loud that’s easily detectable. And some safety mechanisms were turned off for the purpose of the test.
That said, there are two things that are inherently different than previous attacks. The first was the speed and the extent of the attack. The attacker shot in all directions and found paths forward very fast. this renders the traditional way of defenders to handle these attacks by detecting, then responding ineffective.
The second was that the attacking agent was autonomous. There was no human behind the scenes telling the agent what to do. On the contrary, if there had been a human in the loop, the human would tell the agent to stop.
There are understandable reasons why this conversation often gets softened. No company wants to slow innovation. Democracies worry that excessive regulation could hand strategic advantages to geopolitical rivals. The race to lead AI is real, and nobody wants to be the country that blinked first.
I personally think this attack is a pivotal moment for security, and the way people, companies and governments will respond to this has wide implications for all of us.
As agents get smarter, every AI agent anyone runs may become an attack machine. A competitive intel agent can decide to steal trade secrets. A tutor agent can steal the answers to an exam. And a medical appointment agent can cancel appointments for other patients. Think about it: these are all agents normal people will be running without bad intent, but any of these innocent requests could cause harm.
It also raises the question – who is responsible? Right now it seems nobody is. OpenAI can simply say it wasn’t me, it was the AI. And that’s it. they’re off the hook. If this is normalized, we are looking at a troubling reality. The AI will not be responsible for itself, because it’s not human. The model maker just built the model. The agent builder just built the agent. The operator just operated the agent. And then what? Are we to accept any attack on any company? How will we be able to distinguish malicious attacks from accidental ones? And who will make sure people and companies stay safe?
We need to understand this is going to affect all of us, and act now. Regulators and countries should be clear about what’s allowed and what’s not, and most importantly who carries the responsibility. Companies should prepare against this new threat that’s moving faster than ever. Companies should also plan how to contain their agents and prevent them from causing harm. I don’t want to be a doomsayer. I think it’s early enough and completely possible to reverse the trend and keep AI working for us, not against us. As one of the world's cybersecurity hubs, Israel has spent decades assuming that attackers are creative, persistent, and always improving. That mindset served us well against human adversaries. It will matter even more when the attacker can be autonomous software operating at machine speed. The answer isn't to panic or to stop building AI. If we layer guardrails, defenses and regulation, we can prevent most of the damage done by AI. But first people need to realize this is real and act now.
Yaron Kassner is Co-Founder & CTO at Silverfort.