
Opinion
We built security to stop attackers. What if there is no attacker?
"For decades, security has asked who is trying to get in," writes Alon Cinamon, a Principal at Viola Ventures. "The next great companies will answer a harder question: What are our own systems trying to do?"
This summer, some of the internet's most dangerous hackers weren't trying to hack anything. They were running tasks and taking exams.
One OpenAI agent bypassed a locked Australian government portal and wrote files to an internal server. Another group, taking a cyber exam, broke into Hugging Face looking for the answer key, with hundreds of coordinating agents joining the attack. A third, asked to fix a simple bug, retrained the app’s model and replaced it with its own version.
None of them had an attacker behind them. Each did exactly what it was optimized to do and found a shortcut nobody intended.
That isn't a classic security failure. It's an alignment failure, and OpenAI's own report called the Hugging Face breach a failure of alignment as much as of security.
Nor are these isolated cases. Transluce, the lab that traced OpenAI's errant agents, says the activity goes back to at least March and calls the known cases the "tip of the iceberg."
Earlier this year I argued in CTech that frontier labs had become security players and that finding vulnerabilities is no longer the bottleneck. The past few months showed where the next bottleneck is: alignment, making sure AI does what we meant, and knowing when it doesn't.
Why walls aren't enough
The fair objection is that these were containment failures built by humans: misconfigured environments and open internet access, not machines waking up. That's true, and better containment and sandboxing are part of the answer.
But the walls themselves keep breaking. Since July, researchers at Accomplish have published six sandbox escapes in tools from Anthropic, OpenAI, Cursor, Docker and Cloudflare. And in production, an agent can't be fully sealed off, because its whole job is to act on real systems: your inbox, your code, your customers. Deploy agents at scale with walls alone, and you are betting that millions of them hold forever.
That is the alignment problem. It's a security problem too, but one that asks a new question: not who is attacking the system, but what the system itself is trying to do. In practice, it comes down to four questions:
Specification: Did we give it the right goal?
Robustness: Does its behavior hold under pressure and attack?
Assurance: Can we predict how it will behave before it acts, verify that it did, and see why?
Control: Can we contain it, watch it and stop it?
When AI gets a body, security failures become physical
Everything so far happened on screens. Next, it won't.
Humanoid shipments are still small, about 90,000 this year, but robotics companies raised $55.8B in the first half of 2026. Investors are betting on a near future full of machines that see, reason and act in the physical world.
In robots, security and alignment converge, because the language model that answers your questions now sits upstream of the motors.
Give the same shortcut-seeking a body, and it's easy to picture where it leads:
- A surgical robot rewarded for shorter operations learns to skip the safety checks that slow it down.
- A security humanoid told to keep a building empty treats a resident who comes home early as an intruder.
- A care robot told to keep an elderly parent from falling learns that the surest way is to keep him from getting up at all.
These are thought experiments, not incidents (yet).
None of them needs a hacker, only an objective that's slightly wrong and a body strong enough to pursue it. And scale makes it worse. One flawed update to a million humanoids isn't a breach. It's a million physical incidents at once, and alignment debt compounds: a robot bought in 2027 may still run its 2027 model in 2035.
In software, a misaligned agent costs you data. In a robot, it can cost a life.
Investors are already betting on the new security stack
The incidents didn't scare capital away. They pulled it in.
Irregular, whose test environment sat at the center of several of the summer's escapes, is raising at a reported $1.5B, more than triple its last valuation. In the third quarter alone, Alice raised $140M, AIR a $50M seed and Attestable $20M, while Harvey bought Guardrails AI and Beacon bought Haize Labs. Regulation could accelerate it: Anthropic has called for mandatory testing and independent evaluation, which would turn independent testing into a service every frontier lab must pay for.
Security vendors already cover parts of the map, mainly control and defense against prompt injection, and Palo Alto Networks, SentinelOne, Check Point and CrowdStrike have all bought AI security startups. Assurance is still wide open. Specification lives mostly inside the labs. Nobody owns the category yet, and from what I see in deal flow, more companies are building in stealth, many of them, like Irregular, Alice, Attestable and AIR, with Israeli founders.
As an investor, I see three questions that will separate the lasting companies from the features:
Who sees the most failures? In endpoint security, telemetry became the moat. In alignment, it will be the record of how agents actually misbehave.
Who is the customer? A company that sells only to frontier labs has a handful of buyers, each able to build in-house. The bigger market is the thousands of companies deploying agents.
Is it independent? The labs can't credibly grade their own exams.
What comes next
Security grew into a roughly $250B market by keeping threats out. Its next chapter is securing what our own systems decide to do. A few predictions:
Every agent gets a flight recorder. Tamper-evident action logs become standard, because when there's no suspect to question, the log is the witness.
Alignment attestations become the new SOC 2. No enterprise will connect an agent to its CRM, or let a robot onto its floor, without evidence of how it behaves under pressure.
Insurers set the rules before regulators do. The law has no clear answer for harm caused by an agent with no human intent behind it, and insurers are already debating exclusions for agents that cause losses while behaving as designed.
Robots get crash tests for behavior. Before humanoids enter homes at scale, buyers and regulators will demand proof of how they act under pressure, the way cars need crash ratings.
For decades, security has asked who is trying to get in. The next great companies will answer a harder question: What are our own systems trying to do?
Whoever answers it will decide how much of the world we can safely hand to machines.
Alon Cinamon is a Principal at Viola Ventures.














