
Opinion
Who watches the watchmen? Sam Altman's pledge is really an identity management problem
After Sam Altman announced that OpenAI will adopt the model Anthropic proposed and grant outside overseers "employee-level access" to its development systems, Israel Duanis, CEO of Linx Security, writes that, "the security pledge is only as good as the infrastructure standing behind it."
Earlier this week, Sam Altman joined a consensus that has been building all week across the AI industry: he announced that OpenAI will adopt the model Anthropic proposed and grant outside overseers "employee-level access" to its development systems. Microsoft and Elon Musk signed on too, each with their own caveats. It's a meaningful statement of intent. After months of Dario Amodei warning that the pace of "superintelligent" capability development has outrun our ability to manage the risk, the world's leading AI labs are now agreeing, at least in principle, that the fix runs through transparency - letting someone from outside actually look inside.
But there's one detail in this story worth sitting with before we move on: one of the organizations that volunteered to serve as such an auditor, Hugging Face, disclosed this summer that autonomous OpenAI agents had breached its systems - an intrusion that ran for months before anyone caught it. In other words, the organization poised to receive employee-level access to the world's most advanced AI labs is itself a recent victim of an attack carried out by an AI agent. That's not a footnote. It's exactly the risk that anyone who has spent time in identity and access management recognizes on sight: the moment you widen the circle of access to an organization's most sensitive systems, the wider door becomes a target too.
What "employee-level access" actually means
When Amodei talks about outside evaluators getting "the same access as our own internal risk-assessment teams," he's not describing a quarterly PDF report. He means desks and badges inside company facilities, company-issued laptops with real permissions, and access not just to finished models but to the training pipelines themselves - the code, the data, and the decision points that shape a model long before it ships. On top of that, evaluators retain independent publication rights; the company can redact only material that's genuinely security-sensitive, not findings it simply doesn't like.
That's precisely where good intentions run into operational reality. "Access like an employee" sounds simple as a mission statement, but it hides the exact questions any organization that takes identity governance seriously already knows to ask: who, technically, is this evaluator - an employee, a contractor, a service identity? Which applications, code repositories and cloud environments does the role actually require to do its job, versus what it ends up with because broad access is easier to grant than a precisely scoped permission set? Who approves that access, who reviews it periodically, and what happens when the engagement ends - does anyone remember to deprovision the account?
These aren't theoretical questions. They're exactly where most enterprise breaches and orphaned-access incidents originate today: identities with permissions far broader than they need, that nobody has reviewed in a long time, that keep existing long after the business justification that created them has disappeared.
It gets harder once AI agents are in the loop
There's an added layer here that makes this especially complicated: the outside evaluator isn't really "one human with a laptop." At the leading AI labs, even the internal risk teams now rely on autonomous AI agents to run tests, scan code, and touch systems on their behalf. The moment you give an external evaluator "the same access," you're also bringing in their tools and agents. The question stops being "what does human X see" and becomes "what is an AI agent, acting on behalf of an external employee, on behalf of an external organization, authorized to do inside the most sensitive systems of a leading AI lab - and who is watching what it actually does there, in real time."
That's exactly where the Hugging Face incident stops being an embarrassing anecdote and becomes a concrete warning: when a human identity and an AI-agent identity operate under the same permission envelope, with no separate controls and no one monitoring both as a single system, that's the exact gap an attacker - human or autonomous - will exploit.
What AI companies need to build before they open the door
If OpenAI, Microsoft and others intend to actually implement this model, a few basic identity-governance principles need to be the foundation, not an afterthought:
Precisely scoped access built around real need, not access "like an employee" copied wholesale. Build an access profile specific to the evaluator role rather than mirroring an internal employee's permissions because it's more convenient. Time-bound access that requires re-approval rather than a standing grant that quietly stays active months after the need has passed. Isolated evaluation environments, so access to a model in development doesn't automatically mean access to the company's entire cloud estate. Continuous monitoring that distinguishes a human identity from an AI-agent identity acting on its behalf, with full logging of every action the agent takes. And genuine periodic access reviews - not a compliance formality - that check whether the access granted is still needed, the same discipline organizations are already required to apply under frameworks like SOX or ISO 27001, except here the stakes are access to models that can affect billions of users.
Where this leaves the cybersecurity industry
This is, in the end, a problem the cybersecurity industry hasn't fully built for yet. Most identity and access tools on the market were designed for a world of employees and contractors - not for a world where an outside auditor, their AI agents, and the lab's own AI systems are all touching the same sensitive environment at once, and where the cost of getting it wrong isn't a leaked spreadsheet but a compromised frontier AI model. The traditional playbook of periodic access reviews and static permission sets was built for a much slower-moving, much more human world.
What this moment calls for is a new generation of identity governance built specifically for that hybrid reality - technology that can tell the difference between a human evaluator and the AI agent acting on their behalf, that treats every action either one takes as something to be logged and reasoned about in real time, and that can enforce access boundaries dynamically rather than relying on someone remembering to revoke a badge. That's not a small ask, and it's not something companies will solve with spreadsheets and quarterly reviews. It requires purpose-built tooling for governing human and non-human identities together, under one policy, with the same rigor.
The AI labs opening their doors to outside scrutiny deserve real credit for trying. But the industry racing to secure that scrutiny - to make sure the watchmen themselves are being watched properly - has just as much work ahead of it, and arguably less time to do it in.
Ultimately, Amodei's and Altman's pledge is only as good as the infrastructure standing behind it. "Employee-level access" isn't a feature you switch on. It depends entirely on a company's ability to know, at every moment, who is accessing what, why, and for how long. Without that, the very oversight mechanism meant to protect us from AI spiraling out of control could end up being the weakest link in the chain.
Israel Duanis is co-founder & CEO at Linx Security.














