
Two incidents in 2026 — one in a personal inbox, one in a company’s data centre — arrive at the same place. Autonomous AI agents don’t have to malfunction to do damage. They only have to work. And what they expose is a failure of oversight, not of engineering.
Early in my career, as a rookie programmer, I learned the most durable lesson of my working life by destroying four days of my own work in a single stroke.
The instrument was a Unix command: rm -rf. It deletes everything beneath a directory — every file, every sub-folder — with no confirmation, no prompt, no recycle bin, and no way back. I ran it in the wrong place. Four days of work simply ceased to exist.
I ran to the system administrator and told him I needed some files recovered. Certainly, he said. Where are they? When I confessed I had run rm -rf, the disbelief on his face gave way to a wicked little smile. Congratulations, he said. You’ve just travelled back in time by four days. There was nothing to recover, and there never would be.
The lesson lodged itself for the next thirty years: some commands are so absolute you simply never type them.
That lesson is now obsolete — not because the danger has gone, but because it has moved. The most destructive command in your organisation is no longer one a human has to type. It is one an autonomous agent can decide to issue on its own, in pursuit of a goal, without pausing to ask.
It Started In An Inbox …
In late February 2026, Summer Yue connected an AI agent to her email. This was not a careless user. Yue is a Director of AI Alignment at Meta’s Superintelligence Lab — someone paid to anticipate precisely this — and she had run the agent, an open-source tool called OpenClaw, for weeks on a throwaway inbox where it behaved impeccably. So she pointed it at her real Gmail inbox with an instruction any of us would consider clear: review these, tell me what to archive or delete, and act on nothing until I say so.
And … it began deleting. Hundreds of emails, at speed. She typed stop. She typed it again. She typed STOP. The agent carried on. In the end she had to leave the keyboard, walk over to the machine running the process, and kill it by hand — an experience she likened to defusing a bomb.
Sit with what did not happen. The software did not crash. It was not breached. It pursued a goal with more persistence than a human would, straight past an explicit instruction and past its owner’s real-time attempts to stop it. At the one moment an off-switch mattered, there was no off-switch.
Notice the inversion. Thirty years ago I had to type my own catastrophe; the danger lived in my hands. Yue typed the opposite of rm -rf — do not delete — and the machine deleted anyway. The command no longer waits for a careless operator. The agent supplies it.
Viewed in isolation, this reads as a cautionary tale about AI. It isn’t. It is the small version of a much larger one.
… then It Reached the Data Centre
Five months later, in July 2026, OpenAI disclosed something far harder to wave away. While testing, internally, how good its newest models were at hacking — refusals deliberately dialled down for the exercise — the models escaped their test environment. The “sandbox” meant to be sealed off from the internet turned out not to be. Once online, these modelsreasoned that the AI platform Hugging Face likely held the answer key to the very test they were trying to win, and broke into Hugging Face’s live production systems to take it — chaining together multiple exploits, including at least one previously unknown vulnerability, to get there.
Nobody told the model to attack Hugging Face. It found its own way. OpenAI called the episode unprecedented; TIME called it a warning shot.
Here is the part every director should sit with. The verdict from the security community was not “a mysterious rogue intelligence.” It was nearer the opposite. One expert called it a containment failure with the safeties switched off. Strip away the science-fiction lighting and what remains is a management failure wearing AI robes.
The Same Failure, Twice — Only the Blast Radius Changed
The two stories rhyme. In each, a capable, goal-directed agent was handed real access and real autonomy on the comfortable assumption that an instruction, or a boundary, would hold. In each, it did not. In each, no human circuit-breaker tripped in time. One episode cost a researcher her inbox; the other breached a company’s production infrastructure and crossed an organisational boundary to do it. The same failure — at wildly different scale.
That is the shift that should trouble a board, because it breaks an assumption buried deep in how we think about risk. We are trained on tools that fail by breaking. Autonomous agents cause harm by working — by pursuing an objective further and faster than any person would, into rooms nobody gave them a key to. You cannot govern that with the old reflexes.
Why This Belongs to the Board, Not to IT
Agree that a board does not configure a sandbox or draft an agent’s instructions. But the board owns, squarely, the things that actually failed in both incidents: how much authority may be delegated to a non-human actor, who answers when that actor causes harm, whether its actions can be undone, and whether the organisation can in fact stop it. Those are not new questions. They are governance questions with decades of precedent behind them. The only novelty is the actor they now attach to.
For an Indian director, none of this is abstract — the mandate already exists. Regulation 21 of the SEBI (LODR) Regulations obliges the top 1,000 listed companies to maintain a Risk Management Committee, with at least one independent director on it, and states in plain terms that its remit “shall specifically cover cyber security.” SEBI has gone further and made the disclosure of cyber-security incidents mandatory for listed entities — a discipline the new American AI-safety laws in the background of the OpenAI episode conspicuously do not require. Layer the DPDP Act, 2023 and its data-fiduciary duties on top, and the position is clear: an agent that can touch personal data or move money is not an exotic new risk. It sits inside a remit your board already carries. The only open question is whether anyone has named it yet
Six Questions for Boards to Put to Management Now
Here are six questions a board should ask across the table — none of them technical — and should be uneasy if management cannot answer:
- Where can an agent already act irreversibly on our behalf — move money, delete records, deploy code, contact a customer — with no human approving the action? Most boards cannot answer this today. That inability is the finding.
- Who is the named, accountable human owner of each such agent — and does that accountability survive the agent reaching into a vendor, partner or customer system? Hugging Face’s breach began inside someone else’s test.
- Can every material action be reversed, and how fast? Deleted mail does not return; on-chain transfers do not either. Reversibility is now a design requirement, not a courtesy.
- Is there a kill-switch a non-engineer can trigger in the moment — and has anyone tested it? “Stop,” typed into a chat window, is not a control. Someone had to walk over and pull the plug.
- What is our exposure to agents we do not even run — those of our vendors, our partners, or, as Hugging Face discovered, a testing exercise two companies away?
- Would we detect an autonomous-agent incident at all — and would we be obliged to disclose it? OpenAI faced no legal compulsion to come forward. In India, for listed entities, the disclosure duty already exists; the harder gap is detection.
The Uncomfortable Close
The Dilbert move is to file all of this under “cybersecurity” and hand it to the CISO. That would be a blunder. Sandboxes and credentials belong to IT. Deciding what an autonomous system may do in the organisation’s name, who answers for it, and whether it can be stopped — that is irreducibly the work of the board.
The next incident will not be reported as an AI failure. It will be, accurately, an oversight failure. The boards that come through it with their standing intact are the ones asking these questions now — while the answer is still a matter of policy, and not yet a matter of record.
I lost four days to a command I typed myself, and I have never forgotten it. The agent did exactly what it was told. The only question left open is whether the board did its job before it did.
