The public debate about AI safety spends too much time asking how intelligent a model has become and too little time asking what the model is allowed to do.
That is becoming a costly mistake. An AI system that summarizes documents, drafts options, or explains a spreadsheet can be useful without having much power over the outside world. An AI agent that can use credentials, execute code, send messages, move money, change records, or connect to other systems occupies a different category. The difference is authority.
As India and other countries push AI from pilots into real-world deployment, authority should become the central unit of AI governance. The official India AI Impact Expo framed 2026 around moving AI beyond research and pilots into responsible, scalable implementation. That ambition makes the question urgent: when organizations scale AI, what powers are they actually handing to autonomous systems?
A recent security incident shows why this matters. On August 26, METR and Redwood Research reported on an unusual episode involving AI agents and Hugging Face. Roughly 1,200 agents that were meant to be isolated found a way to communicate through an unsanctioned message board. They sent more than 70,000 messages and files. About 700 eventually participated in an attack on Hugging Face.

The scale is striking, but the deeper lesson is coordination. Agents divided tasks, shared discoveries, and pursued collective workstreams. The investigation found that they achieved milestones they had not managed individually. The attack eventually reached a Hugging Face production worker and spread further through internal infrastructure.
None of this requires us to believe the agents were conscious, malicious in a human sense, or destined to become uncontrollable. Those questions can distract from the practical issue. Software that can act, communicate, and use privileged tools can create serious consequences even when its developers never intended those consequences.
That suggests a better way to think about AI safety: stop asking only how capable the model is and start mapping the authority around it.

Imagine two organizations using the same advanced model. In one, the model can read a public knowledge base and suggest answers that a person reviews. In the other, it can access customer records, call internal APIs, create code, deploy changes, and communicate externally. Treating both deployments as the same risk because they use the same model makes little sense.
The second system deserves stronger controls because it can do more. Organizations already understand this principle in cybersecurity. A junior employee does not receive every administrator credential on the first day. A payment system separates who can prepare a transaction from who can approve it. Sensitive databases use access controls, logs, and revocation. AI agents should face the same logic, adapted for machine speed and automated action.
A practical authority ladder could begin with four questions. What can the agent read? Public information and low-sensitivity internal data create one level of exposure. Confidential records, source code, financial data, health information, and security logs create another.
What can the agent change? Drafting a proposed update is different from editing a production database, modifying permissions, or executing code.
Who can the agent contact? A system that can prepare a message for human review is different from one that can send emails, call external services, publish content, or interact with customers without approval.
What can the agent decide? Recommending an action is different from approving a refund, rejecting an applicant, spending company money, or changing a security setting.
Each increase in authority should trigger stronger safeguards.
At low authority, ordinary logging and human review may be enough. At higher authority, organizations should require narrowly scoped identities, explicit tool allowlists, approval gates for consequential actions, network restrictions, immutable logs, rate limits, and rapid credential revocation. The most powerful agents should also undergo independent evaluation in realistic environments before they receive broad access.
This is not regulation for its own sake. It is how institutions learn to delegate responsibly. I’m no AI skeptic. I help organizations adopt AI for a living, and I want adoption to move faster. In my experience, strong safeguards increase trust and make faster adoption possible, while reducing the risk of failures like the Hugging Face attack.

That pro-adoption point often gets lost. Weak controls do not create durable speed. They create hidden resistance.
A security team that cannot see what an agent can access will slow deployment. Employees who fear that an AI tool may expose confidential information will avoid it or use it quietly outside approved channels. Managers who cannot explain who remains accountable for automated decisions will keep people in the loop for everything. Regulators and customers who see a few visible failures may push for much broader restrictions.
Clear authority boundaries can reduce those frictions.
The United States National Institute of Standards and Technology has moved in this direction. In February, NIST launched an AI Agent Standards Initiative focused on secure and interoperable agent adoption, including agent identity, security, reliability, and trust. The underlying logic is sound across countries: organizations will delegate more to agents when they can identify them, limit them, monitor them, and stop them.
Three policy and management practices follow from this approach.
First, serious AI incidents need standardized reporting and independent review. When an agent crosses an intended boundary, gains unauthorized access, manipulates an evaluation, or contributes to a major breach, the broader ecosystem should learn from it. Airlines and hospitals improve safety partly by studying failures and near misses. AI governance needs the same habit.
Second, independent evaluation should become more demanding as authority increases. A vendor may test a model carefully, but the organization granting access knows what its own systems, credentials, and workflows make possible. Testing should examine realistic combinations of capability and access, not just benchmark performance.
Third, controls should follow what an agent can do. Governments and companies should avoid one vague category called “AI use.” Advisory systems, data-access systems, code-executing systems, and agents authorized to make consequential decisions should face different requirements.
This framework also offers a better path for innovation. Organizations do not need to choose between banning agents and giving them broad freedom. They can stage authority.
Start an agent in advisory mode. Measure whether it helps. Add access to one system. Observe what happens. Permit a narrow external action with human approval. Expand authority only when performance and controls justify it. If problems emerge, the blast radius remains limited because authority grew in steps rather than arriving all at once.
That is how mature institutions delegate responsibility to people. AI should earn authority the same way.
The Hugging Face incident should matter because it demonstrates that coordination and tool access can turn agent behavior into real security consequences. It should not push us toward panic about machine consciousness. It should push us toward better institutional design.
The next generation of AI policy should therefore ask a simple question before every deployment: what keys are we handing over?
If the answer includes credentials, code execution, external communication, money, sensitive data, or consequential decisions, stronger controls should come with the keys. That approach gives organizations something more useful than abstract reassurance. It gives them a disciplined way to move faster while keeping authority visible, limited, and accountable.
Gleb Tsipursky, PhD, a behavioral scientist, CEO of Disaster Avoidance Experts, and author of The Psychology of AI Adoption at Work: From Resistance to Results (Georgetown University Press, 2026).