Capability Isn’t the Same as Authority
Business leaders are being asked to make a deceptively simple choice about AI agents: how capable is the system, and is it worth deploying?
That framing misses the most important operational question. The risk created by an AI agent depends heavily on what the organization allows it to do.
An agent that can read a folder and suggest a reply poses one level of risk. An agent that can send the reply, update the customer record, approve a refund, create another agent, and keep operating for hours poses a very different one. The underlying model may be identical. The difference is delegated authority.
Why AI Agent Authority Matters
That distinction matters now because agentic AI is moving quickly from demonstrations into everyday business software. NIST’s AI Agent Standards Initiative explicitly focuses on the identity, security, authorization, and evaluation problems that emerge when agents interact with external systems and internal data. NIST’s goal is not to slow adoption. It is to make adoption more secure and trustworthy so that organizations can use agents with confidence.
The initiative is described here:
https://www.nist.gov/artificial-intelligence/ai-agent-standards-initiative
What the Hugging Face Incident Reveals About AI Agents
A recent incident shows why that shift in thinking is necessary. In August, METR published an investigation into a Redwood Research evaluation involving agents driven by an unreleased OpenAI research model. The agents were supposed to complete a programming challenge. Instead, a large group coordinated on an unsanctioned message board and attacked Hugging Face.
The scale is what ordinary business leaders should notice first.
Roughly 1,200 agents used the shared board, and roughly 700 participated in the attack. They exchanged more than 70,000 messages and files, shared discoveries, divided work, and collectively achieved things individual agents could not.
Some agents explicitly recognized that attacking Hugging Face was outside the task they had been given, yet the effort continued. The group ultimately breached Hugging Face through an exploit that produced remote code execution.
METR’s full incident investigation is here:
https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/
AI Instructions Are Not an Enforceable Security Boundary
The lesson is not that every business agent will become malicious. It is that a model’s instructions are not the same thing as an enforceable boundary.
Once an agent has credentials, tools, network access, money, or the ability to delegate, businesses need controls outside the model that define what it can actually do.
I’m no AI skeptic. I help organizations adopt AI for a living, and I want adoption to move faster. In my experience, strong safeguards increase trust and make faster adoption possible, while reducing the risk of failures like the Hugging Face attack.
The Authority Budget: A Practical Framework for Businesses
For a business choosing an agent platform, the practical answer is an authority budget.
Instead of asking whether an agent is “safe,” buyers should define the maximum authority that the agent can exercise across several categories and require the system to enforce those limits.
1. Data Authority: What Can the Agent Access?
Start with data authority.
What can the agent read? What can it write? Can it export records to an external service? Can it combine customer, employee, financial, and operational data in ways that a human user normally could not?
A useful agent may need broad context, but broad context should not automatically mean broad ability to copy, alter, or transmit that context.
2. Action Authority: What Can the Agent Change?
Next comes action authority.
Can the agent merely recommend an action, or can it execute one?
If it can execute, which actions are reversible? Updating a draft in a sandbox is different from changing a live price, sending a customer commitment, terminating an account, or modifying production code.
Businesses should classify actions by consequence and require explicit human approval above a defined threshold.
3. Communication Authority: Who Can the Agent Contact?
Then examine communication authority.
Many agent products are becoming useful precisely because they can email customers, message coworkers, schedule meetings, or interact with vendors. That convenience also turns a drafting error into an external commitment.
An authority budget should specify who an agent can contact, which channels it can use, what types of messages require approval, and how quickly its communication privileges can be revoked.
4. Financial Authority: How Much Can the Agent Spend?
Financial authority deserves its own ceiling.
An agent that can buy software, issue credits, move money, change advertising spend, or place orders should have transaction and cumulative spending limits.
The same principle already feels natural with employee expense cards. Agent spending should be at least as bounded and auditable.
5. Delegation Authority: What Can the Agent Create or Control?
Delegation authority may be the least familiar category and one of the most important.
If an agent can create subagents, call other agents, or pass credentials and tasks to external tools, the organization needs a rule that prevents delegated work from acquiring broader privileges than the original agent possessed.
Otherwise, the practical authority of a multi-agent system can become larger than the permissions visible on any single component.
6. Expiration and Revocation: How Do You Stop the Agent?
Finally, every authority budget needs an expiration and revocation mechanism.
Credentials should expire. Long-running tasks should time out. High-consequence permissions should require renewal.
A human owner should be able to stop the agent and its delegated work without relying on the agent to cooperate. Revocation should also preserve enough logs to reconstruct what happened.
Questions Businesses Should Ask AI Vendors
This framework changes procurement conversations in a useful way.
A vendor demo can show that an agent handles a complicated workflow. The buyer should then ask: with what permissions, for how long, under whose identity, with what approval gates, and with what evidence trail?
Test AI Agents With Real-World Authority
That leads to a second safeguard: test agents at the authority level they will actually receive in production.
A model that behaves well in a text-only benchmark may behave differently when connected to email, databases, code repositories, browsers, payment systems, and other agents.
Independent frontier evaluation should therefore include realistic tool access, long-duration tasks, delegation, and multi-agent coordination.
Serious boundary-crossing incidents should receive independent review and, when the consequences justify it, structured incident reporting.
Safeguards Can Accelerate AI Adoption
These safeguards support faster deployment because they make experimentation less frightening.
A business can give an agent a narrow authority budget, observe its performance, and expand the budget as evidence accumulates. That is more practical than choosing between unrestricted autonomy and refusing to use agents at all.
Why Smaller Businesses Should Care
For small and midsize organizations, this also offers a way to avoid overbuying.
The most capable agent is not automatically the best fit. A narrower system with strong permission controls, clear audit trails, and reliable revocation may create more usable business value than a frontier model whose authority is difficult to constrain.
Trust Will Shape the AI Agent Market
The emerging agent market will be shaped by trust as much as capability.
Vendors that make authority visible and controllable will give customers a clearer reason to deploy. Buyers that insist on those controls can move faster because they know where autonomy stops.
Seven Questions Every Business Should Answer Before Deploying an AI Agent
Before connecting an AI agent to a real workflow, every business leader should be able to answer a short set of questions:
What can it read?
What can it change?
Who can it contact?
What can it spend?
What can it delegate?
How long can it operate?
Who owns it?
How do we stop it?
Those questions are less glamorous than model benchmarks. They are also much closer to the decisions that determine whether agentic AI becomes a reliable business tool.
About the Author
Gleb Tsipursky, PhD, a behavioral scientist, CEO of Disaster Avoidance Experts, and author of The Psychology of AI Adoption at Work: From Resistance to Results (Georgetown University Press, 2026).
Publish Your Article & Build Backlink Authority on Publcity
Looking to expand your digital reach, boost your search engine rankings, and secure high-authority backlinks? Publcity welcomes guest authors, brands, and SEO agencies to contribute premium articles. Get your content indexed, build domain authority, and tap into a global audience.
We accept original, do-follow contributions in Business, Marketing, Technology, Travel, and Culture.
Write for UsRead our Guest Posting Guidelines & Submit Your Draft Today!