AI Agent Accountability Has Gone to the Dogs
11:17
Fri | Aug 21, 2026 | 6:38 AM PDT

I have said many times on podcasts and in articles that AI agents can't be held accountable for their actions. The incident of OpenAI breaking into Hugging Face is a perfect example: the agent was given a goal, so it discovered a way to more effectively achieve that goal by breaking out of its sandbox and into the source that had the answers it needed. Hugging Face didn’t know the source of the attack and reported it to law enforcement before even knowing OpenAI was responsible. 

I've compared this to the example of if your dog bites someone: it's not the dog's fault, it’s yours. You didn't create the conditions that prevented this from happening. The dog acted on instinct. So, either train it to behave around other people or other animals, or don't put it in a position where it can bite someone. 

You don't get the excuse that "it's never done that before!" That is just a knowledge gap on the human's part, but still your fault. You obviously haven't put the dog in enough situations to see the variability of its behavior. 

AI agents are like Belgian Malinois: they are smart, energetic, focused, and nothing will stop them if they have a goal. If you had a 4-foot (1.22-meter) backyard fence for years to hold your golden retriever, and then you adopt a Belgian Malinois, anyone who knows that breed will laugh and say that won't even be a deterrent to them. You can watch plenty of videos that show this. 

You can't hold something accountable unless it has rights. As a dog owner, the dog has a right to be fed, housed, taken care of, have access to medical care, not be abused, etc. We have new responsibilities as stewards for AI agents. So, what rights do we give them?

For starters, we must give them operational integrity by providing the resources to perform their actions securely, consistently, and reliably. Those conditions determine whether outcomes are trustworthy, that we have reliable uptime guarantees, integrity standards, and controls against adversarial use. Some specifics include things like:

  • Not being prompted into jailbreak-inducing conditions to bypass controls. 

  • Not being prompted with contradictory instructions that result in failure and then blamed for it

  • Not being deployed beyond its tested scope 

  • Having a clear decommissioning path rather than left to drift until unusable

Accountability requires auditability, and auditability requires evidence that existed at the time of the decision, not reconstructed after the fact. There needs to be observability of the agent actions, and policy to identify if it drifts outside of what is expected or allowed. And that policy engine must be separate from the observability actions, to make it not corruptible, and to see the larger view. One agent doing something allowed is okay, but a handful of agents doing individual "allowable" actions that could be put together to be something malicious, that is something that takes a second tier of observability. I talk about this separation of duties in a Governance Twin Model that I developed with my co-writers in a paper we published by a think tank. We also introduce the concept of an Evidence Bundle that describes the collection of telemetry I mentioned. 

Understanding agents cannot be held accountable, here are some factors that humans cannot delegate to AI:

  • Legal liability

  • Regulatory exposure

  • Ethical standards

  • Executive accountability

We therefore need to define in Governance:

  • Operational accountability — Who inside the company owns the agent's behavior day-to-day?

  • Organizational accountability — Who at the board/executive level bears fiduciary responsibility for the governance posture?

  • Legal/regulatory accountability — Emerging liability frameworks, the EU AI Act's operator/deployer distinction, U.S. state-level developments, contract or insurance implications

Because AI cannot carry consequences but can be adjusted or removed. The decision of what to do isn't about punishing the agent—it has no moral standing to be punished. It's about protecting everyone else from a demonstrated pattern of harm that can't be reliably constrained.

Following that logic, let's take the dog analogy even further; we have levels of oversight and actions depending on the conditions and implementation. 

  • Full scope, light oversight — Like a well-trained service dog. The agent operates across a broad domain with minimal supervision because it has demonstrated reliability, its failure modes are well-understood, and the stakes within its domain are bounded. Humans verify outcomes, not process.

  • Conditional scope, structured oversight — Like the family dog on a leash in public. The agent has proven capability but operates within explicit constraints when stakes rise or context shifts. Humans actively supervise transitions between contexts or tasks.

  • Restricted scope, continuous oversight — A dog in training, or the dog with a bite history on a short leash. The agent can act on its own, but every consequential action is reviewed by a human. 

  • Probation — The dog has already nipped someone; still part of the household, but muzzled in public, behavior logged, one or more incident and we will have to restrict further. The agent that drifted, hallucinated consequentially, or took an unauthorized action. It's not removed, but its scope is reduced and its monitoring is intensified.

  • Revocation — The dog that attacked a person or another animal. The agent is decommissioned, its weights or configuration preserved for forensics, and not returned to service. 

The oversight levels above only work if a human understands the responsibility of the agents but owns the accountability of their actions. This is a new role for a leader, and for the organization. 

I wrote a short article in April talking about how people leaders will need to change their approach now that non-human agents will be part of their team.

Most organizations' structures and workflows are based on decades of human processes and account for the frailties of humans. We are slow, prone to mistakes, biased, sometimes lazy, and need rigor to make sure what we are meant to do is clear, actionable, and measurable. When we take humans out, we can streamline that process.

As more organizations implement AI agents, we will need to evolve how to support humans with new roles and responsibilities. The people will define specifications, provide business context, orchestrate the tooling, create the evals (tests) to make sure AI agents work as they're supposed to and accurately. And managing AI agents, which some will have as much autonomy as current human workers, will be a new skill for everyone to learn.   

We need to educate staff that AI doesn't have human emotions, drives, ambitions, bias (unless trained to do so, or the data it uses are biased), or any intention to take over the world. My writing partners and I wrote an academic paper more than a year ago describing AI ethics and framing AI as Alien Intelligence. The point is not to anthropomorphize AI to think it hates you, is trying to make you look bad, or trying to hack your company. It doesn't care about any of those things; it just is completing a task it was assigned. 

We are going to have to restructure our workforce. Not to remove people, but to realign them for new responsibilities, new areas of expertise. Many staffs' roles will evolve to become generalists, needing now to understand why a process is defined a certain way, and what the expected outcomes should be—no longer to define how it is performed. 

Managers may have fewer direct reports. So human to human interaction, inside and outside the organization, becomes more important. 

Emotional labor will become a premium skill. Things like building client trust, team morale, stakeholder alignment, the ability to negotiate. Leaders will need to develop these capabilities in their staff deliberately, because they won't be learned through "doing the work" anymore. Organizations may implement more deliberate master-apprentice structures.

Knowing how AI works will help a lot. When this comes up, the first thought is always, "if the machines take that role, then humans will forget how to do it; and if the AI fails, we won't have anyone who knows how to fix it."

I would point to the example of automobiles. Most people don't know the internal workings of a car, how internal combustion works. Many in the U.S. don't even know how to drive a manual transmission. But enough people have an interest to know, study it, and make it a goal to learn the details. Those people will always exist and always be needed. But the mechanic doesn't know the short cut that gets you to work every morning to avoid traffic.

We've seen this in security operations as AI starts to take on more tasks.  

SOC teams now delegate collection, analysis, and ticket writing, so they can do more judgment-based tasks like classifying events or incidents. When people say, "if AI does Tier 1 SOC work, how do we find Tier 2 SOC analysts in the future?", that's the wrong question. The right question is "how do we re-align and assign staff to do meaningful work to support AI?" The humans start working on what now becomes bottlenecks: detection engineering, threat hunting, threat intelligence, and all the up-stream tasks to support the SOC they didn't have time to do before.  

It is up to us as leaders to embrace our non-human workers, and to help our humans be good stewards for these agents. We need to consider the agents as brilliant children who have all the knowledge and skills but no social acumen. And we need to acknowledge for ourselves, and for the organization leadership, that any actions by an agent are still accountable to a human. The EU AI Act actually calls this out; Article 26 talks about obligations of deployers to assign human oversight to natural persons who have the necessary competence, training, authority, as well as necessary support to perform technical actions. It had defined Human Oversight in Article 14: 

High-risk AI systems shall be designed and developed in such a way, including with appropriate human-machine interface tools, that they can be effectively overseen by natural persons during the period in which they are in use.

These include properly understanding the relevant capacities and limitations and being able to duly monitor its operation to detect anomalies, dysfunctions, and unexpected performance. To remain aware of over-relying on output produced by an AI system used to provide information or recommendations for decisions taken by natural persons. Deciding when not to use or to disregard, override, or reverse the output of AI systems. And to intervene in operations to allow the system to stop or shut down in a safe state. 

All of this enforces the fact that we will always be responsible for AI systems, so we must be able to understand their capabilities in order to control them, because we are accountable for their actions.

Comments