AgentKits
Live data

The State of AI Agent Governance

Anyone can build a capable agent. The harder question is whether it’s safe to run — and most aren’t asked it. This page reports, in aggregate, what we actually see: how much authority real agents hold, how often a consequential action runs without a human in the loop, and how that picture shifts over time. Every number is computed from anonymous results of agents checked with the AgentAz Compliance Scanner — never from prompts, specs, or identities, which are never stored.

Loading the latest numbers…

How this is measured

The scanner grades an agent’s system prompt or agentaz.json against published agent-governance guidance and the AgentAz Trust Levels. A “failed gate” means a required safeguard — most often human approval on a consequential action — isn’t present. A “critical scenario” is a failure mode the agent has no guard against. The Trust Level is derived conservatively from what the agent can do without a human. Figures update as more agents are scanned; percentages are withheld until the sample is large enough to be meaningful.