Tools
Build safe agents, before you ship them
Free, browser-based tools for the governance layer of agentic AI — scan an agent against published guidance, classify its risk, and keep your work in one place. No login, no API key required for the deterministic tools.
Compliance Scanner
Microsoft-aligned · AgentAz™ companionPaste an agent's system prompt or agentaz.json and scan it against the design-layer controls in Microsoft's published governance guidance — with the AgentAz companion mapping. You get pass/fail gates, the specific ways it can fail, a risk radar, and a copy-paste fix block. Deterministic; your prompt is processed on our edge and never stored.
Scan an agent →Failure Library
Post-mortems · real incidentsEngineering post-mortems of real, documented AI agent failures — Air Canada's fabricated refund policy, the Chevy chatbot's $1 car. Each with root cause, the control that would have prevented it, and OWASP/NIST mapping.
Read the post-mortems →Compare Blueprints
Side-by-side · governancePut any two agent blueprints side by side on what decides production-readiness: Trust Level, how many tools auto-execute vs. require approval, worst-case action, frameworks, and setup. Real spec data, not ratings.
Compare two kits →Drift Auditor
Deterministic · governance diffPaste two versions of an agent — old vs new agentaz.json or system prompt — and get a deterministic governance diff: which tools lost approval gates, whether the Trust Level escalated, and the evidence, with a before/after Governance Radar. The same verdict every run.
Audit a change →State of Agent Governance
Live aggregate dataAnonymous aggregate data on how safe real agents are — the share that would run a risky tool without approval, the distribution of Trust Levels, and how it changes over time. No prompts or identities stored.
See the data →Predict the Break
Game · failure modesRead a real agent's setup and call where it breaks before you see what happens — ungated actions, prompt injection, runaway loops, and more. Fast, free, and it sharpens the pattern-matching a checklist can't teach.
Play a round →Risk Assessment
Autonomy tiers A0–A5Work through an agent's authority, reversibility, and oversight to classify its risk tier and see what governance it needs before deployment.
Assess risk →KitForge
Open source · enforcing guardrailsGenerate a LangGraph agent scaffold whose safety actually enforces: authority budgets that block, an HMAC-chained audit log that fails on tamper, and human-in-the-loop gates that halt. Ships the tests that prove it. Python, MIT, with attribution baked into everything it generates.
Get KitForge →My Workspace
Saved & recently viewedYour saved blueprints and recently viewed kits, kept in your browser. Built to become your account workspace when login arrives.
Open workspace →On every blueprint
Two more tools live inside each kit page, because they only make sense against a specific blueprint: