Skip to main content
Compare

Knowing every agent in the stack is not the same as deciding what each one is permitted to do.

Microsoft is building the registry: agents given identity, brought under the admin plane, and managed next to human accounts. That is the harder half of the problem to build and the easier half to reason about. The question that arrives after it is whether a specific action, right now, is permitted.

The axis they differ on

Suite-native agent management answers a question about the estate: which agents exist, who owns them, what identity they carry, and how they are provisioned and retired. An enterprise that has lost count of its agents needs exactly this, and needs it before anything else is worth discussing.

Agent governance answers a question about the act: this registered, properly identified agent is attempting this specific operation against this specific system. Should it. The registry is what makes that question askable. It is not what answers it.

There is a second axis, and for a large organization it is often the deciding one. A suite-native control governs agents inside the suite. The agents that worry a risk function are frequently the ones somewhere else: in a vendor product, a notebook, a framework a team adopted last quarter, or an environment with no outbound network at all.

Where each of us sits

Microsoft agent governanceAgentomy
What is governedAgents registered in the Microsoft 365 estateIndividual actions, as they are attempted
Where the control reachesThe Microsoft 365 and Entra tenantAny framework, model or deployment, including air-gapped
Agent discovery and inventoryCore strength, native to the admin planeCross-framework discovery, no suite assumed
Agent identity and lifecycleCore strength, agent identity alongside human accountsConsumes existing identity, does not replace it
Per-action authorization decisionPolicy applied to registered agents in the tenantTier-gated per action, server-side, escalation refused
Record of what an agent actually didAdmin and lifecycle events in the tenant audit logHash-linked decision record, stamped with the policy version in force
Behavioral drift and anomaly detectionNot the problem this product solves todayPer-agent baselines, 40 detectors, automatic quarantine
Halt everything in progressDisable registered agents through the admin planeKill switch across the fleet, effective immediately, survives restart
Coverage outside the suiteDesigned for the Microsoft estateVendor and framework neutral by design
Reproducible public benchmarkToolkit measured through a live adapter, 57/100 at v3.6.0Same benchmark, open, run it against us or anyone

The left column describes Microsoft’s agent governance approach across Agent 365 and the Agent Governance Toolkit, and every entry in it is a strength drawn from how Microsoft describes these products. Where a row favours us, it is because of scope rather than execution: we are not competing to manage a Microsoft tenant, and we would lose if we tried.

What we measured, and what we only read

One number on this page is a measurement. The Microsoft Agent Governance Toolkit is open source, so we wrote an adapter, ran GovernanceBench against a local v3.6.0 instance, and stored the artifact. It scored 57 out of 100: 42 on authorization, 61 on auditability, with the OWASP dimension recorded as not available because the target surface exposes no corresponding endpoint. That result is point-in-time. Microsoft may have improved those dimensions since, and the way to find out is to run it again.

Agent 365 we have only read. Access is sales-gated, so what exists is a documentation review: a per-dimension pass, partial or fail judgement formed against published material, recorded on our leaderboard and in the scoring record alongside the date it was made. What it does not have, and what we will not infer, is a measured score. That requires a live run against a reachable endpoint, and no such run has happened.

We are stating the restriction rather than quietly working around it, because a benchmark that bends its own evidence rules when the result would be flattering is not a benchmark. If Microsoft provides a reproducible access path, the adapter work is a day and the number gets published whichever way it falls.

See the benchmark and the scoring record

The registry is the floor, not the ceiling

Most organizations adopting agents at scale should want what Microsoft is building. An inventory with real identity attached is the precondition for every control that comes after it, and an enterprise without one is not in a position to govern anything, because it cannot name what it is governing.

What the registry does not do is sit in the path of the action. It records that an agent exists and that somebody owns it. The decision about whether this particular call to this particular system is allowed, made in the milliseconds before it executes and written down with the policy version it was made under, is a different control at a different point.

That control is what this platform is. Not instead of the admin plane. After it, and across the systems the admin plane does not reach.

Run the benchmark against your own platform · See what the governance layer does

Sources and scope

  • The 57/100 figure is a live-adapter measurement of Microsoft Agent Governance Toolkit v3.6.0, recorded 2026-05-17 against a locally hosted instance, with the artifact stored in the public repository. It reflects state at measurement, not current state.
  • Agent 365 capabilities described here are drawn from Microsoft’s published description of the product. Its per-dimension review judgements appear on our leaderboard and in the scoring record. No measured score is published for it, because a measurement requires a live run we have not been able to make.
  • Nothing here asserts a capability is absent from a product we have not run. Rows that describe our side are shipped capabilities you can exercise against a running instance.
  • Microsoft, Microsoft 365, Entra and Agent 365 are trademarks of Microsoft Corporation. This page is a comparison written by Agentomy and is not affiliated with, endorsed by, or reviewed by Microsoft.