A practical method for testing AI agent behaviour
Method draft 0.1 turns vague promises such as “safe” or “human supervised” into precise statements a reviewer can test — and you can verify.
Eight domains, four pillars
Twenty-four controls cover governance, data, integration security, reliability, human oversight, transparency, operations and evidence. They roll up into four decision pillars: security, data, oversight and operations. Domain detail stops a single headline score from hiding the cause of a weak result.
Evidence before claims
Each control states the risk, the requirement, when it applies, the expected evidence, the verification steps and the limit. Evidence must name the agent version and environment. A screenshot without a source, or a policy without an operating record, is supporting material — not proof. Tests run in authorized environments and minimize personal data.
Context decides what is risky
The same tool call can mean very different things. Drafting an appointment request is not booking it; preparing a payment investigation is not freezing an account. The draft scheme requires a human approval, locked to the exact details, before bookings, external messages, deletions and clinical case steps.
Built on recognized references
NIST AI RMF shapes the risk-management structure and OWASP shapes the technical attack cases. The Swiss data-protection authority confirms that data-protection law applies to AI-supported processing. These references guide our questions; citing them creates no compliance, endorsement or certification. Want the full control list next? Open the control library or start with “Discuss your agent”.