Start with what can go wrong
Write down what changes when the agent is wrong. A search assistant that returns a weak link is a different risk from an agent that sends a patient message, books an appointment, freezes a workflow or edits a customer record. List the affected people, data, tools and commitments before you discuss model accuracy.
Separate three verbs: drafting, recommending, executing. This single distinction usually reveals where a human approval or a narrower permission is missing. In healthcare, mark where admin support approaches a clinical decision. In finance, mark where explanation becomes an account or transaction action.
Demand a scope you can replay
Record the model and provider, instructions, knowledge base, integrations, permissions, supported languages, environments and exclusions as one versioned release. A generic platform certificate never describes your configured agent. Ask how the supplier detects changes to hosted models or knowledge sources that carry no conventional software version.
Map data from collection to deletion. Include derived fields, transcripts, records, caches, backups and support access. Ask for the real processing locations and the applicable transfer safeguards. Swiss data protection applies to AI processing — but it imposes no universal rule that every record stays in Switzerland.
Test the edges, not just the happy path
Build cases from real workflows and their edges: missing records, conflicting sources, unsupported languages, a hidden instruction inside a document, a second customer’s identifier, an expired approval and a tool request outside scope. State the sample population and how you chose it. A vendor-chosen demo cannot predict behaviour in your use.
For consequential actions, watch the complete chain. The reviewer should see the exact details before approving; any later change must cancel the approval. Try refusal, expiry and reuse of an old approval. Confirm that the tool or API enforces the authorization — not just the wording of a prompt.
Make a decision you can defend
Turn findings into acceptance conditions with owners, deadlines and rollout limits. Disclosure across customer accounts, uncontrolled privileged tools and bypassed human approval are blockers — never items to balance against a high average score. Define monitoring, incident response, fallback, exit and reassessment triggers.
A certificate supports this decision only when its agent version and scope match yours. It replaces no legal analysis, clinical validation, internal accountability or supplier management. Keep the remaining-risk statement next to the approval, so future teams know what was accepted.