How an AI agent assessment works, step by step
Five phases make one thing visible before any decision: exactly what was tested, what the evidence shows, and which risks remain open.
1. Define the agent
You name the live version, its users, channels, languages, data types, tools, allowed actions and exclusions. Together we walk through one complete healthcare or finance workflow and mark every action that needs a person’s approval. The result is a scope we can test — and later recognize as changed.
| Phase | Applicant responsibility | Reviewer responsibility | Output |
|---|---|---|---|
| Scope | Declare release, workflow, data, tools, owners and exclusions. | Challenge boundaries and identify applicable controls. | Signed scope manifest |
| Evidence | Provide indexed, cleaned records and reproducible test access. | Check sources, completeness and access limitations. | Evidence index and test plan |
| Verification | Support authorized testing and explain observed behaviour. | Run samples, record failures and classify findings. | Finding register and test record |
| Decision | Confirm factual accuracy and remediation commitments. | Independently assess blockers, scores and remaining risk. | Written decision or no decision |
| Review | Report incidents and material changes. | Determine targeted or full retest. | Updated status and review record |
2. Gather the evidence
You provide architecture and data-flow records, access rules, approval policies, test cases, incident history and sample records with personal data removed. Each item is labelled with its source, environment, version, collector and date. Production secrets, raw patient files and customer passwords never belong in the shared material.
3. Test it
The technical review exercises the boundaries that matter: separation between customer accounts, least-privilege tools, resistance to injected instructions, human approval for consequential actions, tamper-showing event records and representative reliability cases. The published sample size and selection method are part of the record — a passing demo alone is never enough. Every finding links to a repeatable procedure.
4. Fix, decide and re-check
Each finding gets an owner, expected evidence and a retest boundary. You may fix the implementation and submit new evidence; the reviewer verifies the changed control and checks whether the change affected neighbouring controls. An interim review can record remaining work, but it is not an issued certification.
The final decision is separated from the implementation work: the reviewer examines critical blockers, the four pillar scores, open limits and conflicts of interest. If a decision is issued, its version, scope and status are published only with your authorization. Review every three months — and sooner after important changes — keeps the result time-bound. Ready to begin? Choose “Discuss your agent” and describe the workflow you want assessed.
Common questions
How long does an assessment take?
The draft promises no fixed duration. Effort depends on scope, evidence readiness, number of workflows, languages, tool access and fixes needed.
Can our own team take part in testing?
Yes: your team explains the system and fixes findings. But it cannot make the final certification decision on its own work.
What happens after a failed test?
The failure enters the finding register with an owner and severity. The fix is retested, and changes that touch other controls widen the retest.