Write a one-page scope first
Name the working agent, release, intended use, users, supported languages, environments, data types, tools and actions. State exclusions in testable words: “drafts appointment requests but cannot confirm a booking” can be checked; “makes no important decisions” cannot.
List the organization operating each component and the person owning the remaining risk. Mark where customer configuration changes the provider’s default control. Buyers must know which evidence belongs to the platform — and which belongs to their own setup.
Organize proof around claims
For each control, give the claim, the applicable configuration, the source record, collection date, version and verification result. Useful records include architecture and data-flow diagrams, access matrices, approval records, test definitions, incident exercises and release notes. Hide unnecessary personal data but keep the facts a reviewer needs.
Policies show intent. Add operating proof: a denied out-of-scope tool call, a detected cross-account attempt, a blocked expired approval and one reconstructed sample transaction. Screenshots without origin, environment or time are hard to rely on.
Show failures and limits
Publish sample size and selection. Show failed cases, fixes, retest results, untested populations and known third-party limits. A clean executive summary that hides exceptions creates more procurement work — reviewers must rediscover them.
State review triggers: model or instruction changes, a new data source, wider tool permission, a new language, a serious incident or changed use. Keep the release inventory linked to regression results, so buyers can judge whether the evidence is still current.
Share in layers
Start with scope, control summary and remaining risk. Give detailed evidence only to authorized reviewers through an access-controlled channel. Never ask prospects to upload production logs, patient files, credentials or confidential prompts to a public contact form.
The finished pack should let an independent reviewer replay a risk-based sample. Your team can explain and fix controls — but it should never be the only party deciding that its own work is certified. The same pack feeds your policy-ready dossier on the insurability track.