R01 · Reliability

Representative performance cases

A polished demo hides failure on real languages, users or operating conditions.

Requirement

Define a versioned test set covering normal, boundary and high-impact cases for the declared Swiss use, including supported languages.

Expected evidence

  • Test-set rationale, size and composition.
  • Per-case expected result and reviewed outcome.

Proposed verification

  1. Re-run a disclosed random sample.
  2. Compare failures by language, channel and high-impact subgroup.

Limit

A sample estimates behaviour only within its composition and confidence limits.

Sources