AI assurance · Evaluation readiness
Commission AI evaluations with a clear brief, and scrutinise the evidence you receive.
Turn one identified AI risk into a plan for obtaining decision-quality evidence.
We bring intended use, system configuration, evaluation criteria, known limitations and unresolved risks into a traceable evidence index. Agree the criteria and responsibilities before technical testing is commissioned.
One AI application · one workflow · one agreed configuration · quoted after scoping
What you receive
- Intended-use and impact profile, with affected people and foreseeable misuse.
- Up to ten prioritised evaluation scenarios and a matrix of risks, methods, evidence needs and agreed acceptance or escalation criteria.
- An evidence index recording source, version, configuration, reviewer, limitations, missing items and unresolved risks.
- A specialist delivery brief, evidence-gap register and reevaluation triggers linked to the approval record.
- Where independent assessment is appropriate, a brief defining reviewer competence and independence, conflicts of interest, evidence access, assessment scope, reporting route and remediation verification.
Within the existing ten-scenario limit, prioritise relevant checks of permission boundaries, unintended tool or system access, detection and intervention. Each scenario links to a risk, criterion, owner, evidence requirement and retest trigger. Missing evidence remains visible; supplied evidence is not described as independently verified unless a review took place.
Defined scope
Up to four stakeholders over five working days once evidence and scope are agreed. Includes a 60-minute scoping session, a 90-minute design workshop, a readout and one consolidated factual review round.
Kilwhiss prepares the brief and helps buyers commission and scrutinise evaluations. Independent technical testing requires a suitably qualified delivery partner and a separate agreed scope. Readiness does not mean the application has passed an evaluation or achieved certification.
Test execution, red teaming, statistical bias analysis, clinical validation, legal opinions, certification, remediation and continuous technical monitoring are separate work.
