Method #7, Risk & compliance, 2 min read
Risk & compliance
Red-teaming, policy adherence, fairness, and evidence for the committee that asks.
Evidence your risk team can read, not a badge.
Contents

What it is
Some failures are not quality in the ordinary sense. They are policy, leakage, or a sentence you cannot defend.
This is for teams whose first evaluation already exists and who now need a scored red team or a policy file.
Red-teaming with a catalog
Adversarial tests are designed from the failure-mode map: jailbreaks, leakage, disallowed advice, prompt injection, retrieval of the wrong document. Findings are a catalog with patches, not a slide of scary examples. A focused red team is this work in concentrated form.
I re-test after the patch. A red team that does not return is a performance. The catalog is versioned with the system it was run against.
I score policy, leakage, and the cases that would fail a release, against the same rubric the rest of the evaluation uses. A PDF of jailbreaks without a threshold is a scare file.
Policy, fairness, safety
Instruction adherence is scored against system, policy, and format constraints. Bias and demographic performance are sliced, not averaged. Content safety is a dimension with a threshold, including “zero critical” where that is the only acceptable line.
Operational risk (policy, leakage, escalation) is this family of failures. A high faithfulness score does not excuse a single critical leak.
Fairness work is sliced evidence, not a slogan. I report the slice that fails and the sample it rests on. I do not publish a single “bias score” that cannot be acted on.
Evidence that can leave the room
Audit-ready packages (eval factsheets, or the format your model-risk function already uses) record who authored the rubric, what data was used, how agreement was measured, and what the release threshold was. If it cannot be shown, it is not evidence.
I write in the register the committee already reads. I do not invent a parallel vocabulary that only the vendor understands.
This is not a legal opinion and not an EU AI Act certificate. Counsel decides what the file means. The point of the work is that they have a scored catalog instead of a vendor narrative.
The file is yours. I do not keep a parallel copy of findings, and I do not sell a compliance badge. If a committee asks “who wrote the rubric, and on which data?”, the factsheet answers in a sentence.
What the file contains
For internal risk, legal, and external review when required.
| Artifact | Reader |
|---|---|
| Failure catalog | Engineering and security |
| Policy adherence scores | Product and compliance |
| Fairness and slice report | Risk and the business owner |
| Eval factsheet | Model risk, audit, regulator |
How I usually run it
I scope the work to policy, leakage, and escalation, once a first evaluation exists. My access is scoped and time-boxed.