martincousseau.com

Method #7, Risk & compliance, 2 min read

Risk & compliance

Red-teaming, policy adherence, fairness, and evidence for the committee that asks.

IN SHORT

Evidence your risk team can read, not a badge.

Dithered 1-bit shield illustration for the risk and compliance guide
PLATE 1. Generated plate · 1-bit · shield

01 · In brief

#

What it is

Some failures are not quality in the ordinary sense. They are policy, leakage, or a sentence you cannot defend.

This is for teams whose first evaluation already exists and who now need a scored red team or a policy file.

02 · Adversarial

#

Red-teaming with a catalog

Adversarial tests are designed from the failure-mode map: jailbreaks, leakage, disallowed advice, prompt injection, retrieval of the wrong document. Findings are a catalog with patches, not a slide of scary examples. A focused red team is this work in concentrated form.

I re-test after the patch. A red team that does not return is a performance. The catalog is versioned with the system it was run against.

I score policy, leakage, and the cases that would fail a release, against the same rubric the rest of the evaluation uses. A PDF of jailbreaks without a threshold is a scare file.

03 · Policy

#

Policy, fairness, safety

Instruction adherence is scored against system, policy, and format constraints. Bias and demographic performance are sliced, not averaged. Content safety is a dimension with a threshold, including “zero critical” where that is the only acceptable line.

Operational risk (policy, leakage, escalation) is this family of failures. A high faithfulness score does not excuse a single critical leak.

Fairness work is sliced evidence, not a slogan. I report the slice that fails and the sample it rests on. I do not publish a single “bias score” that cannot be acted on.

04 · Evidence

#

Evidence that can leave the room

Audit-ready packages (eval factsheets, or the format your model-risk function already uses) record who authored the rubric, what data was used, how agreement was measured, and what the release threshold was. If it cannot be shown, it is not evidence.

I write in the register the committee already reads. I do not invent a parallel vocabulary that only the vendor understands.

This is not a legal opinion and not an EU AI Act certificate. Counsel decides what the file means. The point of the work is that they have a scored catalog instead of a vendor narrative.

The file is yours. I do not keep a parallel copy of findings, and I do not sell a compliance badge. If a committee asks “who wrote the rubric, and on which data?”, the factsheet answers in a sentence.

05 · Artifacts

#

What the file contains

For internal risk, legal, and external review when required.

ArtifactReader
Failure catalogEngineering and security
Policy adherence scoresProduct and compliance
Fairness and slice reportRisk and the business owner
Eval factsheetModel risk, audit, regulator

06 · Engagement

#

How I usually run it

I scope the work to policy, leakage, and escalation, once a first evaluation exists. My access is scoped and time-boxed.