People expect a model audit: weights, bias metrics, red-team results, a note about hallucination rates. Those matter, and plenty of firms do them well. They also answer a question almost nobody at board level is actually asking, which is why the resulting document tends to be filed rather than used.
The question being asked is simpler and harder. We put AI into this process. What happened to the people?
One: what the deployment was optimised for
Every deployment has an objective function, whether or not anyone wrote it down. Sometimes it is in the vendor contract, sometimes in the success metric on the slide, sometimes only in what the team gets praised for. We find it and read it back to you in plain words.
This is usually the moment the room goes quiet. A deployment sold internally as freeing clinicians from admin turns out to be measured on appointments per hour, which is a different thing pointed in a different direction.
A technology does what it is measured on, not what it was announced as.
Two: the Human Return baseline
Before you can say a deployment gave something back, you need to know what was there before. The baseline covers four things we can actually count: hours, skill, human contact, and agency.
- —Hours. Not hours saved — hours received. Where did the time go? Into learning, care, rest, seeing people? Or straight back into throughput, in which case the saving belongs to the organisation and the staff got a busier day.
- —Skill. Which capabilities has the system taken over, and are the humans meant to oversee it still practising them? Delegation without atrophy is a design requirement, not a hope.
- —Human contact. How many conversations, visits and relationships existed before, and how many exist now? For anything sitting between people, this is the metric — not engagement.
- —Agency. When the system is wrong about somebody's money, health, grade or livelihood, who can reach a person, and can that person actually overrule it?
Those four become the net human score: one figure for what the deployment gives back minus what it takes. It is not a survey result. It is calculated from evidence and verified, because a self-declared number is a marketing asset rather than a governance one.
Three: the eight Guardrails, with evidence
Each Guardrail names the evidence it requires, so the assessment is a document review rather than an opinion. What we ask for is mostly unglamorous and already exists.
If a Guardrail cannot be evidenced, we say so and write down what would satisfy it. An audit that produces nothing failable is a brochure.
Four: what the board gets
A short document a non-technical director can read in ten minutes: the net human score with its workings, the Guardrails met and unmet, the EU AI Act Article 50 position, and a small number of changes ranked by how much they move the score. No maturity matrix, no heat map with forty amber squares.
The findings are usually cheaper to act on than people fear. The expensive discovery is not a broken model — it is a metric aimed at the wrong thing, which costs a sentence to change and eighteen months of trust to recover from.
What it is not
It is not a certification, though the same assessment sits underneath one. It is not a compliance sign-off, though it will tell you where you stand on disclosure and human oversight. And it is not an argument for deploying less AI. Most of the organisations we work with are deploying more of it afterwards, with the ROI and the human goal written on the same page.
Scope, timetable and cost are agreed in writing before an engagement begins — they depend on how many deployments are in scope and how much evidence already exists.