Oversight Theater

Your dashboard is green because someone is holding it green

There is a version of AI oversight that every organization can describe and almost none can evidence.

It sounds like this: we have humans in the loop. There is review. Someone signs off. The dashboard is green.

The problem is that all four of those statements can be true while the oversight they describe is decorative. Not fraudulent, not negligent, just unmeasured, and therefore unable to answer the only question that eventually gets asked, which is: who actually decided this, and would anyone have caught it if it were wrong?

The record nobody keeps

Ron Bodkin described the shape of the failure on an earlier episode of this show as the moral crumple zone. Picture a fraud investigator who can clear ten percent of the flags a system generates. He works the queue honestly. Then fraud surfaces in the ninety percent he never reached, and the blame lands on him. He was never empowered to stop an unsafe release before it shipped. He was empowered to look busy in front of a volume he could not absorb.

Patrick Boogaerts, a solutions architect who has spent about a dozen years in the layer between the data a company collects and the decisions it actually makes, reframed that problem in a way I have not stopped thinking about.

The investigator is not failing. He is doing exactly what the system was designed to let him do. The missing infrastructure is the record of what he did not get to.

Most organizations can tell you what was reviewed. Very few can tell you what was auto-approved, by what role, and on whose authority.

Sit with that distinction, because it is the whole thing. Every organization has a record of its human oversight. Almost none has a record of its absence of human oversight. And the second record is the one that determines liability, because it describes the decisions that went out into the world with nobody's judgment attached.

Once you can see it, as Patrick put it, you are no longer arguing about blame. You are looking at a queue.

Confidence is the failure mode, not error

The intuition most leaders carry is that the danger of an AI system is that it will be wrong.

The danger is that it will be confidently wrong.

A system that says it does not know gets escalated. A system that produces a clean, plausible reason gets approved, and nobody goes looking. Wrong gets caught. Confidently wrong gets forwarded.

This is why explainability, as commonly implemented, can make things worse rather than better. A model that emits a fluent justification for its output has not become more accountable. It has become better at ending the conversation. The justification is doing the work that scrutiny used to do.

Patrick's test for whether a screen is real is deceptively simple: what happens on the second question? Plenty of dashboards are built to look right on the first question and fall apart on the second. Executives learn this fast, which is why they stop trusting the screen and start calling the analyst directly.

He described the same pattern from a large enterprise OKR build: the hard problem was never technical. Anything could be rolled up. The failure was that by the time a metric reached the executive layer it had been aggregated so many times that nobody in the room could say what was inside it. A number everyone accepts and nobody can act on.

What fixed it was keeping one path from every top-level number back down to something a person owned. Not a full drill-down on every metric. One honest trail.

The number that cannot be performed

Patrick built a metric he calls the Human Intervention Rate: how often a person has to step in and correct something the system was supposed to handle on its own.

What makes it useful is not the math. It is that it counts work nobody was writing down. Every one of those corrections was already happening. They just were not showing up anywhere a leader could see them.

This is the property that matters. An organization can assert that it has human oversight. It cannot assert a Human Intervention Rate. Either the corrections are being counted or they are not, and a number that has to be produced from records is a number that cannot be performed in a meeting.

Interpreting it takes more care than most people expect. A rising HIR does not automatically mean something is broken. It can also mean automation was pointed at a problem that genuinely requires judgment, and people are correctly stepping in. Both look identical in the aggregate.

What separates them is whether the corrections cluster. Scattered interventions usually mean the work is just hard. Interventions piling up in the same place week after week mean the system has quietly stopped being able to do that job, and someone is absorbing the difference.

That is the fire. And it is usually burning in someone's workload.

What Monday morning actually looks like

Suppose the dashboard is green and the intervention data says a team is quietly fixing the machine all day.

Patrick's answer was not to touch the dashboard. It was to go ask the three or four people doing the correcting what they are actually fixing, because they already know, and in his experience they have usually told someone already. The information almost always exists in the organization before it exists in the system.

Then comes a decision, and leaders should name it out loud. Either fix the process, or accept that this is now a staffed manual step and resource it accordingly. Both are defensible.

What is not defensible is the third option, which is what usually happens: leave it unnamed and let it run on goodwill.

That is oversight theater. The dashboard stays green because people are holding it green. Leadership reads the color as proof the system works, when the color is actually evidence of how much invisible labor is propping it up. And it holds right up until the person doing the propping goes on vacation.

Why this is about to stop being a philosophy question

For anyone operating in a regulated decisioning environment, this argument is about to be presented back to them in the form of a questionnaire.

State insurance regulators are building the examination instrument ahead of any settled definition of what compliance requires. The NAIC's AI Systems Evaluation Tool went into a multistate pilot this year, with adoption anticipated at the Fall National Meeting. It asks carriers to inventory where AI is deployed, to identify their high-risk systems, and to describe their governance in a structured format the regulator designed.

An examiner working from that instrument is not asking whether you have oversight. They are asking you to show the record. Which decisions were reviewed, which were auto-approved, by what role, on whose authority, and what happened when the model was wrong.

A governance program that cannot produce those records has a policy. It does not have a control. And the distinction between the two is invisible right up until the moment someone asks.

The question to ask before you approve anything

Patrick's was the best single test I have heard: who would have caught this if it were wrong?

If the answer is nobody, you are not approving a recommendation. You are absorbing the risk.

The related one is what organizations mean when they say they trust the AI. Usually they mean they have stopped checking. Real trust is earned on a schedule.

And the belief worth retiring is that oversight slows you down. Every fast organization Patrick has worked in was fast because it caught things early, not because it checked less.

The organizations that will win with AI are not the ones with the smartest models. They are the ones that know exactly where their people are still holding the thing together.

Patrick Boogaerts appeared on Episode 6 of the Good Faith EI podcast. The Human Intervention Rate is his framework.

Previous
Previous

The Exam Is Being Written Before the Rule Is

Next
Next

The AI Governance Gap: 14 Companies, One Blind Spot