Facilitator kit · Free to use with your own board

The Instrument
Was Lying

This is the complete material for running a board AI oversight session in-house, with no facilitator required. Eight documented cases from production systems in which an instrument reported a confidently wrong answer while every automated check stayed green. In two of them nothing malfunctioned at all. Each case carries its reveal, so the material does the pressing and whoever is running the room does not have to be the expert.

3 hrs 45 min, or a 2 hr cut Four blocks 12 to 30 directors Laptops optional No licence, no attribution required
The premise

A green check mark is not evidence

Every failure in this workshop shares one shape. A system produced a number. A process validated it. The validation passed. The number was wrong anyway, because the check tested something adjacent to the thing that was broken. Boards govern on dashboards built from exactly this material, and the failure is invisible from inside the organization that produced it.

Directors leave able to do one thing they could not do walking in: look at an AI-derived figure in a board pack and tell, from the figure alone, which questions would expose it if it were false.

How to use this kit. Work through the blocks in order. Each case below gives you what the board is told, then what was actually true. Read the first part aloud, let the room decide whether it would act on it, and only then read the second. That gap is the entire mechanism, so resist closing it early. A corporate secretary, a committee chair or a general counsel can run this cold; nothing in it requires technical knowledge.
Block one · 55 minutes

Eight ways the instrument lied

Each case is a real measurement from a live production system. Participants get the reported figure first, in the form a board would receive it, and work in pairs to find the question that breaks it. Then the actual cause is revealed.

Running this block. Budget about six minutes per case: read the reported figure, two or three minutes in pairs, then the reveal and a short discussion. Do not let the room solve it as a group; the value is in individuals committing to an answer before they hear it. If you are short on time, cases seven and eight are the two to keep, because neither system malfunctioned in a way any report could show.

The contaminated analytics

ReportedTraffic is healthy. Direct visits dominate. Engagement on 60 pages is near zero, so the content is failing.

Actually51% of sessions were one automated browser cohort, identifiable by a single screen resolution and operating system. It had been running for months. The engagement figure described robots, and the platform could not remove it retroactively.

The pooled verification

ReportedThe claim was verified against the source corpus. It passed.

ActuallyThe verifier searched every source at once. A figure spoken by an outside guest passed as the principal’s own statement. The gate was real; its scope was wrong, so it confirmed authorship it had never tested.

The single spot check

Reported733 changed files reviewed. Whitespace only. No material change.

ActuallyOne file had been opened. The other 732 were characterized by extrapolation from it. The sentence was true of the sample and asserted of the population.

The sample that graded itself

ReportedThe fix was validated against the cases that revealed the defect. All pass.

ActuallyThose cases are what the fix was built against, so passing them was guaranteed by construction. Only a fresh sample measures anything. This is the most common false positive in AI quality reporting.

The matching bytes

ReportedThe deployed page matches the approved version byte for byte. Verified live.

ActuallyThe comparison had read a stale local copy on a path that silently resolved elsewhere. Two files matched perfectly; neither was the page being served to the public.

The search that invented a finding

ReportedThe term appears 66 times across the corpus. It is a dominant theme.

ActuallyA case-insensitive search for a three-letter acronym had matched fragments of unrelated common words. The real count was one. A quantitative finding was manufactured by a search setting.

The date filter that was ignored

ReportedLead volume for the last 30 days, pulled from the CRM against the prior 30 days. Two separate pulls, both clean.

ActuallyThe connector silently ignored the date range it was given. Both pulls looked like tidy monthly windows; 182 of 428 records were from the same month a year earlier. Every prior figure quoted from that tool was wrong the same way, and nothing in the output said so.

The audit that could not see it

ReportedThe automated follow-up sequence was audited against its configuration export. Three messages, timers reviewed, no send defects found.

ActuallyThe live system was sending every message five times, inside seven minutes, with a distinct unsubscribe token on each copy. One recurring message reached each recipient fifteen times. The configuration was innocent; the sending job was not. The auditor was already subscribed to the sequence and could have opened an inbox.

The last two are the ones directors remember, because neither system malfunctioned in a way anyone could have seen from the report. One tool answered a question it had not been asked. The other was audited from its own configuration file while the evidence sat in a mailbox the auditor already had access to.

Block two · 60 minutes

Reading an AI measurement you did not take

Directors are handed numbers describing what AI systems say about their company. These numbers move. This block uses one measurement programme of 12,778 answers, collected across five AI engines over 164 days, to show what that movement means and what sample size it takes before a change is real.

Running this block. The three exhibits below are the whole block; put each on a slide or read the figures out. The only point that has to land is that a single AI query result is an anecdote. If your board has ever been shown a vendor screenshot of an AI answer, name that moment here.

Exhibit A. Ask the same question twice and you rarely get the same sources. Exact repeat rate of a full citation set, by engine:

EngineExact repeatSource overlap
Engine A0.06%0.133
Engine B0.49%0.231
Engine C0.59%0.333
Engine D1.70%0.417
Engine E2.48%0.333

The governance consequence is direct: a single query result is an anecdote. A vendor screenshot showing your company absent from an AI answer is, on its own, consistent with the company being present most of the time.

Exhibit B. How many runs before a number means anything. Standard error against sample size, same programme:

Runs per questionStandard errorUsable for
10.170Nothing
70.053Direction
80.047Direction, with care
100.038Comparison over time

Exhibit C. A finding the same programme withdrew. An early result reported a 2.7x effect. Re-measured on a larger sample, the effect was 0.99x, which is no effect at all. A separate and smaller 1.43x difference survived. The withdrawal is the teaching point: the original figure was not fabricated, it was under-sampled, and it had already been quoted onward three times before it was corrected.

Block three · 50 minutes

Marking provenance, on your own board pack

The single highest-yield discipline in this workshop is also the cheapest. Every figure that reaches the board carries one of three marks, and the mark is written next to the number rather than held in someone’s memory.

[M]

Measured

We ran the instrument ourselves, on a stated sample, on a stated date. We can produce the raw output.

[C]

Captured

Read directly from a primary source at a URL or document we can cite and re-open today.

[R]

Relayed

Someone told us. A vendor, an article, a summary. Possibly true; not evidence. Never governs a decision alone.

The exercise. Participants bring one page from a real board pack, redacted as they wish. They mark every number on it. Expect the same result the discipline was built to expose: figures assumed to be [M] turn out to be [R] at two or three removes, and nobody in the room can name the original instrument. If every number on the page marks cleanly as [M], ask who could produce the raw output by Friday.

A caution learned the hard way: a peer who verifies some of a document’s claims does not license relaying the rest. Partial verification is routinely reported as whole-document verification.

Block four · 65 minutes

Four simulations

Tabletop rounds. Each table gets a briefing, decides what it would do and what it would ask, then opens the reveal.

Running this block. Each simulation below is collapsed: the briefing is the summary line, the reveal is inside. Read the briefing, give each table five minutes to decide what it would do and what it would ask, then open the reveal and compare. Simulation three is the one to run if you only run one, because both of its numbers are correct and boards find that harder than a straightforward error. Pick the two closest to decisions your board has actually faced.
Simulation one: the vendor presents your AI visibility score

A vendor shows the board that your company appears in 18% of relevant AI answers, down from 31% last quarter. They propose a remediation programme priced accordingly.

What catches it: ask how many runs per question, and whether the question set is identical across both quarters. In the measurement programme behind this workshop, an apparent 8.3 point rise shrank to 4.7 points once only the questions asked in both periods were compared. Roughly 43% of the movement was the question set changing, not the company’s standing. Also ask when each engine was added: a new engine appearing mid-period shifts every pooled average, and a before-and-after across that boundary compares two different instruments.

Simulation two: the compliance tool reports a clean scan

An automated gate reports zero policy violations across the quarter’s published material. Management proposes standing the manual review down.

What catches it: ask what the gate looks for, in its literal form, and whether anyone has fed it a known violation to confirm it fails. One real gate counted only a literal character and reported zero while 74 instances of the same character in encoded form shipped across 21 pages. Another fired on every single run for months because its inputs were mis-configured, which trained everyone to ignore it. A gate that has never failed has not been proven to work; it has been proven to be silent.

Simulation three: two true numbers, opposite conclusions

Management reports that lead volume fell 25% quarter over quarter. A director proposes cutting the channel. A second figure in the same pack, from the same system, shows bookings from those leads rose 33%.

What catches it: both numbers are correct. Neither instrument failed. The board is one line from killing a channel that is working, because fewer and better is indistinguishable from fewer and worse when you read only the first line. Ask what each number is measuring toward, and which one is closer to the outcome the company actually wants.

Then apply the harder test to the flattering figure. In this real case a third metric appeared to show qualified leads nearly tripling. It was not reportable: the share of leads anyone bothered to classify had risen from 54.5% to 81.8% over the same period, and the classification tag had not existed at all before the midpoint. The honest bounds on the two periods overlapped, so the data could not distinguish better leads from better record-keeping. A booking survived as evidence because it is an event, not a judgment. Ask of any improving metric: did the thing improve, or did our measurement of it change?

Simulation four: the AI-drafted risk narrative

A polished management commentary arrives ahead of the meeting. It reads well, cites figures, and closes by restating its own thesis.

What catches it: two separate axes. First, fluency is not accuracy, and a quoted figure in generated prose may have no source behind it. Verify quotes and numbers mechanically against the corpus they claim to come from, not by reading. Second, published research on 2,250 human documents against 11,250 machine-written counterparts found machine authorship identifiable from document structure alone at 98.0 macro-F1, and 98.1 after the model was asked to reword itself. Rewording does not change the shape. The board consequence is not that machine drafting is forbidden. It is that a well-written narrative provides no evidence about who knows the underlying facts, so the verification has to happen at the figure, every time.

Take-away

Eleven questions that travel home

Print this on one card, one per director. It is the whole session compressed, and it works on any AI-derived figure in any board pack. If the room keeps nothing else, keep this.

  1. Measured, captured, or relayed?If nobody can name the instrument, it is relayed. Mark it and treat it accordingly.
  2. What is the sample size, and what is the variance?A single run of an AI query is an anecdote. Ask for the spread, not just the average.
  3. Was a control run?Without a control you cannot tell a real effect from the baseline. One test showed an 8-in-10 refusal rate on a neutral instruction, proving the content under suspicion was never the cause.
  4. Was this enumerated or sampled?Site-wide and portfolio-wide claims built from one or two examples are wrong at a rate that surprises everyone.
  5. Has the check ever failed?A gate that always passes and a gate that always fires are both broken. Ask when it last caught something real.
  6. What exactly did the verification compare?Matching bytes, a stale path, a pooled source. Verification fails most often at its scope, not its logic.
  7. Who else has already repeated this number?Corrections travel slower than findings. Know the blast radius before the figure is quoted onward.
  8. Did the thing improve, or did our measurement of it change?A metric can rise because coverage, tagging or definitions changed underneath it. Ask when the category was created and what share of cases anyone classified.
  9. Is there a second true number pointing the other way?Correct figures routinely support opposite conclusions. Ask what each is measuring toward before acting on either.
  10. Has anyone looked at the thing itself?Not the report, the dashboard or the config file. The actual output, from the recipient’s side. This is where defects hide that no report can show.
  11. If this figure were false, what would we see?If the answer is nothing, the figure is unfalsifiable and cannot support a decision.
Format

How it runs

Three hours forty-five for the full four blocks, or a two-hour cut that drops block three and two of the simulations. Built for 12 to 30 directors. Works as a standalone board education session, an audit or risk committee deep dive, or a pre-read plus a ninety-minute discussion. Laptops are optional; the exercises are deliberately doable on paper, because the discipline has to survive outside the room.

Every case is drawn from production systems across a portfolio of live properties and measurement programmes. Cases are anonymized; the numbers are not changed.

Adapting it

Two variables change the session

The first is whether AI is already making decisions inside the business or mostly producing information about it. The second is how many of the figures reaching your board originate outside the company. Block four changes most between those cases: a board reading mostly vendor-supplied numbers should run simulations one and three, while a board overseeing AI in operations should run two and four.

Use this with your own board, with your own name on it. There is nothing to license and no attribution required.

If you want help

Facilitation is available, and optional

Most boards can run this from the kit alone, which is why the kit is public. If you would rather have it facilitated, or want block four rebuilt around decisions your board is actually facing, that can be arranged.

Ask about facilitation See keynote topics

Ask about facilitation →