Writing · July 2026

What "independent evaluation required" actually obliges you to produce

Your AI use case has been assessed, and somewhere in the output is an evaluation obligation. Here is what discharges it — and what only looks like it does.

The situation

NSW's assurance machinery has become efficient at telling agencies what they owe. An AI Assessment Framework result arrives with triggered assurance activities. A review committee asks what evidence supports the use case. A gateway review wants to know how benefits will be measured. Someone forwards the obligation to you with a note that says: can you get this ready.

The paperwork is straightforward. The evidence behind it is the part almost nobody has — and the gap between the two is now visible, because the obligations are documented, dated, and increasingly followed up. Recent Auditor-General reports have shown what the follow-up looks like: findings that a system's performance targets were never set, that its accuracy was last measured years ago, that in the absence of defined success outcomes there is no evidence performance is meeting expectations. Those sentences get written when the obligation was discharged with documents instead of measurements.

What the obligation actually asks for

Strip the terminology and every evaluation obligation reduces to one question: can you show what the system actually does? Not what the vendor says it does, not what the business case hopes it does — what it measurably does, on your cases, in your context. Evidence that answers this has a recognisable anatomy:

Claims, stated so they can fail. "The assistant improves the resident experience" is not evaluable. "The assistant answers common enquiries at least as accurately as the current channel, and never gives incorrect answers about payments or deadlines" can be checked — and could come back false. If none of your success criteria could possibly come back false, you don't yet have any.

Thresholds fixed before results exist. A target set after the measurement is a description, not a test. The date on your evaluation design matters more than almost anything in it: it needs to precede the results, provably. This is also what protects you — nobody can accuse a pre-registered test of being shaped around the answer.

A test set with traceable ground truth. Real cases, drawn from your actual demand, with correct answers adjudicated by people qualified to adjudicate them — and a record of who decided what. A ground truth nobody can trace is an opinion with a spreadsheet.

A baseline of the world before the system. Every benefit claim is a comparison. If the current process's accuracy, volumes, and timings aren't measured before go-live, "the system made things better" becomes permanently unverifiable — this is the one measurement that cannot be reconstructed later.

A version stamp. AI systems change: models update, data drifts, scope widens. A result that doesn't record exactly what was measured — model, configuration, knowledge snapshot, date — can't support any claim about the system you're running today. This is also why evaluation is recurring, not a document you file once.

Independence. "Independent evaluation" means the measuring party has no stake in the result. The project team evaluating its own system, or the vendor benchmarking its own product, fails that test by construction — however honest everyone is. Independence isn't an accusation; it's what makes the evidence usable by the people above you.

What doesn't discharge it

Several artifacts routinely get filed against evaluation obligations and survive precisely until someone checks. Vendor benchmark results: measured on the vendor's cases, by the vendor, for the vendor. User acceptance testing: establishes the system runs, not that it's right. A successful demo: a demo is a performance, in both senses. Satisfaction surveys: people can be satisfied with confident wrong answers — that's rather the problem. A completed assessment form: the form is where the obligation is recorded, not where it's met.

None of these are worthless. They're just answers to different questions, and the eventual reader — a review committee, an incoming executive, an auditor — is asking this one: what does the system actually do?

If you're the person this landed on

Three things worth doing regardless of who does the work. First, write down the system's claims as falsifiable statements — this costs an afternoon and immediately reveals how much of the business case is measurable. Second, protect the baseline: if the system isn't live yet, measure the current state now, because that window closes at go-live and never reopens. Third, date everything — whatever evaluation design exists, fix it and version it before results arrive.

If you want to see what the full artifact looks like, we publish a complete worked example of an evaluation design — the claims register, thresholds, invalidation conditions, and test set specification for a fictional agency chatbot. It's the document this obligation is asking you to have.

Have a system that needs evidence?

Thirty minutes, no obligation. We'll tell you if the work isn't needed.

Book a scoping call