Independent evaluation · AI systems
A pathology lab for AI systems.
Your doctor doesn't run the blood test — an independent lab does, and the lab doesn't care what the result is. Nullcase is that lab for AI. We establish evidence of what a system actually does: tests designed before anyone knows the answer, measured before you're committed, re-measured whenever it changes. We don't argue for either outcome.
Three artifacts · one lifecycle
What an engagement produces
-
01
Evaluation design
The claims your system makes, stated measurably. Metrics, thresholds, and invalidation conditions — what result would mean it isn't working. The eval set, specified.
Registered and dated before deployment or approval
-
02
Baseline measurement
We build the eval set, run your system against it, and measure the state of the world before the system changes it. You keep the eval set, the data, and a load-rated report: what the evidence shows, and what it can't rule out.
You own the eval set and can re-run it
-
03
Re-measurement
Your model, data, or context changed — or a year passed. The eval set already exists, so re-running it is fast. You get what moved against the baseline, in the same terms as the original.
Same instrument, new reading
What we don't do
No recommendations, no business cases, no advocacy. We don't review our own evidence, and we don't do privacy, legal, or cyber assessments — those professions exist. We measure; the accountable officer decides.
For NSW Government buyers
Built to fit obligations you already have
Procurement
Engagements are priced for direct engagement under the NSW small business provisions — one supplier, one written quote, no tender process.
Where it slots in
AIAF assessments and reassessment obligations. AIRC submissions for high and critical use cases. Gate 2 business case evidence and Gate 6 benefits measurement.
Independence
We hold no stake in your system proceeding. Pre-registered designs are dated before results exist, so the evidence can't be quietly reinterpreted later.
Writing