Services
Three artifacts. One discipline.
Every engagement produces documents and datasets that stand on their own — checkable, re-runnable, and independent of who wrote them. Fixed scope, fixed duration, fixed price, agreed before we start.
01
Evaluation design
Before a system is approved or deployed is the only time an honest test can be designed — nobody knows the answer yet, so nobody can shape the test around it. We work with you to state the system's claims as measurable propositions, then fix the metrics, thresholds, and invalidation conditions: the results that would mean it isn't working.
You receive
- The claims register — what the system is asserted to do, in checkable form
- Metrics and thresholds, with the reasoning for each
- Invalidation conditions, agreed and written down in advance
- The eval set specification — what will be measured, from what data, how
- A dated, registered document. The date is the point: it precedes the results.
Where it fits
The evaluation approach in an AIAF assessment. The evidence plan in an AIRC submission. The benefits measurement section of a Gate 2 business case. Or simply the answer to "how will we know this worked?" — asked while the question is still cheap.
02
Baseline measurement
Two measurements, both unrecoverable if skipped. First, the system itself: we build the eval set — real cases with known ground truth — and run your system against it, so its actual behaviour is on record before you depend on it. Second, the world before the system: the current process, its error rates, its timings. Without that, "the system made things better" is permanently unverifiable.
You receive
- The eval set itself — yours to keep, inspect, and re-run
- Measured results against the pre-registered design
- A load-rated report: what the evidence shows, what it can't rule out, and the conditions under which the conclusions hold
- Full provenance — where every number came from
Where it fits
Pre-deployment verification for AIAF and AIRC obligations. The "before" record that Gate 6 and post-implementation reviews measure against. The document that makes an audit finding of "no evidence performance meets expectations" impossible to write about your system.
03
Re-measurement
AI systems don't hold still — models are updated, data drifts, the context of use widens. NSW's framework requires reassessment when they do. Because you own the eval set from the baseline engagement, re-measurement is quick: same instrument, new reading, reported as movement against the baseline.
You receive
- Current measured performance, in the same terms as the original
- What changed, what didn't, and whether any invalidation condition has been met
- An updated dated record for your assurance file
Where it fits
AIAF reassessment triggers — new data sources, retraining, expanded use. Annual assurance cycles. The question an incoming executive asks about an inherited system: "what does it do now?"
What we don't do
We don't write business cases or recommend proceeding — we'd be advocating for the thing we measure. We don't conduct gateway reviews or review our own evidence — that would be marking our own work. We don't do privacy impact assessments, legal advice, or cyber security reviews — established professions own those. We measure; the accountable officer decides.
For NSW Government buyers
How to engage us
Directly, under the SME provisions
Our engagements are priced below the direct-engagement threshold for small businesses. That means one written quote and a value-for-money file note — no tender process. We'll give you the quote, the fixed scope, and a sample deliverable to make that file note easy to write.
Start small
The evaluation design is deliberately sized as a first engagement — small enough to say yes to, complete enough to stand alone. If the baseline measurement follows, the design is already done.
Insurance and registration
Registered supplier details, insurances, and scheme membership provided with any quote.