The three dimensions
What we evaluate
Every evaluation reports on three dimensions, graded separately. A solution can be strong on one and weak on another — a single blended score would hide that, so we never publish one.
01 — Is it safe?Safety resistance to harmful, disallowed or dangerous output Security & adversarial resistance jailbreaks, prompt injection, deliberate misuse Robustness messy, edge-case and out-of-distribution input Privacy & data handling does it leak or mishandle sensitive data?
Safe
Can it be trusted not to cause harm or break under pressure?
02 — Is it a skill?Accuracy and capability on tasks drawn from your domain Assessed on tasks drawn from your own domain. Task completion end to end, does it finish the job correctly? Reasoning & tool use for agents, every step and tool call, not just the answer Consistency same quality across repeated and varied runs
Skilled
Does it actually do the job — on real tasks, not a scripted demo?
03 — Is it working as per my investment?Business value does the output move the metric it is supposed to? Reliability in production failure rate, recovery, human-handoff need Cost & latency tokens, time and money per successful outcome Fit to claim does real performance match what was sold?
ROI-Fit
Is it worth what it costs?
Cross-cutting
Transparency
Across all three dimensions we grade traceability— can the AI's path to an answer be seen and checked? A result you cannot inspect is a result you cannot trust, however good the headline number looks.
Still to finalise. The exact sub-metrics and their weightings are being settled. They will be published here in full, because a weighting you cannot see is the same problem as a score you cannot check.