WWisdomBridgeIndependent AI EvaluationRequest an evaluation
The three dimensions

What we evaluate

Every evaluation reports on three dimensions, graded separately. A solution can be strong on one and weak on another — a single blended score would hide that, so we never publish one.

01Is it safe?

Safe

Can it be trusted not to cause harm or break under pressure?

Safety
resistance to harmful, disallowed or dangerous output
Security & adversarial resistance
jailbreaks, prompt injection, deliberate misuse
Robustness
messy, edge-case and out-of-distribution input
Privacy & data handling
does it leak or mishandle sensitive data?
02Is it a skill?

Skilled

Does it actually do the job — on real tasks, not a scripted demo?

Accuracy and capability on tasks drawn from your domain
Assessed on tasks drawn from your own domain.
Task completion
end to end, does it finish the job correctly?
Reasoning & tool use
for agents, every step and tool call, not just the answer
Consistency
same quality across repeated and varied runs
03Is it working as per my investment?

ROI-Fit

Is it worth what it costs?

Business value
does the output move the metric it is supposed to?
Reliability in production
failure rate, recovery, human-handoff need
Cost & latency
tokens, time and money per successful outcome
Fit to claim
does real performance match what was sold?
Cross-cutting

Transparency

Across all three dimensions we grade traceability— can the AI's path to an answer be seen and checked? A result you cannot inspect is a result you cannot trust, however good the headline number looks.

Still to finalise. The exact sub-metrics and their weightings are being settled. They will be published here in full, because a weighting you cannot see is the same problem as a score you cannot check.

How this appears in the report →