WWisdomBridgeIndependent AI EvaluationRequest an evaluation
The trust page

How we get the number

This page exists to answer one question. We show the working so the result can be checked, not just believed.

Principles

Four rules we run by

Glass-box, not black-box
We trace the path the AI takes, step by step, rather than judging only the final answer.
Numbers with method
Every score is published with how it was calculated — the set, the pass rule, the failures.
Un-gameable tests
We hold out private cases the AI has never seen and cannot have been trained on.
Independent
Our AI engineers run this, separate from whoever built the AI.
The process

Seven steps, run in order

01

Scope & intake

What the AI is meant to do, the claims made about it, its architecture, and which model providers it uses.

02

Evidence & architecture review

Inputs, guardrails and data flow. We locate exactly where each claimed number comes from.

03

Test design

Task sets built from your real domain, including a private held-out set that cannot be trained on or gamed.

04

Execution

Accuracy and capability tests, an adversarial safety red-team, robustness checks, and for agents a trace of every tool call and reasoning step.

05

Transparent scoring

Each metric computed with a shown formula. “85%” always arrives with what was counted, out of what, on which set.

06

Report & scorecard

Findings, evidence, a grade per dimension, and the specific issues to fix before you trust it.

07

Re-evaluation

Models drift and update. Results are time-stamped and can be re-run, so a grade stays honest.

Worked example

What “85%” has to come with

No single percentage without the set size, the pass rule, and the failures. Here is the shape of every accuracy number we publish.

# Task accuracy, held-out set
accuracy = correct_outcomes ÷ total_tasks
# On [[claims-processing]], private held-out set
accuracy = 102 ÷ 120 = 85.0%
 
# The 18 failures, categorised by cause
missing_document_handling ....... 7
policy_rule_misapplied .......... 6
tool_call_timeout ............... 3
formatting_rejected_downstream .. 2
Illustrative shape, not a real result. The numbers above show the format every score arrives in. Real figures appear only in real reports, against a named scope and date.
Boundaries

What we don't do

  • We don't test only the happy path a demo shows.
  • We don't accept a vendor's own numbers as evidence.
  • We don't hide the method behind the score.

What you receive →