How we get the number
This page exists to answer one question. We show the working so the result can be checked, not just believed.
Four rules we run by
Seven steps, run in order
Scope & intake
What the AI is meant to do, the claims made about it, its architecture, and which model providers it uses.
Evidence & architecture review
Inputs, guardrails and data flow. We locate exactly where each claimed number comes from.
Test design
Task sets built from your real domain, including a private held-out set that cannot be trained on or gamed.
Execution
Accuracy and capability tests, an adversarial safety red-team, robustness checks, and for agents a trace of every tool call and reasoning step.
Transparent scoring
Each metric computed with a shown formula. “85%” always arrives with what was counted, out of what, on which set.
Report & scorecard
Findings, evidence, a grade per dimension, and the specific issues to fix before you trust it.
Re-evaluation
Models drift and update. Results are time-stamped and can be re-run, so a grade stays honest.
What “85%” has to come with
No single percentage without the set size, the pass rule, and the failures. Here is the shape of every accuracy number we publish.
What we don't do
- We don't test only the happy path a demo shows.
- We don't accept a vendor's own numbers as evidence.
- We don't hide the method behind the score.